why aren't all 11 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Qiu: Pragmatic, smaller models will eventually address the majority of AI workflows
“And I suspect we're going to see something similar where a lot of use cases are going to be able to be addressed by something pretty pragmatic and relatively small.”
Insight
Qiu: Training a giant monolithic model does not magically solve agent reliability
“It's not like, oh, magical, you know, we train a giant model and stick everything into it and then magically it works. Like it does not work. It'll get better at random parts of the agent loop, but that's not what we want.”
Assertion Not checkable as stated
Qiu: Imbue trains state-of-the-art AI models with only 13 or 14 people
“We're kind of like training state, state of the art models with like 14, 13 people.”
Opinion
Qiu: Self-supervised AI may learn representations akin to human cognition
“There's something really interesting here where maybe machines are learning the same kinds of representations or similar representations to what humans are learning. And maybe they can get to a point where they can actually do the types of things that humans a…”
Prediction Not checkable as stated
Qiu: Software output will explode as programming democratizes to non-coders
“Software is just dramatically underwritten because it's so hard to write code today. So, you know, as we said in the future, like, computers will be able to be programmed by regular people. What that means is, like, we're gonna write way, way, way more softwar…”
Insight
Qiu: Autonomous AI agents represent a calculator-to-computer leap in technology
“The diff between this, where we are today, and that is kind of like the diff between the first calculator and where computers are today.”
Insight
Qiu: Chain of thought and tree of thought function as error correction
“Reasoning is one big piece of improving reliability, and second chunk of things is like all of this error correction, and I think like chain of thought, tree of thought, these are error correction techniques.”
Assertion Not checkable as stated
Qiu: The AI industry hasn't pushed data limits on small models
“We're definitely not pushing the bounds of what we can do with data today on small models, and so, you know, smaller things can work well.”
Insight
Qiu: Specialized reasoning planners coordinating sub-agents outperform single monolithic models
“You can kind of have this like more general reasoning layer and also a bunch of sub-agents where it, that, That general reasoning layer is actually very specific. It's a specific planner. It's not that good at, like, browsing the web and things like that, but …”
Insight
Qiu: Evaluating AI models solely on binary correctness loses critical evaluation data
“Part of why a lot of teams try to work on just math or code reasoning is because those are the easiest to evaluate and like the clearest Answers, but just relying on, like, is the output correct or not, that loses a lot of information in the evaluation.”
Insight
Qiu: Developing AI agents today is comparable to writing code in assembly
“So today, like writing agents feels like writing code in assembly, and that really limits the types of agents we can build and also limits the number of people who can build them.”