why aren't all 26 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Douglas: Anthropic believes AGI is reachable in a couple of years
“We think that, you know, AGI is within reach in the next couple of years.”
Prediction Open · timeframe Oct 2028
Douglas: AI industry will reach human-level computer capabilities in 2-3 years
“Which is that in the next two or three years, given the right feedback loops, given the right compute, given the right, you know, elbow grease and this kind of stuff, we think that we as the AI industry are all on track to create something that is at least as …”
Assertion Not checkable as stated
Douglas: Transformers successfully model any domain given sufficient data and compute
“I don't think that's true. I think we haven't yet really found anything that transformers haven't been able to model provided sufficient data and sufficient compute.”
Prediction Not checkable as stated
Douglas predicts DeepMind will lead the world in AI science discoveries
“DeepMind, if you wanted to solve science, is the best place in the world. Like, I think that DeepMind will directly contribute to more scientific discoveries from AI than anything else, right?”
Assertion Not checkable as stated
Douglas: An Anthropic AI agent operated autonomously for 30 hours building apps
“We asked it to build something that looks roughly like a chat app, you know, something like Slack or, you know. And it was, it, the model just worked for 30 hours. Like, it was just spinning there on a computer for 30 hours, and came out with a really good wor…”
Assertion Supported
Douglas: AI autonomous task execution time horizons double every six months
“And so I think it's like every couple of months, the time horizon that the AIs are capable of doing is doubling or something, something crazy. Maybe, maybe every six months the time horizon doubles”
Prediction Not checkable as stated
Douglas: AI application development will see another massive leap next year
“Over the next six months, over the next year, expect dramatic progress here. And like look at where we are now versus where we were a year ago. And the difference is I expect the same jump basically.”
Assertion Not checkable as stated
Douglas: AI coding interventions stem from taste, not raw programming capability
“Right now you need to intervene quite frequently, but it's usually on questions of taste rather than it is questions of, like, raw programming ability.”
Prediction Not checkable as stated
Douglas: AI beating GDP benchmarks won't immediately alter the broader economy
“We'll probably reach like better than human on the GDP eval, and it won't change anything economically because It'll be all the connective tissue, and all the, like, you know, the context, and actually, like, the task won't be representative.”
Prediction Not checkable as stated
Douglas: Individuals will manage 24/7 AI agent teams within two years
“If coding agents progress in the way I've been saying, in a year or two, you'll be able to manage a team, basically, that works 24 seven for you doing work.”
Assertion Not checkable as stated
Douglas: Robotic locomotion is essentially solved using basic reinforcement learning
“Locomotion's kind of solved, to be honest, with basic RL.”
Assertion Not checkable as stated
Douglas: The post-ChatGPT AI compute supercycle begins properly in 2025
“So finally, this year is where the compute, like, super cycle is, like, beginning properly in effect.”
Assertion Not checkable as stated
Douglas: Google LLM inference stack saved hundreds of millions in months
“This ended up saving several hundred million dollars, I think, like, even over the first six months”
Assertion Not checkable as stated
Douglas: Even AI pioneer Noam Shazeer only sees 10% of ideas work
“I once asked this question of Noam Chazier. And he was like, yeah, maybe like 10% of my ideas work, and that's not, right? You know, one of the, you know, an absolute genius, one of the best in the field. So if only 10% of his ideas work, then I think that, yo…”
Assertion Supported
Douglas: Cognition rebuilt Devin's architecture around Anthropic's highly useful Claude Sonnet
“The Cognition folks in Devon found the model so useful, they had to, like, rebuild their architecture around it.”
Prediction Not checkable as stated
Douglas: AI coding agent supervision will drop to 20-minute intervals within months
“Over time you know, over the next couple of months, you're probably gonna end up in a situation where you only need to supervise the models every, you know, 10 minutes, 20 minutes or so.”
Assertion Not checkable as stated
Douglas: The current generation of AI agents are astonishingly good at self-correcting
“I think one of the things that made me remarkable about the current generation of agents is that they can self-correct. In fact, they're astonishingly good at self-correcting.”
Assertion Contradicted
Douglas: Anthropic models autonomously replicated the Claude.ai website in hours
“And in this case, the model replicated Claude.ai with artifacts, with everything else I can't quite remember how long that one took. Maybe a couple hours to do.”
Assertion Partly supported
Douglas: AI progress on METR evaluations plots as a straight line
“Like on the meter eval, if you look at progress over the last two years, you can plot it with a straight line, right?”
Assertion Not checkable as stated
Douglas: Reinforcement learning on language models finally started working in late 2024
“I think also an important change in, you know, in this sort of like era of reasoning models and RL on language models is, at the end of last year, RL on language models finally started to work.”
Assertion Not checkable as stated
Douglas: OpenAI's o1 established test-time compute and RL as a scaling axis
“And I think OpenAI deserves a lot of credit for you know, releasing the first, like, serious RL plus LLMs release with O-one. And I think this really kicked off a pretty, you know, substantial change because it opened up a new axis of scaling, right? There was…”
Assertion Supported
Douglas: Anthropic's mid-tier Sonnet is smarter than its flagship Opus
“One of the interesting things about this most recent release is actually Sonnet is smarter than Opus.”
Assertion Not checkable as stated
Douglas: DeepMind has 1,000 on Gemini and 10,000 on foundational research
“But if you look, Gemini is like, you know, a thousand people, there's a, there's still like 10,000 plus people doing all kinds of really like longterm foundational research at DMI.”
Assertion Supported
Douglas: Sonnet 4.5 pushed SWE-bench scores from roughly 72% to 78%
“We moved recently from roughly 72 to roughly 78 in Sweepbench”
Assertion Partly supported
Douglas: The entire AI industry scored under 20% on SWE-bench last year
“As recently as a year ago, I think we were under 20% or something like that as a field.”
Assertion Supported
Douglas: Anthropic's Claude Opus 4.1 led OpenAI's GDP eval benchmark
“4.1 Opus was the leading model there.”