why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Chen: OpenAI has internal models matching Gemini 3 and better successors coming
“Just looking purely at the benchmarks, you know, we actually felt quite confident you know, we have models internally that perform at the level of Gemini three, and we're pretty confident that we will release them soon, and we can release successor models that…”
Insight
Chen: AI reasoning only emerges at large scale requiring massive compute
“When you look at reasoning you just don't see that happen at small scale, right? There's like a certain scale at which it starts becoming signal bearing and that requires you to have resources, right?”
Prediction Not checkable as stated
Chen: 2025 will be the year of autonomous AI agents
“We see 25 as this year of agents, right? We think of it as a year where models are going to do a lot more autonomous work. You can let them Kind of be unsupervised for much longer periods of time.”
Assertion Supported
Chen: Frontier AI models consistently score over 90% on AIME
“I think one clear example here is the Amy, like probably the hardest auto gradable, like human math eval, at least in the US. And yeah, the models are consistently getting like 90 plus percent on these.”
Insight
Chen: AI reasoning is the key mechanism to make agents reliable
“And I think the reason why we care so much about reasoning is because I think that's the path that we get reliable agents through.”
Assertion Supported
Chen: OpenAI reached 3 million paying business users
“We hit a big milestone. We got I think three million paying business users fairly recently.”
Insight
Mark Chen: Compute Can Scale Heavily into RL Given Right Levers
“I think, like, if you find the right levers, you can really pump a lot of compute into RL as well as pre-training.”