why aren't all 12 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Chen: GPT-4.5 performance jump matches leap from GPT-3.5 to GPT-4
“It signifies an order of magnitude improvement over the last models, kind of commensurate with the jump from 3.5 to four.”
Insight
Chen: AI models cannot learn reasoning from scratch without pre-trained knowledge
“You need knowledge in order to build reasoning on top of it.
Right.
a model can't kind of go in blind and just learn reasoning from scratch.
So we find these two paradigms to be fairly complementary and we think, you know, they have feedback loops on each oth…”
Assertion Not checkable as stated
Chen: GPT-4.5 scaling returns remain consistent with OpenAI's prior projections
“You know, we are seeing the same returns, and I do want to stress that GPT-D 4.5 is that next point on this unsupervised learning paradigm, and, you know, we're very rigorous about how we do this. We make projections based on all the models we've trained befor…”
Assertion Not checkable as stated
Chen: GPT-4.5 hits expected benchmark progression consistent with OpenAI's trajectory
“Well, I really don't think that the accurate characterization is that it doesn't hit the benchmarks that, that we expect it to. So when you look at kind of the development of three to 3.5 to four to 4.5 this does hit the benchmarks that we expect.”
Disclosure
Chen: Focus on reasoning models caused the longer gap before GPT-4.5
“Why there seems to be, you know, a little bit bigger of a gap in release time between four and 4.5, we've been really largely focused on developing the reasoning parallel paradigm as well.”
Prediction Not checkable as stated
Chen: GPT-5 could combine unsupervised scaling with reasoning paradigms
“And so I think, like GPT-V really could be the culmination of a lot of these things coming together.”
Assertion Partly supported
Chen: Users prefer GPT-4.5 over GPT-4o by 60% to 70% margins
“When we look at, kind of, comparisons against GPT-FORO you'll see that everyday use cases, people prefer, you know, by a margin of 60% for actually productivity and knowledge work against GPT-FORO, there's almost like a 70% preference rate.”
Assertion Not checkable as stated
Chen: Nearly all large language models today utilize mixture of experts
“I think pretty much all large language models today use, utilize mixture of experts.”
Opinion
Chen: GPT-4.5 outshines reasoning models like o1 in creative writing
“And, you know, we find that in a lot of areas like creative writing, for instance
Again, this is stuff that we want to test over the next one or two months but we find that there are areas like creative writing where this model outshines reasoning models.”
Assertion Not checkable as stated
Chen: Pausing and restarting training runs is standard across OpenAI models
“Actually, so I think it's interesting that this gets is a point that's attributed to this model because actually in, in, in developing all of our foundation models, right they're all experiments, right? I think you know, running all of the foundation models of…”
Assertion Not checkable as stated
Chen: GPT-4.5 creates ASCII art almost flawlessly, unlike previous models
“If you ask any of the previous models to create ASCII art for you, right? Actually, they mostly just fall down. This one can do it Almost flawless.”
Assertion Not checkable as stated
Chen: OpenAI inference costs dropped orders of magnitude since GPT-4
“The costs have dropped, you know, many orders of magnitude since we first launched GPT-IV.”