why aren't all 21 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Marcus: Models like o1 are not systematically better than GPT-4
“Models like O-I are not systematically better than GPT-IV. There's, they're better in certain use cases. Especially ones where you can create data in advance.”
Prediction Held up
Suleyman: GPT-4 and GPT-4o Efficiency Will Improve 100x
“So I expect that to happen for GPT-IV, GPT-IV-O and all of the other models down the road.”
Opinion
Hendrycks: The next-token prediction paradigm in AI is running out of steam
“That sort of paradigm does seem like it's running out of steam. It has held for many, many orders of magnitude but the returns on doing that are lower.”
Assertion Not checkable as stated
Chen: GPT-4.5 performance jump matches leap from GPT-3.5 to GPT-4
“It signifies an order of magnitude improvement over the last models, kind of commensurate with the jump from 3.5 to four.”
Assertion Not checkable as stated
Chen: GPT-4.5 hits expected benchmark progression consistent with OpenAI's trajectory
“Well, I really don't think that the accurate characterization is that it doesn't hit the benchmarks that, that we expect it to. So when you look at kind of the development of three to 3.5 to four to 4.5 this does hit the benchmarks that we expect.”
Opinion
Saunders: GPT-4 is safe, but GPT-5 or later might be the Titanic
“I don't think that I was working on the Titanic. I don't think that GPT four was the Titanic. I'm more, I'm afraid that like GPT five or GPT six or GPT seven might be the Titanic in, in, in this analogy.”
Insight
If AI scaling stalls, it will likely be due to data memorization
“If, let's say, GPT-Six isn't that much better than GPT-Four, and you had to look back on it and say, like, why, why did that happen? I think the most, the thing I'd expect to say is that right now we are We are kind of fooled by how much data these models cons…”
Opinion
AI scaling will not mysteriously halt halfway through the range of human intelligence
“It would just be bizarre to me that, like, you're halfway through the human range of intelligence, and now it stops getting better, so I do sympathize with Sam's statement in the sense of, like, why would it stop here, right? If it was gonna stop, why, it woul…”
Prediction Held up
Srinivas: Anthropic will release a model better than GPT-4 in 2024
“I actually think they will end up creating a model better than GPT-IV this year. Like, it, it's sort of almost guaranteed to happen. So I believe it's going to happen with CLAW-III.”
Assertion Supported
Ramaswamy: GPT-4 and Claude are a clear step ahead of open source models
“The blunt truth is that the very best of the models out there, whether it's GPT-IV or Claude's biggest model, are a clear step ahead of the pack when it comes to quality. When it comes to reasoning, when it comes to the quality of the text that they produce th…”
Assertion Supported
Lightcap: GPT-5 bakes in tool use and longer-horizon reasoning
“So using tools, for example, is something that really thinks really important for overall intelligence, GPT two and three couldn't really do that as well. GPT-IV could do it in a more nascent way. And now GPT-V, you get that baked in with the benefit of these …”
Assertion Supported
Patel: GPT-4-Level Training Costs Have Dropped 10x to 100x
“If you look at what it costs to train GBT for originally, I think it was like 20,008, 100 over the course of a hundred days. So I think it costs on the order of like half a million to a hundred million dollars, somewhere in that range. And I think you could tr…”
Assertion Not checkable as stated
Patel: Frontier AI cluster costs have scaled from $100M to $10B
“For GPT-IV, it was a few hundred million dollars and it's one building full of GPUs, too. GPT-IV 4.5 and the reasoning models, like, oh, one, oh, three were done in a, in three buildings on the same site, and, you know, billions of dollars to, hey, these next …”
Assertion Supported
Patel: Model inference costs dropped 60x from GPT-4 to DeepSeek-V3
“And likewise, when we look at from GPT-IV to DeepSeq VIII it's fallen roughly 600 X in cost. Right. So we're not quite at that 1200 X, but it has fallen 600 X in cost from 60 dollars to less than you know, to about a dollar. Right. Or to less than a dollar. So…”
Disclosure
Chen: Focus on reasoning models caused the longer gap before GPT-4.5
“Why there seems to be, you know, a little bit bigger of a gap in release time between four and 4.5, we've been really largely focused on developing the reasoning parallel paradigm as well.”
Opinion
Claude 3 and Gemini are not significantly better than GPT-4
“So we've gotten Claude III, we've gotten Gemini. They're not significantly better, if at all, than GPT-IV, and certainly not the newer version of GPT-IV.”
Opinion
Doshi: Image AI is probably a couple years behind a GPT-4 moment
“Images is interesting because it's probably a couple years behind a GPT-IV true moment, right?”
Assertion Not checkable as stated
Chen: Pausing and restarting training runs is standard across OpenAI models
“Actually, so I think it's interesting that this gets is a point that's attributed to this model because actually in, in, in developing all of our foundation models, right they're all experiments, right? I think you know, running all of the foundation models of…”
Prediction Not checkable as stated
Levie: 10x cheaper AI models will drive 100x more usage
“If you could make, again, kind of wave magic wand and you say like we have GPT five or GPT six, and it costs like a 10th
Of what today GPT-IV costs I would argue that you'll probably get a hundred X more usage of AI, not, not just 10 X, you know, more usage of…”
Insight
Multi-hour autonomous coherency is a major bottleneck for agentic AI
“One big one, and this is similar, is that they aren't yet useful in long when you need them to kind of go do a job. You can't be like GPT-IV. I'll be back in a while, but can you like manage my inbox for me in the meantime? Or can you go go book a trip for me?…”
Assertion Not checkable as stated
Chen: OpenAI inference costs dropped orders of magnitude since GPT-4
“The costs have dropped, you know, many orders of magnitude since we first launched GPT-IV.”