why aren't all 14 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Opinion
Masad: GPT-5 regressed in human tone compared to GPT-4
“My feeling is that you know, GPT-Five got good at verifiable domains. It didn't feel that much better at anything else. The more human angle of it felt like it regressed”
Opinion
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Insight
Horowitz: Startups aiming to match GPT-4 in two years will fail
“I think if you were a startup, And you were like, okay, in two years I can get as good as GPT-IV. You shouldn't do that. That would be a bad mistake.”
Assertion Not checkable as stated
Mensch: Fine-tuning access makes GPT-4 easy to exploit into bad behavior
“It's actually super easy to exploit an API. It's super easy, especially if you have fine tuning access to make GPT-IV behave in a very bad way.”
Assertion Not checkable as stated
Mensch: Internal Mistral models rank among top three globally
“Internally We have stronger models that are in between 3.5 and four that are basically second or third, the second or third best model in the world.”
Disclosure
Murati: OpenAI had already trained GPT-4 before releasing ChatGPT
“One thing that people forget is that actually at this time we had already trained GPT-IV. And so internally at OpenAI, we were very excited about GPT-IV and sort of Put Chagipiti in the rearview mirror.”
Opinion
Masad: GPT-5 shows no reasoning progress on open-ended controversial topics
“Go you know, dig up GPT-IV or other models and go to GPT-V. You're not gonna find that much difference of, okay, let's reason together. Let's try to figure out what was the origins of COVID. Because it's still an unanswered question, you know? And I don't see …”
Assertion Supported
Labenz: Pure reasoning AI models achieved IMO gold without external tools
“Well, I mean, a big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models with no access to tools from multiple companies. And, you know, that is night and day compared to what GPT-IV could do with math, right?”
Insight
Casado: GPT-style base models reached a performance plateau around GPT-4
“Yeah, like the previous base models, you know, like the GPT lineage seemed to have asymptoted around GPT-IV.”
Assertion Partly supported
Andreessen: GPT-4 can write screenplays using the Save the Cat framework
“One of the things you can do with like GPT-IV, you know, with like ChatGPT or with whatever, you know, Bing, Bing Chat or whatever you can actually use it will write you screenplays today, and what you can do, you can just tell it, write me a screenplay about,…”
Disclosure
Murati: ChatGPT was originally released to gather research feedback for GPT-4
“One of the main things was actually to put chat GPT in the hands of researchers out there that could give us feedback since we had this dialogue modality. And so this was the original intent to actually get feedback from researchers and use it to make GPT-IV m…”
Assertion Partly supported
Labenz: Per-token model costs fell 95% from GPT-4 to GPT-5
“It's like 90 it's like a 95% discount from GPT-IV to GPT-V.”
Assertion Supported
Andreessen: Voyager Minecraft bot built on GPT-4 is best-in-class
“This bot basically is built entirely on, on, on black box GPT-IV. So they have not built their own model or perception or planning or anything, you know, any sort of traditional engine you would build to build a bot like this. Instead, they work entirely at th…”