why aren't all 13 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Opinion
Masad: GPT-5 regressed in human tone compared to GPT-4
“My feeling is that you know, GPT-Five got good at verifiable domains. It didn't feel that much better at anything else. The more human angle of it felt like it regressed”
Assertion Not checkable as stated
Labenz: OpenAI's router failure caused bad initial GPT-5 outputs
“The problem at launch was that that router was broken. So all of the queries were going to the dumb model, and so a lot of people literally just got Bad outputs, which were worse than oh three because they were getting non thinking responses.”
Insight
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Opinion
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Opinion
Masad: GPT-5 shows no reasoning progress on open-ended controversial topics
“Go you know, dig up GPT-IV or other models and go to GPT-V. You're not gonna find that much difference of, okay, let's reason together. Let's try to figure out what was the origins of COVID. Because it's still an unanswered question, you know? And I don't see …”
Insight
Benchmark-topping AI models are not necessarily what consumers want for chat
“I don't necessarily think the like smartest model that scores the best on sort of all of these objective benchmarks of intelligence will be the model that people want to chat with.”
Assertion Not checkable as stated
Kim: GPT-5 internal testers felt insulted by instant answers to hard questions
“I think we hear this with GPT-Five internally when people are testing and they're like, oh, I thought I asked like a really hard question. I feel like a little bit insulted that I thought for like two seconds or like when it doesn't even want to think at all.”
Opinion
Kim: GPT-5 front-end coding is a massive leap over o3
“If you compare it to O three's front end coding capability, this is just totally next level.”
Assertion Supported
OpenAI's GPT-5 scored highest on the physician-trained HealthBench medical benchmark
“They talked about how GPT-V was kind of the highest scoring model on this thing called HealthBench, which is a benchmark they trained with like, 250 plus physicians. To measure how good an LLM is at answering medical questions.”
Prediction Not checkable as stated
Kim: AI prompt-based app generation will spur surge in indie businesses
“I think we're just gonna have a lot more, I would expect, like, maybe a lot more, like, indie type of, like, Businesses built around this because of the fact that, like, you just need to have the idea, write a simple prompt, and then you get the full fledged a…”
Opinion
Kim: GPT-5's creative writing capability is tender and touching
“That's one of my favorite improvements in GBT five. The writing, I honestly find it's very tender and touching, especially for a lot of the creative writing that we want to do.”
Disclosure
Kim: GPT-5 is a step change for personal coding and writing
“I use it for coding and writing all the time, and it's just a huge stuff change.”