Roy: GPT-5 felt like a massive dud compared to expectations
“Again, GPT-V, we all know, felt like a massive dud that was supposed to be that magic moment for everyone”
Frontier models are cheaper for agentic tasks because they require fewer turns
“Interestingly, in Tau Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the some of the GBD five, for instance got to the answer faster. And …”
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Fioca: OpenAI evaluates GPT-5 coding models on behavioral software engineering practices
“And so these are just best software engineering practices that turn out to be behavior characteristics, and we can measure the model's performance on those behaviors and grade it that way.”
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”
Kaiser: Pre-training between GPT-4 and GPT-5 focused on reducing costs
“The pre-training part in that timeframe was mostly about making things cheaper. Not making things better.”
Kaiser: OpenAI model names are now detached from technical milestones
“Now the naming is by capability, right? GPT-Five is a capable model. 5.1 is a more capable model. Mini is the smaller model that's slightly less capable, but faster and cheaper. And the thinking models are the ones that do more research, right? In that sense, …”
Altman: OpenAI will eventually build low-power devices running GPT-6 locally
“Someday we will make an incredible consumer device that can run a GPT-Five or GPT-Six capable model completely locally at a low power draw.”
Masad: GPT-5 regressed in human tone compared to GPT-4
“My feeling is that you know, GPT-Five got good at verifiable domains. It didn't feel that much better at anything else. The more human angle of it felt like it regressed”
Masad: GPT-5 shows no reasoning progress on open-ended controversial topics
“Go you know, dig up GPT-IV or other models and go to GPT-V. You're not gonna find that much difference of, okay, let's reason together. Let's try to figure out what was the origins of COVID. Because it's still an unanswered question, you know? And I don't see …”
Tworek: GPT-5 can effectively be considered an iteration like 'o3.1'
“Like GPT-Five in some way I can be considered as like, oh, 3.1. It's a little bit of like, you know, iteration of like the same thing and the same concept”
Isenberg: GenSpark simultaneously queries GPT-5, Claude Sonnet 4, and Gemini 2.5 Flash
“What it is doing is it is pinging GPT-Five. It is pinging Claude Sonet-Four. It is pinging Gemini and 2.5 flash. So instead of me having to go individually to those products, I have it right there.”
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Labenz: OpenAI's router failure caused bad initial GPT-5 outputs
“The problem at launch was that that router was broken. So all of the queries were going to the dumb model, and so a lot of people literally just got Bad outputs, which were worse than oh three because they were getting non thinking responses.”
OpenAI's latest model leveled Anthropic's dominant lead among Cursor developers
“GPT five, the model opening I released a month ago at this point, I believe two months ago, it really changed that. You know, people have talked to a cursor say that, you know, it completely Leveled the kind of usage between OpenAI and Anthropics models in a w…”
Patel: GPT-5 Is Roughly the Same Size and Cost as GPT-4o
“GPT-V. It's basically the same size as four O and roughly the same cost. That's actually a little bit cheaper potentially.”
Taneja: Newer AI model capabilities outweigh a year of GTM head start
“So the go to market advantage you may have created in a year having started on GPT versus five might be anemic compared to the technology advantage that you have. If you start in the GPT five or the choices you make and how fast you can move because the models…”
Wu: GPT-5 solves coding problems no other AI model can solve
“Especially for like coding use cases, especially at the, you know, at the, when it thinks for a while, it'll usually solve problems that no other models can solve.”
Wu: GPT-5 hallucinations dropped to near zero on certain benchmark evaluations
“I think there was an eval that showed that hallucinations basically went to zero for a lot of this.”
Agrawal: OpenAI's Latest Release Is a Unified System with Dynamic Routing
“It was the first time they released a system, not a single model. It has a model router. It makes, bakes in a lot of thinking of how much compute resources to allocate to a query coming in.”
Altman: GPT-5 will integrate all reasoning models into a single interface
“It's an integrated model, so you don't have to, like, pick in our model switcher and know if you should use GPT-IV or O-III or O-IV or any of the complicated things. It's just one thing that, that works”
Altman: GPT-5 is like having 24/7 PhD-level experts across all fields
“It is like having PhD-level experts in every field available to you, 24 seven, for whatever you need. Not only to ask anything, but also to do anything for you. So if you know, need a piece of software created, it can kind of do it from scratch all at once.”
Altman: GPT-5's robustness enables long, complex agentic workflows
“The sort of, the robustness and reliability has greatly increased and that, that's very helpful for agentic workflows. So I'm like very impressed by how long and complex of a task it can carry out.”
Altman: OpenAI added Indian user feedback into GPT-5 and ChatGPT
“We've taken a lot of feedback from users in India about what they'd like from us. Better support for languages, more affordable access much more, and we've been able to put that into this model and upgrades to ChatGPT.”
Altman: GPT-5 builds small software much better than any prior model
“GPT-V is quite good at helping to create small pieces of software very quickly. And much better than any model that I've used.”
Benchmark-topping AI models are not necessarily what consumers want for chat
“I don't necessarily think the like smartest model that scores the best on sort of all of these objective benchmarks of intelligence will be the model that people want to chat with.”
OpenAI's GPT-5 scored highest on the physician-trained HealthBench medical benchmark
“They talked about how GPT-V was kind of the highest scoring model on this thing called HealthBench, which is a benchmark they trained with like, 250 plus physicians. To measure how good an LLM is at answering medical questions.”
Sonwalkar: GPT-5 costs half as much as o3
“Also it's half the cost of O three. So it's much cheaper. So it helps you, helps your margins.”
Coogan: Model auto-routing gives OpenAI customer leverage and higher margins
“This change is more bullish for the business because it shows that, that OpenAI is a dominant consumer app and they have increasing leverage over the customer to route to cheaper models that will save money and be higher margin.”
Coogan: Power users criticizing GPT-5 will still use OpenAI within a month
“Let's check in with that person and see what app they have on their home row in a month. Almost certainly open AI. Almost certainly. I would be very shocked if they're like, I'm daily driving something else.”
Baker: GPT-5 is OpenAI's first frontier model that isn't decisively best
“This is the first time that, you know, open AI has released a new model that was not decisively the best.”
Baker: Grok 4 scored 44.4% on Humanity's Last Exam versus GPT-5's 42%
“GroK IV when it came out and it is improved. Since, since its release roughly a month ago, was at 44.4%, and GPT five was at 42.”
Turley: GPT-5 achieves state-of-the-art results on math and reasoning benchmarks
“One way to look at that is academic benchmarks on many of the standard ones whether or not it's math or reasoning or, you know, just raw intelligence, this model is state of the art.”
Turley: GPT-5 dynamically chooses when to reason without manual user prompting
“It thinks too, just like O three did, but you don't have to manually, you know, tell it to do that. It'll just dynamically decide to think when it needs to. And when it doesn't need to think, it just responds instantly.”
Turley: GPT-5 is amazing at building front-end applications
“GPT-Five is amazing at making great front end applications.”
Turley: GPT-5 will be free and default, with legacy models for Pro
“Just use the product. You don't even have to pay. Should be your default model starting tomorrow. And just use it and don't think about models anymore. Unless you want to and you're a pro user, in which case you get all the old models, so rest assured.”
Kim: GPT-5 internal testers felt insulted by instant answers to hard questions
“I think we hear this with GPT-Five internally when people are testing and they're like, oh, I thought I asked like a really hard question. I feel like a little bit insulted that I thought for like two seconds or like when it doesn't even want to think at all.”
Kim: GPT-5 is a step change for personal coding and writing
“I use it for coding and writing all the time, and it's just a huge stuff change.”
Kim: GPT-5 front-end coding is a massive leap over o3
“If you compare it to O three's front end coding capability, this is just totally next level.”
Kim: AI prompt-based app generation will spur surge in indie businesses
“I think we're just gonna have a lot more, I would expect, like, maybe a lot more, like, indie type of, like, Businesses built around this because of the fact that, like, you just need to have the idea, write a simple prompt, and then you get the full fledged a…”
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Kim: GPT-5's creative writing capability is tender and touching
“That's one of my favorite improvements in GBT five. The writing, I honestly find it's very tender and touching, especially for a lot of the creative writing that we want to do.”
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Lightcap: Base GPT-5 beats GPT-4o even without added reasoning time
“Even though if you don't allow any thinking time you still get a typically net better answer than you would for one of our non-thinking models like GPT-IV-I.”
Lightcap: GPT-5 beats previous models on SWE-bench and health benchmarks
“It scores better on things like Sweebench. It scores better on all the kind of academic evals that we put it through. This one in particular, we actually made a real emphasis to have it score better on certain health benchmarks. So It's better at medical reaso…”
Lightcap: Post-training and test-time compute act as force multipliers
“And that continues to hold true, but we now have this kind of other category of training, which is post-training and being able to use test time compute in more interesting ways than we used to as almost kind of a second stage of training. And so we think that…”
Lightcap: GPT-5 bakes in tool use and longer-horizon reasoning
“So using tools, for example, is something that really thinks really important for overall intelligence, GPT two and three couldn't really do that as well. GPT-IV could do it in a more nascent way. And now GPT-V, you get that baked in with the benefit of these …”
Lightcap: GPT-5 does not qualify as an AGI system
“And so I do, I think we're at a system that I would call AGI. No. But I think we see, we start to see the traces and the pieces of that overall system for generalized learning start to come together in models like GPT-V and I suspect suspect in its successors.”
Lightcap: GPT-5 capability overhang would fuel ten years of product building
“I think you could pause AI progress right here for 10 years, and you'd still have about a decade worth of new products to get built, of new ways that people figure out how to use the models even at a GPT-V level model in interesting products and interesting pr…”
Lightcap: OpenAI is bringing GPT-5 to ChatGPT's free tier
“And we're bringing GPT-V to our free tier”
Lightcap: GPT-5 will feel dramatically different to average users, not power users
“And so we expect that like for, yeah, for the average user, it will feel dramatically different. Maybe for the kind of upper echelon of power user, it may not feel as different.”
Lightcap: OpenAI prioritized healthcare applications during GPT-5 training
“We focused on health a lot with this release because that was one of the consistently common things that we heard from people as a starting point for how they've used powerful AI was in, when they're navigating a health journey. And so we really wanted to make…”
Lightcap: GPT-5 is four to five times more accurate than predecessor models
“GBD-V, I think depends on how you measure it, but it's, you know, four to five times more accurate than its predecessors.”
OpenAI Tested GPT-5 With Uber, Amgen, Cursor, and JetBrains Pre-Release
“We've worked with large enterprises and small startups and the entire spectrum in between on testing these models and GPT-V specifically before release. And we get a lot of feedback from companies like Uber and Amgen and Harvey and Cursor lovable you know JetB…”
Early GPT-5 testers report noticeable gains across coding, science, and writing
“And the story that we wrote, we kind of talked about how at least the people that we've talked to who tested it so far have been pretty impressed. They seem to think that it's been, you know, there's been improvements in a number of domains and both like scien…”
Morris: OpenAI will not declare GPT-5 as AGI
“I hate to make bold predictions, especially in tech, because you can often be spectacularly wrong, but I do not think they're going to say GPT-V is anything approaching AGI, however you choose to define it.”
Kantrowitz: There is a decent chance OpenAI calls GPT-5 AGI
“I think there is a decent chance. I'm not saying it's for sure going to happen. I think there's a decent chance that they are going to say it's AGI that GPT five is AGI and by they, I mean, open AI.”
Marcus: OpenAI's Project Orion failed and became GPT-4.5
“So OpenAI tried to build GPT-V and they had a thing called Project Orion and it actually failed. And eventually got released as GPT four and a half. So what they thought was going to be GPT five just didn't meet expectations.”
Patel: OpenAI's Orion training run failed to reach GPT-5 performance levels
“There were hopes that Orion could be used for GPT-V but its improvement was, like, not enough to be, like, really a GPT-V. Furthermore, it was trained on the classical method, which is, like which is a ton of pre-training, and then some reinforcement learning …”
Patel: GPT-5 will simultaneously scale pre-training and post-training reasoning
“And so now GPT-Five, as Sam calls it, is, is gonna be a model that has huge pre-training scale, right? Like GPT-Five, but also huge post-training scale, Like O-one and O-three and continuing to scale that up, right? This would be the first time we see a model …”