Nathan: Artifact quality improved dramatically over GPT-5.4 and GPT-5.5
“One of the big, like, pushes that we made for this launch was, like, artifacts, right? Like, both on the model side, like, I think if you compare this with 5.5 and 5.4 before that, you'll see that there's been pretty dramatic improvements in the quality of the…”
Claude Mythos and GPT-5.5 score 20% and 15% on ExploitGym benchmark
“Claude Mythos preview successfully exploited a 157 of the 898 instances, and OpenAI's GPT 5.5 exploited one 20 within, ah, 120 of the eight 98, so you have, like, roughly 20% performance for Mythos, and 5.5 got, like, 15% or something like that, but whenever y…”
Dan Shipper: GPT-5.6 Accurately Handles 90% of Routine Emails
“I think that five five was fine, but not quite good enough. Like, I would say generally the email responses here, so like this email draft, for example, is like, It's as if I wrote it, like it's, maybe there's something that I would slightly change, and I migh…”
Brown: GPT-5.5 is far more compute-efficient than GPT-5.4
“It turned out that 5.5 is just much more efficient with its thinking. If you run it at max settings, 5.4 is thinking for a lot longer. It takes longer to get back a response than 5.5. And once you control for the amount of thinking time, actually you can see t…”
Brown: Modern AI models can reason for weeks before plateauing
“What we're seeing today with the modern models is that 5.5 and other models can think for, if you scaffold them reasonably well, can think for weeks even before having performance plateau on some of these benchmarks.”
Awais: Devs bootstrap Taste files with frontier models, then execute cheaply
“A lot of people, what they're doing is they're building one project with a really high quality LLM, like Opus or GPT, 5.5, right? They're building a taste file, and then they're using, you know, super cheap models to continuously build on that more.”
Dubois: GPT-5.5 performs most tasks roughly two times faster
“Most of the tasks can be basically performed, I would say like two X faster now with this model.”
Dubois: GPT-5.5 succeeded by combining inference efficiency and latency optimizations
“And the final thing that people care about is latency on x-axis, performance on y-axis, and this is where everything comes together, and this is really what happened with 5.5.”
Dubois: Test-time compute scaling exhibits logarithmic, diminishing returns
“We, we've seen again and again, the longer the model think for the better answers we will get. The problem is that this, these curves that we're talking about are not, are definitely not linear, and like they, there's some plateauing effect, and they kind of l…”
GPT-5.5 uses Spud, OpenAI's first base model upgrade in a year
“GPT-Five. Five is based on a new base model called Spud, which is the first base model upgrade they've done in, I don't know, over a year, and having a new base model will pave the way for future improvements as well.”
Brown: GPT-5.5 API Costs Double That of GPT-5.4
“Yeah, if you use it via the API, it's twice as expensive as Ah, the recent model, which is, ah, GPT-Five-Point-Four. It's twice as expensive.”
Brown: GPT-5.5 API Costs 20% More Than Claude Opus 4.7
“And then it's also 20% more expensive than Opus 4.7.”
Kantrowitz: OpenAI gave all Nvidia employees access to GPT-5.5
“They also did something interesting, which is that they gave access to GPT, 5.5, to the entire company of NVIDIA.”
OpenAI's GPT-5.5 matches GPT-5.4 per-token latency in real-world serving
“GPT 5.5 delivers this step up in intelligence without compromising on speed. Matches GPT 5.4 per token latency in real world serving.”
Brockman: GPT-5.5 can now autonomously operate computers and web browsers
“The fact that it's now really crossed the threshold of usefulness for general kinds of applications, and so it's much better at creating slides, spreadsheets, much better at computer use, using your browser, being able to kind of click through applications tha…”