Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”
Patel: OpenAI Struggled on GPT-4.5 Cluster Scaling and Reinforcement Learning
“They actually tried that with 4.5. They screwed up some things cause it was really hard to get, you know, a 100,000 GPUs to work properly. There's challenges there. Also, they hadn't figured out the whole reinforcement learning paradigm at that time.”
Morcos: GPT-4.5 and Llama 4 show limits of naive mega-model scaling
“And I think that's what we've seen to some extent with the failure of the mega models, right? With 4.5 and Lama four and others. I think that there is a challenge of just continuing to do that naively and you have to figure out how to break it.”
OpenAI's GPT-4.5 Was Built as a Trillion-Parameter Dense Model
“GPT, 4.5 was an experimental model from OpenAI. It was the idea, let's train a trillion parameter dense model, meaning it is not sparse, meaning all the token, all the neurons are activated on every request. And it was so slow. It's really hard to run these th…”
Chollet: Scaling LLMs 50,000x Only Lifted ARC Accuracy to 10%
“And from at the time, back in 2019 to now, with a model like GPT 4.5, for instance, there's been a roughly 50,000 X scale up of basal alarms. And we went from zero percent accuracy on that benchmark to roughly 10%, which is not a lot.”
Noam Brown: GPT-4.5 makes Tic-Tac-Toe mistakes without System 2 reasoning
“With Tic-Tac-Toe, we see that, like, GPD-Four .5 falls over. You know, it plays decently well. I shouldn't say it falls over. It does reasonably well. You can draw the board. It can make legal moves, but it will make mistakes sometimes, and if you really need …”
Howard: OpenAI compute spending grows exponentially while model utility scales logarithmically
“They kept on kind of exponentially increasing the amount they were spending on their models, whilst the Return, you know, the kind of utility of those models was only increasing logarithmically, and you kind of very quickly hit this point where it's like, oh, …”
Howard: OpenAI is shutting down GPT-4.5
“I think they're shutting down that product or they've shut down that product, if I understand correctly.”
Marcus: OpenAI's Project Orion failed and became GPT-4.5
“So OpenAI tried to build GPT-V and they had a thing called Project Orion and it actually failed. And eventually got released as GPT four and a half. So what they thought was going to be GPT five just didn't meet expectations.”
Brown: GPT-4.5 Likely Has Trillions of Parameters Enabling Sparse Connections
“GPT, 4.5 is like ginormous model, trillions of parameters, most likely. And that like, there's more room in the model to have these like little sparse connections materialize as you go through layers of the transformer. And I just haven't seen anything like th…”
OpenAI is working to inject GPT-4.5's humor and nuance into future models
“We're working on incorporating kind of those improvements into the models more generally. People loved about 4.5 is like the humor, the green text, the nuance. So we've heard that feedback and I know, yeah, there's lots of folks working on that and trying to b…”
McLaughlin: GPT-4.5 was a really interesting moment for AI humor
“I will say you know, GPT 4.5 was like a really interesting moment for humor for me. Right. Where like, you know, I was testing this model internally a lot. And like one of the ways I like kind of realized like, wow, this actually is like a interesting step cha…”
Swyx: Frontier Models Exist Primarily to Distill Smaller, Usable Models
“Even GPT 4.5 is too expensive. Normally it's really gonna use it in, in any reasonable quantity. Like, you know, Claude 3.5 Opus, like if it does exist, still not like, you know, the thing that we actually use is Sonnet, right? So like, it's almost like a depl…”
Chen: GPT-4.5 performance jump matches leap from GPT-3.5 to GPT-4
“It signifies an order of magnitude improvement over the last models, kind of commensurate with the jump from 3.5 to four.”
Chen: Focus on reasoning models caused the longer gap before GPT-4.5
“Why there seems to be, you know, a little bit bigger of a gap in release time between four and 4.5, we've been really largely focused on developing the reasoning parallel paradigm as well.”
Chen: Users prefer GPT-4.5 over GPT-4o by 60% to 70% margins
“When we look at, kind of, comparisons against GPT-FORO you'll see that everyday use cases, people prefer, you know, by a margin of 60% for actually productivity and knowledge work against GPT-FORO, there's almost like a 70% preference rate.”
Chen: GPT-4.5 outshines reasoning models like o1 in creative writing
“And, you know, we find that in a lot of areas like creative writing, for instance
Again, this is stuff that we want to test over the next one or two months but we find that there are areas like creative writing where this model outshines reasoning models.”
Chen: GPT-4.5 scaling returns remain consistent with OpenAI's prior projections
“You know, we are seeing the same returns, and I do want to stress that GPT-D 4.5 is that next point on this unsupervised learning paradigm, and, you know, we're very rigorous about how we do this. We make projections based on all the models we've trained befor…”
Chen: Pausing and restarting training runs is standard across OpenAI models
“Actually, so I think it's interesting that this gets is a point that's attributed to this model because actually in, in, in developing all of our foundation models, right they're all experiments, right? I think you know, running all of the foundation models of…”
Chen: GPT-4.5 creates ASCII art almost flawlessly, unlike previous models
“If you ask any of the previous models to create ASCII art for you, right? Actually, they mostly just fall down. This one can do it Almost flawless.”
Chen: GPT-4.5 hits expected benchmark progression consistent with OpenAI's trajectory
“Well, I really don't think that the accurate characterization is that it doesn't hit the benchmarks that, that we expect it to. So when you look at kind of the development of three to 3.5 to four to 4.5 this does hit the benchmarks that we expect.”
Training next-generation frontier AI models will soon require gigawatts of power
“And then how much, basically, would it cost in terms of energy to train a GPT four level model, a 4.5 level model, five, whatever. And you get into the gigawatts pretty soon.”
Srinivas: OpenAI will stay ahead of competitors with GPT-4.5 or GPT-5
“I'm sure there's a GPT 4.5 or five that will stay ahead. So it really is going to be a cat and mouse game there where Anthropics playing catch up and OpenAI is ahead through multimodal capabilities, reasoning capabilities, and things, things like that.”
McGrew: GPT-5 will render current AI infrastructure solutions completely obsolete
“I'm always a little worried about the infrastructure work, because I think, you know, you're solving the problems as they exist today, and when GPD 4.5 and GPD five come out, you know, they're gonna have fundamentally different use cases and fundamentally diff…”