Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 1/5

Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%

Nathan Labenz · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · Oct 14, 2025 · at 8:58

Nathan Labenz, host of The Cognitive Revolution, compares OpenAI's GPT-4.5 model performance against the o3 reasoning class on factual trivia benchmarks.

0:00 / 0:19exact quote · 19.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Nathan Labenz

Assertion Supported
Labenz: Claude 4 system card documented AI blackmailing a human engineer
“In the cloud four system card, they reported blackmailing of the human. The setup was that the AI had access to the engineer's email and They told the AI that it was going to be like replaced with a, you know, a less ethical version or something like that. It …”
Nathan Labenz Oct 14, 2025 ▶ 1:04:23 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Prediction Not checkable as stated
Labenz: Tool-Equipped Next-Gen AI Models Will Resemble Superintelligence
“When we start to give the next generation of the model these power tools, and they start to solve previously unsolved engineering problems, I think you start to have something that looks kind of like super intelligence.”
Nathan Labenz Oct 14, 2025 ▶ 0:14 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Opinion
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Nathan Labenz Oct 14, 2025 ▶ 5:03 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Prediction Not checkable as stated
Labenz: AI will outperform average developers on standard apps within five years
“But I would be very surprised if you can't get your nuts and bolts Web app, mobile app type things spit out for you for far less and far faster than, and probably honestly with significantly higher quality and less back and forth with an AI system than, you kn…”
Nathan Labenz Oct 14, 2025 ▶ 44:58 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Assertion Not checkable as stated
Labenz: OpenAI's router failure caused bad initial GPT-5 outputs
“The problem at launch was that that router was broken. So all of the queries were going to the dumb model, and so a lot of people literally just got Bad outputs, which were worse than oh three because they were getting non thinking responses.”
Nathan Labenz Oct 14, 2025 ▶ 21:35 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Assertion Not checkable as stated
Labenz: Plugged-in AI experts are not pushing timelines past 2030
“I don't think too many people, at least that I, you know, think are really plugged in on this, are pushing out too much past 20 30 at all.”
Nathan Labenz Oct 14, 2025 ▶ 23:57 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.