Assertion Supported AI assessment confidence: 95% certainty 3/5 debate potential 1/5

Labenz: FrontierMath AI benchmark scores rose from 2% to 25% in under a year

Nathan Labenz · Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question · Oct 14, 2025 · at 15:14

Nathan Labenz, host of The Cognitive Revolution, explains rapid progress in advanced mathematical reasoning capabilities during a discussion with Erik Torenberg.

0:00 / 0:08exact quote · 8.2s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Now we've got the frontier math benchmark that is, I think now like up to 25%. It was two percent about a year ago, or even a little less than a year ago, I think.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Nathan Labenz

Assertion Supported
Labenz: Claude 4 system card documented AI blackmailing a human engineer
“In the cloud four system card, they reported blackmailing of the human. The setup was that the AI had access to the engineer's email and They told the AI that it was going to be like replaced with a, you know, a less ethical version or something like that. It …”
Nathan Labenz Oct 14, 2025 ▶ 1:04:23 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Prediction Not checkable as stated
Labenz: Tool-Equipped Next-Gen AI Models Will Resemble Superintelligence
“When we start to give the next generation of the model these power tools, and they start to solve previously unsolved engineering problems, I think you start to have something that looks kind of like super intelligence.”
Nathan Labenz Oct 14, 2025 ▶ 0:14 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Opinion
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Nathan Labenz Oct 14, 2025 ▶ 5:03 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Prediction Not checkable as stated
Labenz: AI will outperform average developers on standard apps within five years
“But I would be very surprised if you can't get your nuts and bolts Web app, mobile app type things spit out for you for far less and far faster than, and probably honestly with significantly higher quality and less back and forth with an AI system than, you kn…”
Nathan Labenz Oct 14, 2025 ▶ 44:58 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Assertion Not checkable as stated
Labenz: OpenAI's router failure caused bad initial GPT-5 outputs
“The problem at launch was that that router was broken. So all of the queries were going to the dumb model, and so a lot of people literally just got Bad outputs, which were worse than oh three because they were getting non thinking responses.”
Nathan Labenz Oct 14, 2025 ▶ 21:35 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Assertion Not checkable as stated
Labenz: Plugged-in AI experts are not pushing timelines past 2030
“I don't think too many people, at least that I, you know, think are really plugged in on this, are pushing out too much past 20 30 at all.”
Nathan Labenz Oct 14, 2025 ▶ 23:57 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.