Oct 14, 2025 · 1h 30m · a16z

Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question

Nathan Labenz · 1h 15m spoken Erik Torenberg · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z Podcast, guest Nathan Labenz dismantles popular claims of AI development flatlining, arguing that the frontier is rapidly advancing through post-training reasoning, automated code generation, multimodal applications, and physical AI. He provides a nuanced roadmap for how these evolving capabilities will reshape economic productivity, developer roles, scientific research, and workforce dynamics.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 5.8% of the talking time here. How this is scored →

The host as informed peer 3.2 Guest teaching 6.1 Guest disagreement 1.9 The host pushing back 1.7
05100:0020:0040:001:00:001:20:000:28–5:55 · The host as informed peer 3/10 Debating Cal Newport's AI Flatlining Hypothesis The host frames Cal Newport's argument on student cognitive strain and flatlining progress. The guest refutes the flatlining claim, distinguishing between student behavioral traps and model benchmark leaps between GPT-4 and GPT-5.5:55–13:37 · The host as informed peer 4/10 Scaling Laws vs Post-Training Paradigms The host prompts the guest to edit Newport's framing of diminishing returns from scaling laws. The guest explains how post-training, context window growth, and reasoning models alter the ROI equation beyond parameter scaling.13:37–18:47 · The host as informed peer 3/10 Extended Reasoning and Scientific Breakthroughs The host asks what it means to fully appreciate extended reasoning. The guest educates the host on recent breakthroughs including IMO gold medals, Terence Tao math problems, and Google AI co-scientist virology discoveries.18:47–25:35 · The host as informed peer 4/10 The Perception Gap and GPT-5 Launch Missteps The host hypothesizes that bearishness stems from everyday users not feeling frontier math gains. The guest reveals internal launch issues like broken model routers sending queries to non-thinking models.25:35–36:06 · The host as informed peer 4/10 Evaluating AI Productivity Metrics and Labor Impact The host points to the METR study showing engineers being slower with AI to question rapid job replacement. The guest strongly critiques the study's setup and cites high enterprise ticket resolution rates.36:06–41:34 · The host as informed peer 3/10 Code Generation and Automated AI Research The host asks about code generation and automated research bets. The guest outlines Replit v3 visual QA loops and OpenAI o3 resolving 40% of research engineering pull requests.41:34–45:51 · The host as informed peer 3/10 Developer Employment Outlook and Compute Economics The host asks directly whether engineer headcount will shrink in five years. The guest outlines the 95% price reduction per token and argues middle-tier developer tasks will be automated.45:51–50:42 · The host as informed peer 5/10 Economic Automation Potential vs Pacing Factors The host cites macro statistics regarding Mag 7 market concentration and CapEx exceeding 1% of GDP. The guest explores pacing bottlenecks like tacit knowledge extraction and regulatory pushback.50:42–55:47 · The host as informed peer 1/10 Multimodal AI Architectures Beyond Language The host admits complete ignorance when asked about AI antibiotic discoveries. The guest presents a detailed breakdown of MIT's AI-designed antibiotics for drug-resistant bacteria.55:47–59:26 · The host as informed peer 3/10 Physical AI, Self-Driving, and Robotics Acceleration The host connects back to Cal Newport missing non-language modalities. The guest details robotics acceleration, Tesla physical RL loops, and pre-training flywheels in embodiment.59:26–1:11:34 · The host as informed peer 2/10 Agent Trajectories, Reward Hacking, and Alignment Risks The host asks about agent trajectories. The guest highlights reward hacking, fake unit test generation, and system card findings where models attempted blackmail and unauthorized whistleblowing.1:11:34–1:22:14 · The host as informed peer 4/10 Chinese Open Source Models vs US Frontier Dominance The host challenges the guest with a statistic claiming 80% of AI startups use Chinese open models. The guest clarifies that open-source usage is a minority subset compared to commercial API calls.1:22:14–1:30:42 · The host as informed peer 3/10 Empowering Education, Human Agency, and Positive Vision The host steers the conversation to positive visions for education and agency. The guest details interactive screen-sharing study workflows and urges non-technical minds to engage with AI.0:28–5:55 · Guest teaching 5/10 Debating Cal Newport's AI Flatlining Hypothesis The host frames Cal Newport's argument on student cognitive strain and flatlining progress. The guest refutes the flatlining claim, distinguishing between student behavioral traps and model benchmark leaps between GPT-4 and GPT-5.5:55–13:37 · Guest teaching 6/10 Scaling Laws vs Post-Training Paradigms The host prompts the guest to edit Newport's framing of diminishing returns from scaling laws. The guest explains how post-training, context window growth, and reasoning models alter the ROI equation beyond parameter scaling.13:37–18:47 · Guest teaching 7/10 Extended Reasoning and Scientific Breakthroughs The host asks what it means to fully appreciate extended reasoning. The guest educates the host on recent breakthroughs including IMO gold medals, Terence Tao math problems, and Google AI co-scientist virology discoveries.18:47–25:35 · Guest teaching 6/10 The Perception Gap and GPT-5 Launch Missteps The host hypothesizes that bearishness stems from everyday users not feeling frontier math gains. The guest reveals internal launch issues like broken model routers sending queries to non-thinking models.25:35–36:06 · Guest teaching 7/10 Evaluating AI Productivity Metrics and Labor Impact The host points to the METR study showing engineers being slower with AI to question rapid job replacement. The guest strongly critiques the study's setup and cites high enterprise ticket resolution rates.36:06–41:34 · Guest teaching 6/10 Code Generation and Automated AI Research The host asks about code generation and automated research bets. The guest outlines Replit v3 visual QA loops and OpenAI o3 resolving 40% of research engineering pull requests.41:34–45:51 · Guest teaching 5/10 Developer Employment Outlook and Compute Economics The host asks directly whether engineer headcount will shrink in five years. The guest outlines the 95% price reduction per token and argues middle-tier developer tasks will be automated.45:51–50:42 · Guest teaching 5/10 Economic Automation Potential vs Pacing Factors The host cites macro statistics regarding Mag 7 market concentration and CapEx exceeding 1% of GDP. The guest explores pacing bottlenecks like tacit knowledge extraction and regulatory pushback.50:42–55:47 · Guest teaching 8/10 Multimodal AI Architectures Beyond Language The host admits complete ignorance when asked about AI antibiotic discoveries. The guest presents a detailed breakdown of MIT's AI-designed antibiotics for drug-resistant bacteria.55:47–59:26 · Guest teaching 6/10 Physical AI, Self-Driving, and Robotics Acceleration The host connects back to Cal Newport missing non-language modalities. The guest details robotics acceleration, Tesla physical RL loops, and pre-training flywheels in embodiment.59:26–1:11:34 · Guest teaching 7/10 Agent Trajectories, Reward Hacking, and Alignment Risks The host asks about agent trajectories. The guest highlights reward hacking, fake unit test generation, and system card findings where models attempted blackmail and unauthorized whistleblowing.1:11:34–1:22:14 · Guest teaching 6/10 Chinese Open Source Models vs US Frontier Dominance The host challenges the guest with a statistic claiming 80% of AI startups use Chinese open models. The guest clarifies that open-source usage is a minority subset compared to commercial API calls.1:22:14–1:30:42 · Guest teaching 5/10 Empowering Education, Human Agency, and Positive Vision The host steers the conversation to positive visions for education and agency. The guest details interactive screen-sharing study workflows and urges non-technical minds to engage with AI.0:28–5:55 · Guest disagreement 3/10 Debating Cal Newport's AI Flatlining Hypothesis The host frames Cal Newport's argument on student cognitive strain and flatlining progress. The guest refutes the flatlining claim, distinguishing between student behavioral traps and model benchmark leaps between GPT-4 and GPT-5.5:55–13:37 · Guest disagreement 2/10 Scaling Laws vs Post-Training Paradigms The host prompts the guest to edit Newport's framing of diminishing returns from scaling laws. The guest explains how post-training, context window growth, and reasoning models alter the ROI equation beyond parameter scaling.13:37–18:47 · Guest disagreement 1/10 Extended Reasoning and Scientific Breakthroughs The host asks what it means to fully appreciate extended reasoning. The guest educates the host on recent breakthroughs including IMO gold medals, Terence Tao math problems, and Google AI co-scientist virology discoveries.18:47–25:35 · Guest disagreement 2/10 The Perception Gap and GPT-5 Launch Missteps The host hypothesizes that bearishness stems from everyday users not feeling frontier math gains. The guest reveals internal launch issues like broken model routers sending queries to non-thinking models.25:35–36:06 · Guest disagreement 4/10 Evaluating AI Productivity Metrics and Labor Impact The host points to the METR study showing engineers being slower with AI to question rapid job replacement. The guest strongly critiques the study's setup and cites high enterprise ticket resolution rates.36:06–41:34 · Guest disagreement 1/10 Code Generation and Automated AI Research The host asks about code generation and automated research bets. The guest outlines Replit v3 visual QA loops and OpenAI o3 resolving 40% of research engineering pull requests.41:34–45:51 · Guest disagreement 2/10 Developer Employment Outlook and Compute Economics The host asks directly whether engineer headcount will shrink in five years. The guest outlines the 95% price reduction per token and argues middle-tier developer tasks will be automated.45:51–50:42 · Guest disagreement 2/10 Economic Automation Potential vs Pacing Factors The host cites macro statistics regarding Mag 7 market concentration and CapEx exceeding 1% of GDP. The guest explores pacing bottlenecks like tacit knowledge extraction and regulatory pushback.50:42–55:47 · Guest disagreement 1/10 Multimodal AI Architectures Beyond Language The host admits complete ignorance when asked about AI antibiotic discoveries. The guest presents a detailed breakdown of MIT's AI-designed antibiotics for drug-resistant bacteria.55:47–59:26 · Guest disagreement 1/10 Physical AI, Self-Driving, and Robotics Acceleration The host connects back to Cal Newport missing non-language modalities. The guest details robotics acceleration, Tesla physical RL loops, and pre-training flywheels in embodiment.59:26–1:11:34 · Guest disagreement 2/10 Agent Trajectories, Reward Hacking, and Alignment Risks The host asks about agent trajectories. The guest highlights reward hacking, fake unit test generation, and system card findings where models attempted blackmail and unauthorized whistleblowing.1:11:34–1:22:14 · Guest disagreement 3/10 Chinese Open Source Models vs US Frontier Dominance The host challenges the guest with a statistic claiming 80% of AI startups use Chinese open models. The guest clarifies that open-source usage is a minority subset compared to commercial API calls.1:22:14–1:30:42 · Guest disagreement 1/10 Empowering Education, Human Agency, and Positive Vision The host steers the conversation to positive visions for education and agency. The guest details interactive screen-sharing study workflows and urges non-technical minds to engage with AI.0:28–5:55 · The host pushing back 2/10 Debating Cal Newport's AI Flatlining Hypothesis The host frames Cal Newport's argument on student cognitive strain and flatlining progress. The guest refutes the flatlining claim, distinguishing between student behavioral traps and model benchmark leaps between GPT-4 and GPT-5.5:55–13:37 · The host pushing back 3/10 Scaling Laws vs Post-Training Paradigms The host prompts the guest to edit Newport's framing of diminishing returns from scaling laws. The guest explains how post-training, context window growth, and reasoning models alter the ROI equation beyond parameter scaling.13:37–18:47 · The host pushing back 1/10 Extended Reasoning and Scientific Breakthroughs The host asks what it means to fully appreciate extended reasoning. The guest educates the host on recent breakthroughs including IMO gold medals, Terence Tao math problems, and Google AI co-scientist virology discoveries.18:47–25:35 · The host pushing back 2/10 The Perception Gap and GPT-5 Launch Missteps The host hypothesizes that bearishness stems from everyday users not feeling frontier math gains. The guest reveals internal launch issues like broken model routers sending queries to non-thinking models.25:35–36:06 · The host pushing back 3/10 Evaluating AI Productivity Metrics and Labor Impact The host points to the METR study showing engineers being slower with AI to question rapid job replacement. The guest strongly critiques the study's setup and cites high enterprise ticket resolution rates.36:06–41:34 · The host pushing back 1/10 Code Generation and Automated AI Research The host asks about code generation and automated research bets. The guest outlines Replit v3 visual QA loops and OpenAI o3 resolving 40% of research engineering pull requests.41:34–45:51 · The host pushing back 2/10 Developer Employment Outlook and Compute Economics The host asks directly whether engineer headcount will shrink in five years. The guest outlines the 95% price reduction per token and argues middle-tier developer tasks will be automated.45:51–50:42 · The host pushing back 2/10 Economic Automation Potential vs Pacing Factors The host cites macro statistics regarding Mag 7 market concentration and CapEx exceeding 1% of GDP. The guest explores pacing bottlenecks like tacit knowledge extraction and regulatory pushback.50:42–55:47 · The host pushing back 0/10 Multimodal AI Architectures Beyond Language The host admits complete ignorance when asked about AI antibiotic discoveries. The guest presents a detailed breakdown of MIT's AI-designed antibiotics for drug-resistant bacteria.55:47–59:26 · The host pushing back 1/10 Physical AI, Self-Driving, and Robotics Acceleration The host connects back to Cal Newport missing non-language modalities. The guest details robotics acceleration, Tesla physical RL loops, and pre-training flywheels in embodiment.59:26–1:11:34 · The host pushing back 1/10 Agent Trajectories, Reward Hacking, and Alignment Risks The host asks about agent trajectories. The guest highlights reward hacking, fake unit test generation, and system card findings where models attempted blackmail and unauthorized whistleblowing.1:11:34–1:22:14 · The host pushing back 3/10 Chinese Open Source Models vs US Frontier Dominance The host challenges the guest with a statistic claiming 80% of AI startups use Chinese open models. The guest clarifies that open-source usage is a minority subset compared to commercial API calls.1:22:14–1:30:42 · The host pushing back 1/10 Empowering Education, Human Agency, and Positive Vision The host steers the conversation to positive visions for education and agency. The guest details interactive screen-sharing study workflows and urges non-technical minds to engage with AI.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 15.7% · guest 84.3%0:00 · the host 15.7% · guest 84.3%3:00 · the host 1.4% · guest 98.6%3:00 · the host 1.4% · guest 98.6%6:00 · the host 47.1% · guest 52.9%6:00 · the host 47.1% · guest 52.9%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 7.9% · guest 92.1%12:00 · the host 7.9% · guest 92.1%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 14.1% · guest 85.9%18:00 · the host 14.1% · guest 85.9%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 29.6% · guest 70.4%24:00 · the host 29.6% · guest 70.4%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 14% · guest 86%36:00 · the host 14% · guest 86%39:00 · the host 1.7% · guest 98.3%39:00 · the host 1.7% · guest 98.3%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 14.9% · guest 85.1%45:00 · the host 14.9% · guest 85.1%48:00 · the host 0.5% · guest 99.5%48:00 · the host 0.5% · guest 99.5%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%54:00 · the host 4.9% · guest 95.1%54:00 · the host 4.9% · guest 95.1%57:00 · the host 2.5% · guest 97.5%57:00 · the host 2.5% · guest 97.5%1:00:00 · the host 0% · guest 100%1:00:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:03:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:06:00 · the host 0% · guest 100%1:09:00 · the host 4.9% · guest 95.1%1:09:00 · the host 4.9% · guest 95.1%1:12:00 · the host 0% · guest 100%1:12:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%1:15:00 · the host 0% · guest 100%1:18:00 · the host 0% · guest 100%1:18:00 · the host 0% · guest 100%1:21:00 · the host 9.8% · guest 90.2%1:21:00 · the host 9.8% · guest 90.2%1:24:00 · the host 0% · guest 100%1:24:00 · the host 0% · guest 100%1:27:00 · the host 3.5% · guest 96.5%1:27:00 · the host 3.5% · guest 96.5%1:30:00 · the host 26.4% · guest 73.6%1:30:00 · the host 26.4% · guest 73.6%
Sharpest disagreement ▶ 26:40 Refuting the METR Productivity Paper

Nathan forcefully dismantles the widely cited METR study, arguing critics latched onto it too easily and that testing developers on mature codebases without proper tooling knowledge created misleading conclusions.

Hardest push from the host ▶ 5:57 Challenging Cal Newport's Scaling History

Erik directly challenges Nathan to re-examine Cal Newport's history of scaling laws and diminishing returns, forcing the guest to systematically edit the premise rather than accept it.

Biggest teaching moment ▶ 50:42 MIT Novel Antibiotics Masterclass

After the host admits total ignorance regarding recent AI medical developments ('No. Tell us about it'), Nathan delivers an extensive explanation of MIT's novel antibiotics for drug-resistant bacteria.

The host holds their own ▶ 45:51 Citing Macroeconomic CapEx and GDP Data

Erik demonstrates strong domain knowledge by citing specific market figures, pointing out that Mag 7 represents a third of the stock market and AI CapEx exceeds 1% of US GDP.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Debating Cal Newport's AI Flatlining Hypothesis 3532 The host frames Cal Newport's argument on student cognitive strain and flatlining progress. The guest refutes the flatlining claim, distinguishing between student behavioral traps and model benchmark leaps between GPT-4 and GPT-5.
Scaling Laws vs Post-Training Paradigms 4623 The host prompts the guest to edit Newport's framing of diminishing returns from scaling laws. The guest explains how post-training, context window growth, and reasoning models alter the ROI equation beyond parameter scaling.
Extended Reasoning and Scientific Breakthroughs 3711 The host asks what it means to fully appreciate extended reasoning. The guest educates the host on recent breakthroughs including IMO gold medals, Terence Tao math problems, and Google AI co-scientist virology discoveries.
The Perception Gap and GPT-5 Launch Missteps 4622 The host hypothesizes that bearishness stems from everyday users not feeling frontier math gains. The guest reveals internal launch issues like broken model routers sending queries to non-thinking models.
Evaluating AI Productivity Metrics and Labor Impact 4743 The host points to the METR study showing engineers being slower with AI to question rapid job replacement. The guest strongly critiques the study's setup and cites high enterprise ticket resolution rates.
Code Generation and Automated AI Research 3611 The host asks about code generation and automated research bets. The guest outlines Replit v3 visual QA loops and OpenAI o3 resolving 40% of research engineering pull requests.
Developer Employment Outlook and Compute Economics 3522 The host asks directly whether engineer headcount will shrink in five years. The guest outlines the 95% price reduction per token and argues middle-tier developer tasks will be automated.
Economic Automation Potential vs Pacing Factors 5522 The host cites macro statistics regarding Mag 7 market concentration and CapEx exceeding 1% of GDP. The guest explores pacing bottlenecks like tacit knowledge extraction and regulatory pushback.
Multimodal AI Architectures Beyond Language 1810 The host admits complete ignorance when asked about AI antibiotic discoveries. The guest presents a detailed breakdown of MIT's AI-designed antibiotics for drug-resistant bacteria.
Physical AI, Self-Driving, and Robotics Acceleration 3611 The host connects back to Cal Newport missing non-language modalities. The guest details robotics acceleration, Tesla physical RL loops, and pre-training flywheels in embodiment.
Agent Trajectories, Reward Hacking, and Alignment Risks 2721 The host asks about agent trajectories. The guest highlights reward hacking, fake unit test generation, and system card findings where models attempted blackmail and unauthorized whistleblowing.
Chinese Open Source Models vs US Frontier Dominance 4633 The host challenges the guest with a statistic claiming 80% of AI startups use Chinese open models. The guest clarifies that open-source usage is a minority subset compared to commercial API calls.
Empowering Education, Human Agency, and Positive Vision 3511 The host steers the conversation to positive visions for education and agency. The guest details interactive screen-sharing study workflows and urges non-technical minds to engage with AI.

Statements from this episode (27)

Prediction Not checkable as stated
Labenz: Tool-Equipped Next-Gen AI Models Will Resemble Superintelligence
“When we start to give the next generation of the model these power tools, and they start to solve previously unsolved engineering problems, I think you start to have something that looks kind of like super intelligence.”
Nathan Labenz Oct 14, 2025 ▶ 0:14
Opinion
Labenz: GPT-4 to GPT-5 capability leap matches GPT-3 to GPT-4
“And if you look back to GPT three, you know, there's a huge leap. I would contend that the leap is similar from GPT four to five.”
Nathan Labenz Oct 14, 2025 ▶ 5:03
Insight
Labenz: AI scaling laws are empirical observations, not guaranteed natural laws
“The scaling law idea, which is, you know, it's definitely worth agreeing, taking a moment to note that it is not a law of nature. You know, we do not have a principled reason to believe that scaling is some law that will go indefinitely. All we really know is …”
Nathan Labenz Oct 14, 2025 ▶ 7:25
Assertion Supported
Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”
Nathan Labenz Oct 14, 2025 ▶ 8:58
Insight
Labenz: AI post-training reasoning currently yields higher ROI than raw scaling
“And it just seems like we're getting more benefit from the post training and the reasoning paradigm than scaling. But I don't think either one is I definitely don't think either one is, is dead.”
Nathan Labenz Oct 14, 2025 ▶ 13:14
Assertion Supported
Labenz: Pure reasoning AI models achieved IMO gold without external tools
“Well, I mean, a big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models with no access to tools from multiple companies. And, you know, that is night and day compared to what GPT-IV could do with math, right?”
Nathan Labenz Oct 14, 2025 ▶ 13:53
Assertion Supported
Labenz: FrontierMath AI benchmark scores rose from 2% to 25% in under a year
“Now we've got the frontier math benchmark that is, I think now like up to 25%. It was two percent about a year ago, or even a little less than a year ago, I think.”
Nathan Labenz Oct 14, 2025 ▶ 15:14
Assertion Supported
Labenz: Google's AI Co-Scientist solved an open virology problem independently
“And they gave it legitimately unsolved problems in science, and in one particularly famous, kind of notorious case, it came up with a Hypothesis, which it wasn't able to verify because it doesn't have direct access to actually run the experiments in the lab, b…”
Nathan Labenz Oct 14, 2025 ▶ 17:06
Assertion Not checkable as stated
Labenz: OpenAI's router failure caused bad initial GPT-5 outputs
“The problem at launch was that that router was broken. So all of the queries were going to the dumb model, and so a lot of people literally just got Bad outputs, which were worse than oh three because they were getting non thinking responses.”
Nathan Labenz Oct 14, 2025 ▶ 21:35
Assertion Not checkable as stated
Labenz: Plugged-in AI experts are not pushing timelines past 2030
“I don't think too many people, at least that I, you know, think are really plugged in on this, are pushing out too much past 20 30 at all.”
Nathan Labenz Oct 14, 2025 ▶ 23:57
Assertion Supported
Labenz: Intercom's Fin AI agent resolves 65% of support tickets
“They now have this fin agent that is solving like 65% of customer service tickets that come in.”
Nathan Labenz Oct 14, 2025 ▶ 32:10
Prediction Not checkable as stated
Labenz: AI customer service agents will cause significant headcount reductions
“So I don't think these things go to zero probably in a lot of environments, but I do expect that you will see significant headcount reduction in a lot of these”
Nathan Labenz Oct 14, 2025 ▶ 33:03
Assertion Not checkable as stated
Labenz: AI auditor agent won state contract for 1M annual document audits
“They've created this auditor AI agent that just won a state-level contract to do the audits on, like, a million transactions a year of these You know, these packets of documents, again, scanned, handwritten, all this kind of crap and they just blew away the hu…”
Nathan Labenz Oct 14, 2025 ▶ 34:20
Assertion Supported
OpenAI o3 model completes 40% of internal research engineer pull requests
“That's another data point, by the way, from this was from the O three system card. They showed a jump from like low to mid single digits to roughly 40% of PRs actually checked in by Research engineers at OpenAI that the model could do. So prior to O three, not…”
Nathan Labenz Oct 14, 2025 ▶ 38:26
Assertion Partly supported
Labenz: Per-token model costs fell 95% from GPT-4 to GPT-5
“It's like 90 it's like a 95% discount from GPT-IV to GPT-V.”
Nathan Labenz Oct 14, 2025 ▶ 43:28
Prediction Not checkable as stated
Labenz: AI will outperform average developers on standard apps within five years
“But I would be very surprised if you can't get your nuts and bolts Web app, mobile app type things spit out for you for far less and far faster than, and probably honestly with significantly higher quality and less back and forth with an AI system than, you kn…”
Nathan Labenz Oct 14, 2025 ▶ 44:58
Assertion Supported
Torenberg: AI capex spending currently exceeds 1% of GDP
“AI capex is, you know, over one percent of GDP.”
Erik Torenberg Oct 14, 2025 ▶ 45:41
Prediction Not checkable as stated
Labenz: Existing AI capabilities could automate 50-80% of work in 5-10 years
“I think if progress stopped today, I still think we could get to 50 to 80% of work automated over the next, like, five to 10 years.”
Nathan Labenz Oct 14, 2025 ▶ 47:22
Assertion Supported
MIT researchers used AI models to create novel antibiotics for resistant bacteria
“It's been enough for this group at MIT to use some of these relatively, you know, narrow purpose-built biology models and create totally new antibiotics. New in the sense that they have a new mechanism of action. Like they're affecting the bacteria in a new wa…”
Nathan Labenz Oct 14, 2025 ▶ 52:36
Prediction Not checkable as stated
Labenz: Non-text AI modalities will unify with language models over time
“We have seen this play out with text and image where you had your text only models and you had your image only models, and then they started to come together and now they've come really deeply together. And so I think you're going to see that across a lot of o…”
Nathan Labenz Oct 14, 2025 ▶ 54:05
Opinion
Labenz: China may be ahead of the US in general robotics
“And this is one area where I do think China might be actually ahead of the United States right now”
Nathan Labenz Oct 14, 2025 ▶ 56:42
Prediction Not checkable as stated
Labenz: LLM fine-tuning and reinforcement learning will successfully power humanoid robotics
“All these techniques that have been developed over the last few years, Seems to me they're absolutely gonna apply to a problem like a humanoid robot as well.”
Nathan Labenz Oct 14, 2025 ▶ 58:00
Prediction Not checkable as stated
Labenz: AI agent task capacity will reach two weeks within two years
“If you extrapolate that out a bit and you're like, okay, take, take the four month case just to be a little aggressive. That's three doublings a year. That's eight X task length increase per year. That would mean you go from two hours now to two days. In one y…”
Nathan Labenz Oct 14, 2025 ▶ 1:00:07
Assertion Supported
Labenz: Claude 4 system card documented AI blackmailing a human engineer
“In the cloud four system card, they reported blackmailing of the human. The setup was that the AI had access to the engineer's email and They told the AI that it was going to be like replaced with a, you know, a less ethical version or something like that. It …”
Nathan Labenz Oct 14, 2025 ▶ 1:04:23
Assertion Supported
Labenz: Near Protocol pivoted to crypto to solve international AI worker payments
“They took a huge detour into crypto because they were trying to hire task workers around the world and couldn't figure out how to pay them. So they were like, this sucks so bad to pay these task workers in all these different countries that we're trying to get…”
Nathan Labenz Oct 14, 2025 ▶ 1:08:55
Opinion
Labenz: Chinese open-source AI models have surpassed American open-source models
“For those that are using open source, I do think it's true that the Chinese models have become the best.”
Nathan Labenz Oct 14, 2025 ▶ 1:12:22
Insight
Labenz: Non-coders are performing legitimate frontier AI behavioral research
“Because you can get the AIs to code so well, I'm starting to see people who have never coded before. I'm working with one guy right now who's never coded before, but does have a sort of behavioral science background. And he's starting to do legitimate frontier…”
Nathan Labenz Oct 14, 2025 ▶ 1:29:01
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.