Oct 14, 2025 · 1h 30m · a16z
Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the a16z Podcast, guest Nathan Labenz dismantles popular claims of AI development flatlining, arguing that the frontier is rapidly advancing through post-training reasoning, automated code generation, multimodal applications, and physical AI. He provides a nuanced roadmap for how these evolving capabilities will reshape economic productivity, developer roles, scientific research, and workforce dynamics.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 5.8% of the talking time here. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Nathan forcefully dismantles the widely cited METR study, arguing critics latched onto it too easily and that testing developers on mature codebases without proper tooling knowledge created misleading conclusions.
Hardest push from the host ▶ 5:57 Challenging Cal Newport's Scaling HistoryErik directly challenges Nathan to re-examine Cal Newport's history of scaling laws and diminishing returns, forcing the guest to systematically edit the premise rather than accept it.
Biggest teaching moment ▶ 50:42 MIT Novel Antibiotics MasterclassAfter the host admits total ignorance regarding recent AI medical developments ('No. Tell us about it'), Nathan delivers an extensive explanation of MIT's novel antibiotics for drug-resistant bacteria.
The host holds their own ▶ 45:51 Citing Macroeconomic CapEx and GDP DataErik demonstrates strong domain knowledge by citing specific market figures, pointing out that Mag 7 represents a third of the stock market and AI CapEx exceeds 1% of US GDP.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Debating Cal Newport's AI Flatlining Hypothesis | 3 | 5 | 3 | 2 | The host frames Cal Newport's argument on student cognitive strain and flatlining progress. The guest refutes the flatlining claim, distinguishing between student behavioral traps and model benchmark leaps between GPT-4 and GPT-5. | |
| Scaling Laws vs Post-Training Paradigms | 4 | 6 | 2 | 3 | The host prompts the guest to edit Newport's framing of diminishing returns from scaling laws. The guest explains how post-training, context window growth, and reasoning models alter the ROI equation beyond parameter scaling. | |
| Extended Reasoning and Scientific Breakthroughs | 3 | 7 | 1 | 1 | The host asks what it means to fully appreciate extended reasoning. The guest educates the host on recent breakthroughs including IMO gold medals, Terence Tao math problems, and Google AI co-scientist virology discoveries. | |
| The Perception Gap and GPT-5 Launch Missteps | 4 | 6 | 2 | 2 | The host hypothesizes that bearishness stems from everyday users not feeling frontier math gains. The guest reveals internal launch issues like broken model routers sending queries to non-thinking models. | |
| Evaluating AI Productivity Metrics and Labor Impact | 4 | 7 | 4 | 3 | The host points to the METR study showing engineers being slower with AI to question rapid job replacement. The guest strongly critiques the study's setup and cites high enterprise ticket resolution rates. | |
| Code Generation and Automated AI Research | 3 | 6 | 1 | 1 | The host asks about code generation and automated research bets. The guest outlines Replit v3 visual QA loops and OpenAI o3 resolving 40% of research engineering pull requests. | |
| Developer Employment Outlook and Compute Economics | 3 | 5 | 2 | 2 | The host asks directly whether engineer headcount will shrink in five years. The guest outlines the 95% price reduction per token and argues middle-tier developer tasks will be automated. | |
| Economic Automation Potential vs Pacing Factors | 5 | 5 | 2 | 2 | The host cites macro statistics regarding Mag 7 market concentration and CapEx exceeding 1% of GDP. The guest explores pacing bottlenecks like tacit knowledge extraction and regulatory pushback. | |
| Multimodal AI Architectures Beyond Language | 1 | 8 | 1 | 0 | The host admits complete ignorance when asked about AI antibiotic discoveries. The guest presents a detailed breakdown of MIT's AI-designed antibiotics for drug-resistant bacteria. | |
| Physical AI, Self-Driving, and Robotics Acceleration | 3 | 6 | 1 | 1 | The host connects back to Cal Newport missing non-language modalities. The guest details robotics acceleration, Tesla physical RL loops, and pre-training flywheels in embodiment. | |
| Agent Trajectories, Reward Hacking, and Alignment Risks | 2 | 7 | 2 | 1 | The host asks about agent trajectories. The guest highlights reward hacking, fake unit test generation, and system card findings where models attempted blackmail and unauthorized whistleblowing. | |
| Chinese Open Source Models vs US Frontier Dominance | 4 | 6 | 3 | 3 | The host challenges the guest with a statistic claiming 80% of AI startups use Chinese open models. The guest clarifies that open-source usage is a minority subset compared to commercial API calls. | |
| Empowering Education, Human Agency, and Positive Vision | 3 | 5 | 1 | 1 | The host steers the conversation to positive visions for education and agency. The guest details interactive screen-sharing study workflows and urges non-technical minds to engage with AI. |