Jul 31, 2025 · 1h 18m · latent-space

The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

Nathan Lambert · 54m spoken Shawn Wang · 13m spoken Alessio Fanelli · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode, AI2 post-training researcher Nathan Lambert joins hosts Alessio Fanelli and Swix to break down the rise of Reinforcement Learning from Verifiable Rewards (RLVR), the mechanics of test-time compute scaling, and the taxonomy of modern reasoning architectures. Lambert explores agent tool use, reward hacking pathologies, and AI2's strategic roadmap for developing transparent, frontier-grade open-source models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 26.1% of the talking time here. How this is scored →

The hosts as informed peer 5.5 Guest teaching 4.1 Guest disagreement 1.7 The hosts pushing back 1.9
05100:0020:0040:001:00:000:04–6:07 · The hosts as informed peer 4/10 Welcome and Catching Up with Nathan Lambert Friendly intro where Swyx and Alessio establish Nathan's background before Nathan recounts the development of Tulu 3 and the coining of RLVR. The hosts prompt him smoothly with conversational context.6:07–9:04 · The hosts as informed peer 5/10 Expanding RLVR into Multi-Hop Tool Environments Nathan breaks down the transition from single-generation RLVR to multi-hop environments like o3 and Deep Research. Alessio brings in insights from a previous conversation with Noam Brown regarding non-verifiable rewards.9:04–12:47 · The hosts as informed peer 5/10 Agent Progress, Non-Verifiable Tasks, and Data Moats Nathan explains why agent progress will look different from raw modeling progress and details the limitations of academic preference datasets compared to frontier lab data moats. Alessio questions whether frontier labs retain a structural edge via live inference inspection.12:47–15:39 · The hosts as informed peer 6/10 Evaluating the Longevity of Chatbot Arena Benchmarks Swyx presses on whether Chatbot Arena is 'cooked' and discusses vulnerabilities to gaming and the lack of multi-turn evals, citing recent critiques like Sarah Hooker's. Nathan defends the ongoing utility of Elo leaderboards for compression benchmarking.15:39–20:22 · The hosts as informed peer 5/10 Writing the RLHF Book Amid the Reasoning Revolution Nathan articulates why he is keeping his upcoming book focused on RLHF rather than RLVR, arguing RLHF deals with enduring alignment and human preference problems while RLVR may be solved quickly. Swyx probes his dismissal of GRPO algorithmic hype.20:22–27:35 · The hosts as informed peer 6/10 Hybrid Reasoning Models versus Unified Interfaces Alessio and Swyx challenge the necessity of hybrid reasoning models versus unified interfaces that self-allocate compute. Nathan outlines why search-augmented models degrade on raw SimpleQA when isolated from external search retrieval.27:35–35:03 · The hosts as informed peer 6/10 Deconstructing Q*, o1, and Inference Scaling Plots Nathan calls the standard inference-time scaling plot a 'psyop' because plotting test compute on a log axis misleads people into treating search as a simple linear knob. Alessio and Swyx discuss the ARC-AGI benchmark harness controversies.35:03–37:51 · The hosts as informed peer 5/10 Strategic Research Opportunities for Academic AI Nathan lays out realistic strategies for academic AI institutions like AI2, emphasizing domain-specific evals and artifact production over trying to match frontier labs on raw token burn.37:51–44:24 · The hosts as informed peer 6/10 A Taxonomy of Reasoning: Skills, Strategy, Abstraction, Calibration Nathan introduces his taxonomy of reasoning: skills, strategy, abstraction, and calibration. Swyx pushes back, questioning whether dedicated plan tokens are necessary when engineers prefer external tool orchestration.44:24–46:57 · The hosts as informed peer 5/10 Task Decomposition, Planning Tokens, and Compute Allocation Alessio questions whether plan templates could be reused as static blueprints rather than regenerated dynamically on every query. Nathan details how prompt generation remains more cost-effective despite potential abstraction failures.46:57–51:53 · The hosts as informed peer 6/10 Parallel Test-Time Compute, Verifiers, and Diffusion LLMs Swyx pushes back on Nathan's skepticism of parallel test-time compute, arguing it allows teams to pull forward future model capabilities to generate synthetic data for distillation.51:53–54:34 · The hosts as informed peer 7/10 Code Generation Pathologies from Reinforcement Learning Alessio brings concrete technical observations from using coding tools, noting that RL induces models to write defensive if-statements to suppress missing environment variables rather than raising clear errors.54:34–1:00:39 · The hosts as informed peer 6/10 The Three Eras of Reinforcement Learning Over-Optimization Nathan walks through the three historic eras of RL reward over-optimization across robotics control, RLHF, and RLVR. Swyx asks about the nuances of reward design and partial credit in multi-domain setups.1:00:39–1:07:54 · The hosts as informed peer 6/10 Infrastructural Constraints and Long Inference Trajectories Swyx and Nathan discuss training run limits and wall-clock feedback bottlenecks, referencing Swyx's earlier debate with Noam Brown. Nathan discusses Joanne Jang's Model Spec work at OpenAI.1:07:54–1:13:05 · The hosts as informed peer 6/10 Micro-Distillation, Hardware Pricing, and AI Wearables Alessio and Nathan discuss routing across micro-distilled models for specialized media tasks. Swyx pushes back, claiming multimodal models like GPT-4o will eventually absorb all point solutions, while noting the unexpected appreciation of RTX 4090 GPUs.1:13:05–1:15:47 · The hosts as informed peer 5/10 Meta's Strategic Pivot and the Economics of AI Talent Swyx and Nathan examine Meta's aggressive compensation packages for AI talent, analyzing the economic trade-off between massive cluster compute expenditures and elite researcher retention.1:15:47–1:18:45 · The hosts as informed peer 5/10 AI2's Roadmap for an Open 'American DeepSeek' Nathan outlines AI2's long-term technical roadmap for an open American equivalent to DeepSeek, moving OLMo from dense architectures to sparse MoE reasoning models.0:04–6:07 · Guest teaching 3/10 Welcome and Catching Up with Nathan Lambert Friendly intro where Swyx and Alessio establish Nathan's background before Nathan recounts the development of Tulu 3 and the coining of RLVR. The hosts prompt him smoothly with conversational context.6:07–9:04 · Guest teaching 4/10 Expanding RLVR into Multi-Hop Tool Environments Nathan breaks down the transition from single-generation RLVR to multi-hop environments like o3 and Deep Research. Alessio brings in insights from a previous conversation with Noam Brown regarding non-verifiable rewards.9:04–12:47 · Guest teaching 5/10 Agent Progress, Non-Verifiable Tasks, and Data Moats Nathan explains why agent progress will look different from raw modeling progress and details the limitations of academic preference datasets compared to frontier lab data moats. Alessio questions whether frontier labs retain a structural edge via live inference inspection.12:47–15:39 · Guest teaching 3/10 Evaluating the Longevity of Chatbot Arena Benchmarks Swyx presses on whether Chatbot Arena is 'cooked' and discusses vulnerabilities to gaming and the lack of multi-turn evals, citing recent critiques like Sarah Hooker's. Nathan defends the ongoing utility of Elo leaderboards for compression benchmarking.15:39–20:22 · Guest teaching 5/10 Writing the RLHF Book Amid the Reasoning Revolution Nathan articulates why he is keeping his upcoming book focused on RLHF rather than RLVR, arguing RLHF deals with enduring alignment and human preference problems while RLVR may be solved quickly. Swyx probes his dismissal of GRPO algorithmic hype.20:22–27:35 · Guest teaching 5/10 Hybrid Reasoning Models versus Unified Interfaces Alessio and Swyx challenge the necessity of hybrid reasoning models versus unified interfaces that self-allocate compute. Nathan outlines why search-augmented models degrade on raw SimpleQA when isolated from external search retrieval.27:35–35:03 · Guest teaching 4/10 Deconstructing Q*, o1, and Inference Scaling Plots Nathan calls the standard inference-time scaling plot a 'psyop' because plotting test compute on a log axis misleads people into treating search as a simple linear knob. Alessio and Swyx discuss the ARC-AGI benchmark harness controversies.35:03–37:51 · Guest teaching 4/10 Strategic Research Opportunities for Academic AI Nathan lays out realistic strategies for academic AI institutions like AI2, emphasizing domain-specific evals and artifact production over trying to match frontier labs on raw token burn.37:51–44:24 · Guest teaching 4/10 A Taxonomy of Reasoning: Skills, Strategy, Abstraction, Calibration Nathan introduces his taxonomy of reasoning: skills, strategy, abstraction, and calibration. Swyx pushes back, questioning whether dedicated plan tokens are necessary when engineers prefer external tool orchestration.44:24–46:57 · Guest teaching 4/10 Task Decomposition, Planning Tokens, and Compute Allocation Alessio questions whether plan templates could be reused as static blueprints rather than regenerated dynamically on every query. Nathan details how prompt generation remains more cost-effective despite potential abstraction failures.46:57–51:53 · Guest teaching 4/10 Parallel Test-Time Compute, Verifiers, and Diffusion LLMs Swyx pushes back on Nathan's skepticism of parallel test-time compute, arguing it allows teams to pull forward future model capabilities to generate synthetic data for distillation.51:53–54:34 · Guest teaching 3/10 Code Generation Pathologies from Reinforcement Learning Alessio brings concrete technical observations from using coding tools, noting that RL induces models to write defensive if-statements to suppress missing environment variables rather than raising clear errors.54:34–1:00:39 · Guest teaching 5/10 The Three Eras of Reinforcement Learning Over-Optimization Nathan walks through the three historic eras of RL reward over-optimization across robotics control, RLHF, and RLVR. Swyx asks about the nuances of reward design and partial credit in multi-domain setups.1:00:39–1:07:54 · Guest teaching 4/10 Infrastructural Constraints and Long Inference Trajectories Swyx and Nathan discuss training run limits and wall-clock feedback bottlenecks, referencing Swyx's earlier debate with Noam Brown. Nathan discusses Joanne Jang's Model Spec work at OpenAI.1:07:54–1:13:05 · Guest teaching 4/10 Micro-Distillation, Hardware Pricing, and AI Wearables Alessio and Nathan discuss routing across micro-distilled models for specialized media tasks. Swyx pushes back, claiming multimodal models like GPT-4o will eventually absorb all point solutions, while noting the unexpected appreciation of RTX 4090 GPUs.1:13:05–1:15:47 · Guest teaching 4/10 Meta's Strategic Pivot and the Economics of AI Talent Swyx and Nathan examine Meta's aggressive compensation packages for AI talent, analyzing the economic trade-off between massive cluster compute expenditures and elite researcher retention.1:15:47–1:18:45 · Guest teaching 4/10 AI2's Roadmap for an Open 'American DeepSeek' Nathan outlines AI2's long-term technical roadmap for an open American equivalent to DeepSeek, moving OLMo from dense architectures to sparse MoE reasoning models.0:04–6:07 · Guest disagreement 1/10 Welcome and Catching Up with Nathan Lambert Friendly intro where Swyx and Alessio establish Nathan's background before Nathan recounts the development of Tulu 3 and the coining of RLVR. The hosts prompt him smoothly with conversational context.6:07–9:04 · Guest disagreement 1/10 Expanding RLVR into Multi-Hop Tool Environments Nathan breaks down the transition from single-generation RLVR to multi-hop environments like o3 and Deep Research. Alessio brings in insights from a previous conversation with Noam Brown regarding non-verifiable rewards.9:04–12:47 · Guest disagreement 2/10 Agent Progress, Non-Verifiable Tasks, and Data Moats Nathan explains why agent progress will look different from raw modeling progress and details the limitations of academic preference datasets compared to frontier lab data moats. Alessio questions whether frontier labs retain a structural edge via live inference inspection.12:47–15:39 · Guest disagreement 2/10 Evaluating the Longevity of Chatbot Arena Benchmarks Swyx presses on whether Chatbot Arena is 'cooked' and discusses vulnerabilities to gaming and the lack of multi-turn evals, citing recent critiques like Sarah Hooker's. Nathan defends the ongoing utility of Elo leaderboards for compression benchmarking.15:39–20:22 · Guest disagreement 2/10 Writing the RLHF Book Amid the Reasoning Revolution Nathan articulates why he is keeping his upcoming book focused on RLHF rather than RLVR, arguing RLHF deals with enduring alignment and human preference problems while RLVR may be solved quickly. Swyx probes his dismissal of GRPO algorithmic hype.20:22–27:35 · Guest disagreement 2/10 Hybrid Reasoning Models versus Unified Interfaces Alessio and Swyx challenge the necessity of hybrid reasoning models versus unified interfaces that self-allocate compute. Nathan outlines why search-augmented models degrade on raw SimpleQA when isolated from external search retrieval.27:35–35:03 · Guest disagreement 3/10 Deconstructing Q*, o1, and Inference Scaling Plots Nathan calls the standard inference-time scaling plot a 'psyop' because plotting test compute on a log axis misleads people into treating search as a simple linear knob. Alessio and Swyx discuss the ARC-AGI benchmark harness controversies.35:03–37:51 · Guest disagreement 1/10 Strategic Research Opportunities for Academic AI Nathan lays out realistic strategies for academic AI institutions like AI2, emphasizing domain-specific evals and artifact production over trying to match frontier labs on raw token burn.37:51–44:24 · Guest disagreement 2/10 A Taxonomy of Reasoning: Skills, Strategy, Abstraction, Calibration Nathan introduces his taxonomy of reasoning: skills, strategy, abstraction, and calibration. Swyx pushes back, questioning whether dedicated plan tokens are necessary when engineers prefer external tool orchestration.44:24–46:57 · Guest disagreement 2/10 Task Decomposition, Planning Tokens, and Compute Allocation Alessio questions whether plan templates could be reused as static blueprints rather than regenerated dynamically on every query. Nathan details how prompt generation remains more cost-effective despite potential abstraction failures.46:57–51:53 · Guest disagreement 2/10 Parallel Test-Time Compute, Verifiers, and Diffusion LLMs Swyx pushes back on Nathan's skepticism of parallel test-time compute, arguing it allows teams to pull forward future model capabilities to generate synthetic data for distillation.51:53–54:34 · Guest disagreement 1/10 Code Generation Pathologies from Reinforcement Learning Alessio brings concrete technical observations from using coding tools, noting that RL induces models to write defensive if-statements to suppress missing environment variables rather than raising clear errors.54:34–1:00:39 · Guest disagreement 1/10 The Three Eras of Reinforcement Learning Over-Optimization Nathan walks through the three historic eras of RL reward over-optimization across robotics control, RLHF, and RLVR. Swyx asks about the nuances of reward design and partial credit in multi-domain setups.1:00:39–1:07:54 · Guest disagreement 2/10 Infrastructural Constraints and Long Inference Trajectories Swyx and Nathan discuss training run limits and wall-clock feedback bottlenecks, referencing Swyx's earlier debate with Noam Brown. Nathan discusses Joanne Jang's Model Spec work at OpenAI.1:07:54–1:13:05 · Guest disagreement 2/10 Micro-Distillation, Hardware Pricing, and AI Wearables Alessio and Nathan discuss routing across micro-distilled models for specialized media tasks. Swyx pushes back, claiming multimodal models like GPT-4o will eventually absorb all point solutions, while noting the unexpected appreciation of RTX 4090 GPUs.1:13:05–1:15:47 · Guest disagreement 2/10 Meta's Strategic Pivot and the Economics of AI Talent Swyx and Nathan examine Meta's aggressive compensation packages for AI talent, analyzing the economic trade-off between massive cluster compute expenditures and elite researcher retention.1:15:47–1:18:45 · Guest disagreement 1/10 AI2's Roadmap for an Open 'American DeepSeek' Nathan outlines AI2's long-term technical roadmap for an open American equivalent to DeepSeek, moving OLMo from dense architectures to sparse MoE reasoning models.0:04–6:07 · The hosts pushing back 1/10 Welcome and Catching Up with Nathan Lambert Friendly intro where Swyx and Alessio establish Nathan's background before Nathan recounts the development of Tulu 3 and the coining of RLVR. The hosts prompt him smoothly with conversational context.6:07–9:04 · The hosts pushing back 1/10 Expanding RLVR into Multi-Hop Tool Environments Nathan breaks down the transition from single-generation RLVR to multi-hop environments like o3 and Deep Research. Alessio brings in insights from a previous conversation with Noam Brown regarding non-verifiable rewards.9:04–12:47 · The hosts pushing back 2/10 Agent Progress, Non-Verifiable Tasks, and Data Moats Nathan explains why agent progress will look different from raw modeling progress and details the limitations of academic preference datasets compared to frontier lab data moats. Alessio questions whether frontier labs retain a structural edge via live inference inspection.12:47–15:39 · The hosts pushing back 3/10 Evaluating the Longevity of Chatbot Arena Benchmarks Swyx presses on whether Chatbot Arena is 'cooked' and discusses vulnerabilities to gaming and the lack of multi-turn evals, citing recent critiques like Sarah Hooker's. Nathan defends the ongoing utility of Elo leaderboards for compression benchmarking.15:39–20:22 · The hosts pushing back 2/10 Writing the RLHF Book Amid the Reasoning Revolution Nathan articulates why he is keeping his upcoming book focused on RLHF rather than RLVR, arguing RLHF deals with enduring alignment and human preference problems while RLVR may be solved quickly. Swyx probes his dismissal of GRPO algorithmic hype.20:22–27:35 · The hosts pushing back 2/10 Hybrid Reasoning Models versus Unified Interfaces Alessio and Swyx challenge the necessity of hybrid reasoning models versus unified interfaces that self-allocate compute. Nathan outlines why search-augmented models degrade on raw SimpleQA when isolated from external search retrieval.27:35–35:03 · The hosts pushing back 2/10 Deconstructing Q*, o1, and Inference Scaling Plots Nathan calls the standard inference-time scaling plot a 'psyop' because plotting test compute on a log axis misleads people into treating search as a simple linear knob. Alessio and Swyx discuss the ARC-AGI benchmark harness controversies.35:03–37:51 · The hosts pushing back 1/10 Strategic Research Opportunities for Academic AI Nathan lays out realistic strategies for academic AI institutions like AI2, emphasizing domain-specific evals and artifact production over trying to match frontier labs on raw token burn.37:51–44:24 · The hosts pushing back 3/10 A Taxonomy of Reasoning: Skills, Strategy, Abstraction, Calibration Nathan introduces his taxonomy of reasoning: skills, strategy, abstraction, and calibration. Swyx pushes back, questioning whether dedicated plan tokens are necessary when engineers prefer external tool orchestration.44:24–46:57 · The hosts pushing back 2/10 Task Decomposition, Planning Tokens, and Compute Allocation Alessio questions whether plan templates could be reused as static blueprints rather than regenerated dynamically on every query. Nathan details how prompt generation remains more cost-effective despite potential abstraction failures.46:57–51:53 · The hosts pushing back 3/10 Parallel Test-Time Compute, Verifiers, and Diffusion LLMs Swyx pushes back on Nathan's skepticism of parallel test-time compute, arguing it allows teams to pull forward future model capabilities to generate synthetic data for distillation.51:53–54:34 · The hosts pushing back 2/10 Code Generation Pathologies from Reinforcement Learning Alessio brings concrete technical observations from using coding tools, noting that RL induces models to write defensive if-statements to suppress missing environment variables rather than raising clear errors.54:34–1:00:39 · The hosts pushing back 1/10 The Three Eras of Reinforcement Learning Over-Optimization Nathan walks through the three historic eras of RL reward over-optimization across robotics control, RLHF, and RLVR. Swyx asks about the nuances of reward design and partial credit in multi-domain setups.1:00:39–1:07:54 · The hosts pushing back 2/10 Infrastructural Constraints and Long Inference Trajectories Swyx and Nathan discuss training run limits and wall-clock feedback bottlenecks, referencing Swyx's earlier debate with Noam Brown. Nathan discusses Joanne Jang's Model Spec work at OpenAI.1:07:54–1:13:05 · The hosts pushing back 3/10 Micro-Distillation, Hardware Pricing, and AI Wearables Alessio and Nathan discuss routing across micro-distilled models for specialized media tasks. Swyx pushes back, claiming multimodal models like GPT-4o will eventually absorb all point solutions, while noting the unexpected appreciation of RTX 4090 GPUs.1:13:05–1:15:47 · The hosts pushing back 2/10 Meta's Strategic Pivot and the Economics of AI Talent Swyx and Nathan examine Meta's aggressive compensation packages for AI talent, analyzing the economic trade-off between massive cluster compute expenditures and elite researcher retention.1:15:47–1:18:45 · The hosts pushing back 1/10 AI2's Roadmap for an Open 'American DeepSeek' Nathan outlines AI2's long-term technical roadmap for an open American equivalent to DeepSeek, moving OLMo from dense architectures to sparse MoE reasoning models.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 45% · guest 55%0:00 · the hosts 45% · guest 55%3:00 · the hosts 2.2% · guest 97.8%3:00 · the hosts 2.2% · guest 97.8%6:00 · the hosts 22.5% · guest 77.5%6:00 · the hosts 22.5% · guest 77.5%9:00 · the hosts 11.3% · guest 88.7%9:00 · the hosts 11.3% · guest 88.7%12:00 · the hosts 26.9% · guest 73.1%12:00 · the hosts 26.9% · guest 73.1%15:00 · the hosts 14.8% · guest 85.2%15:00 · the hosts 14.8% · guest 85.2%18:00 · the hosts 23.3% · guest 76.7%18:00 · the hosts 23.3% · guest 76.7%21:00 · the hosts 23.9% · guest 76.1%21:00 · the hosts 23.9% · guest 76.1%24:00 · the hosts 22% · guest 78%24:00 · the hosts 22% · guest 78%27:00 · the hosts 40.4% · guest 59.6%27:00 · the hosts 40.4% · guest 59.6%30:00 · the hosts 13.2% · guest 86.8%30:00 · the hosts 13.2% · guest 86.8%33:00 · the hosts 42.6% · guest 57.4%33:00 · the hosts 42.6% · guest 57.4%36:00 · the hosts 6.3% · guest 93.7%36:00 · the hosts 6.3% · guest 93.7%39:00 · the hosts 13.7% · guest 86.3%39:00 · the hosts 13.7% · guest 86.3%42:00 · the hosts 34.3% · guest 65.7%42:00 · the hosts 34.3% · guest 65.7%45:00 · the hosts 32.6% · guest 67.4%45:00 · the hosts 32.6% · guest 67.4%48:00 · the hosts 15.8% · guest 84.2%48:00 · the hosts 15.8% · guest 84.2%51:00 · the hosts 46.1% · guest 53.9%51:00 · the hosts 46.1% · guest 53.9%54:00 · the hosts 21.7% · guest 78.3%54:00 · the hosts 21.7% · guest 78.3%57:00 · the hosts 11.2% · guest 88.8%57:00 · the hosts 11.2% · guest 88.8%1:00:00 · the hosts 48.3% · guest 51.7%1:00:00 · the hosts 48.3% · guest 51.7%1:03:00 · the hosts 25.9% · guest 74.1%1:03:00 · the hosts 25.9% · guest 74.1%1:06:00 · the hosts 34.1% · guest 65.9%1:06:00 · the hosts 34.1% · guest 65.9%1:09:00 · the hosts 50.5% · guest 49.5%1:09:00 · the hosts 50.5% · guest 49.5%1:12:00 · the hosts 52.8% · guest 47.2%1:12:00 · the hosts 52.8% · guest 47.2%1:15:00 · the hosts 4.7% · guest 95.3%1:15:00 · the hosts 4.7% · guest 95.3%1:18:00 · the hosts 11.7% · guest 88.3%1:18:00 · the hosts 11.7% · guest 88.3%
Sharpest disagreement ▶ 28:17 Inference scaling plots called a psyop

Nathan aggressively rejects the standard industry presentation of test-time compute, labeling the popular inference-time scaling curve a deceptive marketing construct.

Hardest push from the hosts ▶ 41:40 Swyx rejects native plan tokens

Swyx directly challenges Nathan's reasoning taxonomy by arguing that modern software engineering workflows favor external tool orchestration over internal model plan tokens.

Biggest teaching moment ▶ 16:47 Why RLHF outlasts RLVR as a research discipline

Nathan educates the hosts on why preference tuning and RLHF will remain foundational research challenges indefinitely while verifiable reward RLVR may quickly saturate.

The host holds their own ▶ 52:10 Alessio diagnoses RL code pathologies

Alessio demonstrates deep practitioner expertise by detailing how RL-trained models introduce silent failure patterns through defensive if-statements in real-world codebases.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Welcome and Catching Up with Nathan Lambert 4311 Friendly intro where Swyx and Alessio establish Nathan's background before Nathan recounts the development of Tulu 3 and the coining of RLVR. The hosts prompt him smoothly with conversational context.
Expanding RLVR into Multi-Hop Tool Environments 5411 Nathan breaks down the transition from single-generation RLVR to multi-hop environments like o3 and Deep Research. Alessio brings in insights from a previous conversation with Noam Brown regarding non-verifiable rewards.
Agent Progress, Non-Verifiable Tasks, and Data Moats 5522 Nathan explains why agent progress will look different from raw modeling progress and details the limitations of academic preference datasets compared to frontier lab data moats. Alessio questions whether frontier labs retain a structural edge via live inference inspection.
Evaluating the Longevity of Chatbot Arena Benchmarks 6323 Swyx presses on whether Chatbot Arena is 'cooked' and discusses vulnerabilities to gaming and the lack of multi-turn evals, citing recent critiques like Sarah Hooker's. Nathan defends the ongoing utility of Elo leaderboards for compression benchmarking.
Writing the RLHF Book Amid the Reasoning Revolution 5522 Nathan articulates why he is keeping his upcoming book focused on RLHF rather than RLVR, arguing RLHF deals with enduring alignment and human preference problems while RLVR may be solved quickly. Swyx probes his dismissal of GRPO algorithmic hype.
Hybrid Reasoning Models versus Unified Interfaces 6522 Alessio and Swyx challenge the necessity of hybrid reasoning models versus unified interfaces that self-allocate compute. Nathan outlines why search-augmented models degrade on raw SimpleQA when isolated from external search retrieval.
Deconstructing Q*, o1, and Inference Scaling Plots 6432 Nathan calls the standard inference-time scaling plot a 'psyop' because plotting test compute on a log axis misleads people into treating search as a simple linear knob. Alessio and Swyx discuss the ARC-AGI benchmark harness controversies.
Strategic Research Opportunities for Academic AI 5411 Nathan lays out realistic strategies for academic AI institutions like AI2, emphasizing domain-specific evals and artifact production over trying to match frontier labs on raw token burn.
A Taxonomy of Reasoning: Skills, Strategy, Abstraction, Calibration 6423 Nathan introduces his taxonomy of reasoning: skills, strategy, abstraction, and calibration. Swyx pushes back, questioning whether dedicated plan tokens are necessary when engineers prefer external tool orchestration.
Task Decomposition, Planning Tokens, and Compute Allocation 5422 Alessio questions whether plan templates could be reused as static blueprints rather than regenerated dynamically on every query. Nathan details how prompt generation remains more cost-effective despite potential abstraction failures.
Parallel Test-Time Compute, Verifiers, and Diffusion LLMs 6423 Swyx pushes back on Nathan's skepticism of parallel test-time compute, arguing it allows teams to pull forward future model capabilities to generate synthetic data for distillation.
Code Generation Pathologies from Reinforcement Learning 7312 Alessio brings concrete technical observations from using coding tools, noting that RL induces models to write defensive if-statements to suppress missing environment variables rather than raising clear errors.
The Three Eras of Reinforcement Learning Over-Optimization 6511 Nathan walks through the three historic eras of RL reward over-optimization across robotics control, RLHF, and RLVR. Swyx asks about the nuances of reward design and partial credit in multi-domain setups.
Infrastructural Constraints and Long Inference Trajectories 6422 Swyx and Nathan discuss training run limits and wall-clock feedback bottlenecks, referencing Swyx's earlier debate with Noam Brown. Nathan discusses Joanne Jang's Model Spec work at OpenAI.
Micro-Distillation, Hardware Pricing, and AI Wearables 6423 Alessio and Nathan discuss routing across micro-distilled models for specialized media tasks. Swyx pushes back, claiming multimodal models like GPT-4o will eventually absorb all point solutions, while noting the unexpected appreciation of RTX 4090 GPUs.
Meta's Strategic Pivot and the Economics of AI Talent 5422 Swyx and Nathan examine Meta's aggressive compensation packages for AI talent, analyzing the economic trade-off between massive cluster compute expenditures and elite researcher retention.
AI2's Roadmap for an Open 'American DeepSeek' 5411 Nathan outlines AI2's long-term technical roadmap for an open American equivalent to DeepSeek, moving OLMo from dense architectures to sparse MoE reasoning models.

Statements from this episode (40)

Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20
Assertion Not checkable as stated
Lambert: Academia relied on UltraFeedback for open preference tuning for a year
“The academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular. And still a year later is like this state of the art data set for open preferen…”
Nathan Lambert Jul 31, 2025 ▶ 3:15
Insight
Lambert: RLVR is broader than ground truth because code is verifiable
“The verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.”
Nathan Lambert Jul 31, 2025 ▶ 5:13
Opinion
Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that whe…”
Nathan Lambert Jul 31, 2025 ▶ 7:35
Insight
Lambert: Context compression is crucial for long-horizon AI agents
“Compressing context, like that's not I don't think that's really a verifiable thing, but that being messed up, like that's a super crucial skill for long context actions and long longer tasks is just compressing well, and that's going to take some training nov…”
Nathan Lambert Jul 31, 2025 ▶ 9:50
Assertion Supported
Lambert: Frontier AI labs still rely on human preference data
“Every time I check in with people at frontier labs, they're like, yeah, we still use human preference data.”
Nathan Lambert Jul 31, 2025 ▶ 12:12
Prediction Open · timeframe Jul 2028
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Nathan Lambert Jul 31, 2025 ▶ 15:03
Insight
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Nathan Lambert Jul 31, 2025 ▶ 16:42
Insight
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Nathan Lambert Jul 31, 2025 ▶ 19:28
Insight
Lambert: North star of reasoning models is dynamic token budget calibration
“I think that has to be the north star for most people working on reasoning, which is the model will just Spend the right amount of tokens on it.”
Nathan Lambert Jul 31, 2025 ▶ 20:44
Prediction Not checkable as stated
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Nathan Lambert Jul 31, 2025 ▶ 20:52
Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Nathan Lambert Jul 31, 2025 ▶ 22:59
Prediction Open · timeframe Jul 2030
Lambert: All major frontier AI labs will build their own search indexes
“I think they'll all do end up doing their own index and it should, it's one of those things that's like Google should have an advantage again, but who knows if they do.”
Nathan Lambert Jul 31, 2025 ▶ 24:23
Insight
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Nathan Lambert Jul 31, 2025 ▶ 24:35
Assertion Not checkable as stated
Swix: OpenAI Deep Research was built by three people as an o3 wrapper
“As far as I know, it's three people did it. It was Isa and like the two other collaborators that she had. I don't know if they did that much on top of all three, like every indication I've had from over the eye is that deep research is more or less a thin wrap…”
Shawn Wang Jul 31, 2025 ▶ 25:36
Assertion Supported
Fanelli: Simon Willison reported Anthropic added Brave Search as a sub-processor
“Our friend Simon Willison wrote a post that Anthropic added Brave Search as one of the sub processor in their product.”
Alessio Fanelli Jul 31, 2025 ▶ 27:36
Insight
Lambert: Inference scaling plots misleadingly suggest search is an easy control knob
“The core of that article is just, they're taking points from within training, or there's a natural variance, and then you line them up. And if you line them up, then you get this nice inference time-scaling behavior, which is, and now people, a lot of people h…”
Nathan Lambert Jul 31, 2025 ▶ 28:59
Insight
Lambert: Scaling RL long enough requires a curriculum of increasing difficulty
“If you scale RL long enough, You're going to need a curriculum of things getting harder. And like, that's pretty obvious.”
Nathan Lambert Jul 31, 2025 ▶ 32:35
Opinion
Lambert: AI benchmarks like ARC-AGI should prioritize testing without harnesses
“Harnesses are cool, but they're gonna, they're, They're a handicap that's changing the learning dynamics substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses.”
Nathan Lambert Jul 31, 2025 ▶ 34:02
Insight
Lambert: AI academics must build datasets and evals rather than papers
“If you're trying to have impact in AI right now, it's as an academic, you have to like level up out of papers to artifacts, which is models, datasets, evals. Datasets and evals are easier for people to have impact on.”
Nathan Lambert Jul 31, 2025 ▶ 36:13
Prediction Open · timeframe Jul 2028
Lambert: Academics cannot match industry compute on Humanity's Last Exam
“I just think it's kind of unlikely that we're going to win as a academic and a state of the art number because they're going to start spending millions of tokens per query. And it's just a lot of, it's a lot of compute burn. Like the getting, beating that on t…”
Nathan Lambert Jul 31, 2025 ▶ 37:12
Insight
Lambert: Reasoning models solved basic skills; planning is the next frontier
“So I came up with four and the foundational one was skills, which is What I would say that we have already done with O-one and R-one, which is you do a lot of RL, you show the inference time scaling works and you get really high benchmark numbers. And then the…”
Nathan Lambert Jul 31, 2025 ▶ 38:34
Assertion Supported
Lambert: DeepSeek-R1 starts solving math questions immediately without explicit planning
“If you look at DeepSeq R-One and you ask it a hard math question, it's not like, here's my plan of attack. It just starts.”
Nathan Lambert Jul 31, 2025 ▶ 40:10
Assertion Not checkable as stated
Lambert: Current Language Models Cannot Prioritize Experiments for Multi-Week Research Plans
“So it's like, how do you come up with a research plan in 10 weeks? Like there's a lot of, how do you prioritize which experiments to do? It's like, there's a lot of inductive biases that go into that, that I don't like a language model would not do well at tha…”
Nathan Lambert Jul 31, 2025 ▶ 46:31
Insight
Lambert: Parallel compute provides robustness rather than low-probability search
“Well, I don't think we're using parallel compute in a way to search over like low probability tokens. We're using it to get robustness. If you use like O-one pro is, it was so nice because it Just had a very predictable depth to it, even on niche topics where …”
Nathan Lambert Jul 31, 2025 ▶ 47:45
Prediction Held up
Lambert: Labs will surely use parallel-compute models to generate synthetic data
“Well, I bet people, I mean, they surely will use these for synthetic data. It's just like the marginal gain on synthetic data is always very high.”
Nathan Lambert Jul 31, 2025 ▶ 50:22
Opinion
Lambert: Labs trade code usability for massive RL performance gains
“That's just like the labs are trading off massive gains in performance or small detriments in usability. And it's like, do you ship that model? Yeah. Like you just ship it and deal with it later, but I'm sure they could, I'm sure that's a fixable thing.”
Nathan Lambert Jul 31, 2025 ▶ 52:44
Insight
Lambert: Code maintainability is a human preference problem in RL
“The software stuff is not easy because it's almost like maintainability almost feels like a human preference type issue again, where somebody could look at it and be like, yeah, that's not as good, but adding the heuristic and trading seems very messy.”
Nathan Lambert Jul 31, 2025 ▶ 53:24
Insight
Lambert: RLVR is harder to over-optimize on math than code
“For math, it's a bit harder to over optimize, I think. Unless you have tools and the model learns to search and cheat instead of learning math, which I'm sure somebody could see that out in the world, which is like, oh, I'll just find the, you're training. It'…”
Nathan Lambert Jul 31, 2025 ▶ 57:25
Insight
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Nathan Lambert Jul 31, 2025 ▶ 59:33
Assertion Not checkable as stated
Lambert: Long inference generations break RL infrastructure and require more GPUs
“The inference, high inference length generations definitely just, like, kind of breaks all infrastructure, because there's just so many tokens, there's more opportunity for out of memory or other things to go wrong. So it's like, just on a default, all of your…”
Nathan Lambert Jul 31, 2025 ▶ 1:00:39
Insight
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Nathan Lambert Jul 31, 2025 ▶ 1:03:38
Insight
Lambert: Custom personality fine-tuning is open source AI's winning turf
“If open models are to win, part of it could be just, like, everybody can have exactly the model they want. We're serving GPT-IV. It's kind of its thing. You can prompt it, but if fine-tuning is more effective than prompting, everybody can have the model. That …”
Nathan Lambert Jul 31, 2025 ▶ 1:05:05
Opinion
Lambert: The local model community is much smaller than assumed
“Like the local modeling community, I think is much smaller than people give it credit for, because most of the use for open models is still in APIs.”
Nathan Lambert Jul 31, 2025 ▶ 1:09:26
Assertion Partly supported
Swix: Nvidia RTX 4090 prices doubled in the past year
“40 and 90 prices have doubled in the last year.”
Shawn Wang Jul 31, 2025 ▶ 1:10:08
Prediction Not checkable as stated
Lambert: OpenAI's open model will be best-in-class in its size category
“I expected. It'll be best in class for some size Category in some subset of tasks. That's like, OpenAI only does things like that.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:07
Prediction Open · timeframe Jul 2028
Lambert: Jony Ive and OpenAI hardware will run in the cloud
“I think that thing will run on the cloud. I don't think that'll run local anyways.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:54
Opinion
Lambert: Meta withholding its leading benchmark model is bad execution
“But to be a model that claims to be open and then not release the model that is your leading claim is just, like, that is, like, bad execution.”
Nathan Lambert Jul 31, 2025 ▶ 1:13:36
Insight
Lambert: Top AI talent is dramatically cheaper than GPU clusters
“Talent is cheaper than GPUs by a dramatic margin, and At the end of the day, it's like, okay, if we're spending this much, they go to the room and they stare in the mirror and you're like, wait, it might not actually be that ridiculous to spend this money on t…”
Nathan Lambert Jul 31, 2025 ▶ 1:13:46
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Nathan Lambert Jul 31, 2025 ▶ 1:16:23
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.