Jul 31, 2025 · 1h 18m · latent-space
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode, AI2 post-training researcher Nathan Lambert joins hosts Alessio Fanelli and Swix to break down the rise of Reinforcement Learning from Verifiable Rewards (RLVR), the mechanics of test-time compute scaling, and the taxonomy of modern reasoning architectures. Lambert explores agent tool use, reward hacking pathologies, and AI2's strategic roadmap for developing transparent, frontier-grade open-source models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 26.1% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Nathan aggressively rejects the standard industry presentation of test-time compute, labeling the popular inference-time scaling curve a deceptive marketing construct.
Hardest push from the hosts ▶ 41:40 Swyx rejects native plan tokensSwyx directly challenges Nathan's reasoning taxonomy by arguing that modern software engineering workflows favor external tool orchestration over internal model plan tokens.
Biggest teaching moment ▶ 16:47 Why RLHF outlasts RLVR as a research disciplineNathan educates the hosts on why preference tuning and RLHF will remain foundational research challenges indefinitely while verifiable reward RLVR may quickly saturate.
The host holds their own ▶ 52:10 Alessio diagnoses RL code pathologiesAlessio demonstrates deep practitioner expertise by detailing how RL-trained models introduce silent failure patterns through defensive if-statements in real-world codebases.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Welcome and Catching Up with Nathan Lambert | 4 | 3 | 1 | 1 | Friendly intro where Swyx and Alessio establish Nathan's background before Nathan recounts the development of Tulu 3 and the coining of RLVR. The hosts prompt him smoothly with conversational context. | |
| Expanding RLVR into Multi-Hop Tool Environments | 5 | 4 | 1 | 1 | Nathan breaks down the transition from single-generation RLVR to multi-hop environments like o3 and Deep Research. Alessio brings in insights from a previous conversation with Noam Brown regarding non-verifiable rewards. | |
| Agent Progress, Non-Verifiable Tasks, and Data Moats | 5 | 5 | 2 | 2 | Nathan explains why agent progress will look different from raw modeling progress and details the limitations of academic preference datasets compared to frontier lab data moats. Alessio questions whether frontier labs retain a structural edge via live inference inspection. | |
| Evaluating the Longevity of Chatbot Arena Benchmarks | 6 | 3 | 2 | 3 | Swyx presses on whether Chatbot Arena is 'cooked' and discusses vulnerabilities to gaming and the lack of multi-turn evals, citing recent critiques like Sarah Hooker's. Nathan defends the ongoing utility of Elo leaderboards for compression benchmarking. | |
| Writing the RLHF Book Amid the Reasoning Revolution | 5 | 5 | 2 | 2 | Nathan articulates why he is keeping his upcoming book focused on RLHF rather than RLVR, arguing RLHF deals with enduring alignment and human preference problems while RLVR may be solved quickly. Swyx probes his dismissal of GRPO algorithmic hype. | |
| Hybrid Reasoning Models versus Unified Interfaces | 6 | 5 | 2 | 2 | Alessio and Swyx challenge the necessity of hybrid reasoning models versus unified interfaces that self-allocate compute. Nathan outlines why search-augmented models degrade on raw SimpleQA when isolated from external search retrieval. | |
| Deconstructing Q*, o1, and Inference Scaling Plots | 6 | 4 | 3 | 2 | Nathan calls the standard inference-time scaling plot a 'psyop' because plotting test compute on a log axis misleads people into treating search as a simple linear knob. Alessio and Swyx discuss the ARC-AGI benchmark harness controversies. | |
| Strategic Research Opportunities for Academic AI | 5 | 4 | 1 | 1 | Nathan lays out realistic strategies for academic AI institutions like AI2, emphasizing domain-specific evals and artifact production over trying to match frontier labs on raw token burn. | |
| A Taxonomy of Reasoning: Skills, Strategy, Abstraction, Calibration | 6 | 4 | 2 | 3 | Nathan introduces his taxonomy of reasoning: skills, strategy, abstraction, and calibration. Swyx pushes back, questioning whether dedicated plan tokens are necessary when engineers prefer external tool orchestration. | |
| Task Decomposition, Planning Tokens, and Compute Allocation | 5 | 4 | 2 | 2 | Alessio questions whether plan templates could be reused as static blueprints rather than regenerated dynamically on every query. Nathan details how prompt generation remains more cost-effective despite potential abstraction failures. | |
| Parallel Test-Time Compute, Verifiers, and Diffusion LLMs | 6 | 4 | 2 | 3 | Swyx pushes back on Nathan's skepticism of parallel test-time compute, arguing it allows teams to pull forward future model capabilities to generate synthetic data for distillation. | |
| Code Generation Pathologies from Reinforcement Learning | 7 | 3 | 1 | 2 | Alessio brings concrete technical observations from using coding tools, noting that RL induces models to write defensive if-statements to suppress missing environment variables rather than raising clear errors. | |
| The Three Eras of Reinforcement Learning Over-Optimization | 6 | 5 | 1 | 1 | Nathan walks through the three historic eras of RL reward over-optimization across robotics control, RLHF, and RLVR. Swyx asks about the nuances of reward design and partial credit in multi-domain setups. | |
| Infrastructural Constraints and Long Inference Trajectories | 6 | 4 | 2 | 2 | Swyx and Nathan discuss training run limits and wall-clock feedback bottlenecks, referencing Swyx's earlier debate with Noam Brown. Nathan discusses Joanne Jang's Model Spec work at OpenAI. | |
| Micro-Distillation, Hardware Pricing, and AI Wearables | 6 | 4 | 2 | 3 | Alessio and Nathan discuss routing across micro-distilled models for specialized media tasks. Swyx pushes back, claiming multimodal models like GPT-4o will eventually absorb all point solutions, while noting the unexpected appreciation of RTX 4090 GPUs. | |
| Meta's Strategic Pivot and the Economics of AI Talent | 5 | 4 | 2 | 2 | Swyx and Nathan examine Meta's aggressive compensation packages for AI talent, analyzing the economic trade-off between massive cluster compute expenditures and elite researcher retention. | |
| AI2's Roadmap for an Open 'American DeepSeek' | 5 | 4 | 1 | 1 | Nathan outlines AI2's long-term technical roadmap for an open American equivalent to DeepSeek, moving OLMo from dense architectures to sparse MoE reasoning models. |