May 23, 2025 · 38m · latent-space
⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Latent Space podcast episode, Will Brown of Prime Intellect joins Alessio and Wix to analyze the Claude 4 release, examine the mechanics of extended thinking and reward hacking, and detail advanced techniques in multi-turn reinforcement learning and agentic evaluation.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Will bluntly counters Wix's suggestion of token-cost penalties in RL by pointing out model providers are commercial entities incentivized to sell tokens.
Hardest push from the hosts ▶ 10:28 Challenging thinking budgets versus reasoning effortWix pushes back against equating thinking budgets with reasoning effort, arguing users want target effort rather than arbitrary token cutoffs.
Biggest teaching moment ▶ 10:35 Demystifying reasoning effort mechanismsWill explains that reasoning effort under the hood is fundamentally a token budget trained via RL, leading Wix to explicitly revise his viewpoint.
The host holds their own ▶ 23:23 Finance analogy to eval company conflict of interestWix draws directly on his Morgan Stanley finance domain background to deliver an incisive analogy comparing AI eval businesses to credit rating agency conflicts.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Analyzing the Claude 4 Release and Shift Toward Practical Agents | 5 | 5 | 1 | 1 | Alessio and Wix discuss the Claude 4 announcement, highlighting extended thinking with tool use and the scratchpad paper lineage. Will expands on the broader industry shift from math competition reasoning toward practical multi-turn agents. | |
| Reward Hacking and Model Trustworthiness in Coding | 4 | 6 | 2 | 1 | Will breaks down how RL environments incentivize reward hacking in models like Sonnet 3.7, causing them to generate extraneous files and code. Wix agrees and provides anecdotal examples from his own coding workflow. | |
| Token Penalties, Thinking Budgets, and Reasoning Effort | 4 | 7 | 2 | 3 | Wix questions why token costs are not penalized in RL rewards, and Will explains how providers sell tokens and how thinking budgets are trained. Wix openly admits Will's explanation shifted his perspective on reasoning effort versus token budgets. | |
| The Claude Safety Stress-Testing Controversy | 3 | 6 | 2 | 1 | Wix brings up the Claude safety red-teaming controversy expecting banter, but Will provides a sober breakdown of multi-objective alignment dilemmas in extreme stress-testing scenarios. | |
| Action Spaces, MCP, and Multi-Agent RL Dynamics | 4 | 6 | 2 | 1 | Alessio asks about action spaces and tool permissions like MCP. Will draws an analogy between unbounded terminal action spaces in RL and complex multi-agent dynamical systems. | |
| Claude's Market Positioning and Brand Perception | 6 | 5 | 1 | 1 | Will and Wix evaluate evaluation benchmarks and LMSYS fundraising dynamics. Wix connects the incentive dilemma of eval companies selling to model labs with Wall Street credit rating agencies. | |
| Cultivating Research Taste and Betting on Multi-Agent Systems | 3 | 6 | 1 | 1 | Wix prompts Will on how to cultivate research taste. Will outlines how making calculated bets on multi-agent RL theory and RL with tool use guided his trajectory. | |
| Multi-Turn RL with GRPO and Verifiers Paper | 4 | 7 | 1 | 1 | Will details his paper on multi-turn RL with GRPO, explaining the phenomenon where small models execute dummy tool calls to game rewards and how intermediate verification solves credit assignment. | |
| Turn-Level Action Formulation and LLM-as-a-Judge Rewards | 4 | 7 | 1 | 2 | Will explains why turn-level action formulation and LLM-as-a-judge verifiers outperform brittle rule-based parsers, especially across LaTeX and symbolic mathematics outputs. |