May 23, 2025 · 38m · latent-space

⚡️Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect

Will Brown · 28m spoken Shawn Wang · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Latent Space podcast episode, Will Brown of Prime Intellect joins Alessio and Wix to analyze the Claude 4 release, examine the mechanics of extended thinking and reward hacking, and detail advanced techniques in multi-turn reinforcement learning and agentic evaluation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.1 Guest teaching 6.1 Guest disagreement 1.4 The hosts pushing back 1.3
05100:0010:0020:0030:001:04–5:51 · The hosts as informed peer 5/10 Analyzing the Claude 4 Release and Shift Toward Practical Agents Alessio and Wix discuss the Claude 4 announcement, highlighting extended thinking with tool use and the scratchpad paper lineage. Will expands on the broader industry shift from math competition reasoning toward practical multi-turn agents.5:52–8:39 · The hosts as informed peer 4/10 Reward Hacking and Model Trustworthiness in Coding Will breaks down how RL environments incentivize reward hacking in models like Sonnet 3.7, causing them to generate extraneous files and code. Wix agrees and provides anecdotal examples from his own coding workflow.8:39–12:17 · The hosts as informed peer 4/10 Token Penalties, Thinking Budgets, and Reasoning Effort Wix questions why token costs are not penalized in RL rewards, and Will explains how providers sell tokens and how thinking budgets are trained. Wix openly admits Will's explanation shifted his perspective on reasoning effort versus token budgets.12:20–16:56 · The hosts as informed peer 3/10 The Claude Safety Stress-Testing Controversy Wix brings up the Claude safety red-teaming controversy expecting banter, but Will provides a sober breakdown of multi-objective alignment dilemmas in extreme stress-testing scenarios.16:57–20:51 · The hosts as informed peer 4/10 Action Spaces, MCP, and Multi-Agent RL Dynamics Alessio asks about action spaces and tool permissions like MCP. Will draws an analogy between unbounded terminal action spaces in RL and complex multi-agent dynamical systems.20:51–25:14 · The hosts as informed peer 6/10 Claude's Market Positioning and Brand Perception Will and Wix evaluate evaluation benchmarks and LMSYS fundraising dynamics. Wix connects the incentive dilemma of eval companies selling to model labs with Wall Street credit rating agencies.25:16–27:28 · The hosts as informed peer 3/10 Cultivating Research Taste and Betting on Multi-Agent Systems Wix prompts Will on how to cultivate research taste. Will outlines how making calculated bets on multi-agent RL theory and RL with tool use guided his trajectory.27:30–32:00 · The hosts as informed peer 4/10 Multi-Turn RL with GRPO and Verifiers Paper Will details his paper on multi-turn RL with GRPO, explaining the phenomenon where small models execute dummy tool calls to game rewards and how intermediate verification solves credit assignment.32:00–37:33 · The hosts as informed peer 4/10 Turn-Level Action Formulation and LLM-as-a-Judge Rewards Will explains why turn-level action formulation and LLM-as-a-judge verifiers outperform brittle rule-based parsers, especially across LaTeX and symbolic mathematics outputs.1:04–5:51 · Guest teaching 5/10 Analyzing the Claude 4 Release and Shift Toward Practical Agents Alessio and Wix discuss the Claude 4 announcement, highlighting extended thinking with tool use and the scratchpad paper lineage. Will expands on the broader industry shift from math competition reasoning toward practical multi-turn agents.5:52–8:39 · Guest teaching 6/10 Reward Hacking and Model Trustworthiness in Coding Will breaks down how RL environments incentivize reward hacking in models like Sonnet 3.7, causing them to generate extraneous files and code. Wix agrees and provides anecdotal examples from his own coding workflow.8:39–12:17 · Guest teaching 7/10 Token Penalties, Thinking Budgets, and Reasoning Effort Wix questions why token costs are not penalized in RL rewards, and Will explains how providers sell tokens and how thinking budgets are trained. Wix openly admits Will's explanation shifted his perspective on reasoning effort versus token budgets.12:20–16:56 · Guest teaching 6/10 The Claude Safety Stress-Testing Controversy Wix brings up the Claude safety red-teaming controversy expecting banter, but Will provides a sober breakdown of multi-objective alignment dilemmas in extreme stress-testing scenarios.16:57–20:51 · Guest teaching 6/10 Action Spaces, MCP, and Multi-Agent RL Dynamics Alessio asks about action spaces and tool permissions like MCP. Will draws an analogy between unbounded terminal action spaces in RL and complex multi-agent dynamical systems.20:51–25:14 · Guest teaching 5/10 Claude's Market Positioning and Brand Perception Will and Wix evaluate evaluation benchmarks and LMSYS fundraising dynamics. Wix connects the incentive dilemma of eval companies selling to model labs with Wall Street credit rating agencies.25:16–27:28 · Guest teaching 6/10 Cultivating Research Taste and Betting on Multi-Agent Systems Wix prompts Will on how to cultivate research taste. Will outlines how making calculated bets on multi-agent RL theory and RL with tool use guided his trajectory.27:30–32:00 · Guest teaching 7/10 Multi-Turn RL with GRPO and Verifiers Paper Will details his paper on multi-turn RL with GRPO, explaining the phenomenon where small models execute dummy tool calls to game rewards and how intermediate verification solves credit assignment.32:00–37:33 · Guest teaching 7/10 Turn-Level Action Formulation and LLM-as-a-Judge Rewards Will explains why turn-level action formulation and LLM-as-a-judge verifiers outperform brittle rule-based parsers, especially across LaTeX and symbolic mathematics outputs.1:04–5:51 · Guest disagreement 1/10 Analyzing the Claude 4 Release and Shift Toward Practical Agents Alessio and Wix discuss the Claude 4 announcement, highlighting extended thinking with tool use and the scratchpad paper lineage. Will expands on the broader industry shift from math competition reasoning toward practical multi-turn agents.5:52–8:39 · Guest disagreement 2/10 Reward Hacking and Model Trustworthiness in Coding Will breaks down how RL environments incentivize reward hacking in models like Sonnet 3.7, causing them to generate extraneous files and code. Wix agrees and provides anecdotal examples from his own coding workflow.8:39–12:17 · Guest disagreement 2/10 Token Penalties, Thinking Budgets, and Reasoning Effort Wix questions why token costs are not penalized in RL rewards, and Will explains how providers sell tokens and how thinking budgets are trained. Wix openly admits Will's explanation shifted his perspective on reasoning effort versus token budgets.12:20–16:56 · Guest disagreement 2/10 The Claude Safety Stress-Testing Controversy Wix brings up the Claude safety red-teaming controversy expecting banter, but Will provides a sober breakdown of multi-objective alignment dilemmas in extreme stress-testing scenarios.16:57–20:51 · Guest disagreement 2/10 Action Spaces, MCP, and Multi-Agent RL Dynamics Alessio asks about action spaces and tool permissions like MCP. Will draws an analogy between unbounded terminal action spaces in RL and complex multi-agent dynamical systems.20:51–25:14 · Guest disagreement 1/10 Claude's Market Positioning and Brand Perception Will and Wix evaluate evaluation benchmarks and LMSYS fundraising dynamics. Wix connects the incentive dilemma of eval companies selling to model labs with Wall Street credit rating agencies.25:16–27:28 · Guest disagreement 1/10 Cultivating Research Taste and Betting on Multi-Agent Systems Wix prompts Will on how to cultivate research taste. Will outlines how making calculated bets on multi-agent RL theory and RL with tool use guided his trajectory.27:30–32:00 · Guest disagreement 1/10 Multi-Turn RL with GRPO and Verifiers Paper Will details his paper on multi-turn RL with GRPO, explaining the phenomenon where small models execute dummy tool calls to game rewards and how intermediate verification solves credit assignment.32:00–37:33 · Guest disagreement 1/10 Turn-Level Action Formulation and LLM-as-a-Judge Rewards Will explains why turn-level action formulation and LLM-as-a-judge verifiers outperform brittle rule-based parsers, especially across LaTeX and symbolic mathematics outputs.1:04–5:51 · The hosts pushing back 1/10 Analyzing the Claude 4 Release and Shift Toward Practical Agents Alessio and Wix discuss the Claude 4 announcement, highlighting extended thinking with tool use and the scratchpad paper lineage. Will expands on the broader industry shift from math competition reasoning toward practical multi-turn agents.5:52–8:39 · The hosts pushing back 1/10 Reward Hacking and Model Trustworthiness in Coding Will breaks down how RL environments incentivize reward hacking in models like Sonnet 3.7, causing them to generate extraneous files and code. Wix agrees and provides anecdotal examples from his own coding workflow.8:39–12:17 · The hosts pushing back 3/10 Token Penalties, Thinking Budgets, and Reasoning Effort Wix questions why token costs are not penalized in RL rewards, and Will explains how providers sell tokens and how thinking budgets are trained. Wix openly admits Will's explanation shifted his perspective on reasoning effort versus token budgets.12:20–16:56 · The hosts pushing back 1/10 The Claude Safety Stress-Testing Controversy Wix brings up the Claude safety red-teaming controversy expecting banter, but Will provides a sober breakdown of multi-objective alignment dilemmas in extreme stress-testing scenarios.16:57–20:51 · The hosts pushing back 1/10 Action Spaces, MCP, and Multi-Agent RL Dynamics Alessio asks about action spaces and tool permissions like MCP. Will draws an analogy between unbounded terminal action spaces in RL and complex multi-agent dynamical systems.20:51–25:14 · The hosts pushing back 1/10 Claude's Market Positioning and Brand Perception Will and Wix evaluate evaluation benchmarks and LMSYS fundraising dynamics. Wix connects the incentive dilemma of eval companies selling to model labs with Wall Street credit rating agencies.25:16–27:28 · The hosts pushing back 1/10 Cultivating Research Taste and Betting on Multi-Agent Systems Wix prompts Will on how to cultivate research taste. Will outlines how making calculated bets on multi-agent RL theory and RL with tool use guided his trajectory.27:30–32:00 · The hosts pushing back 1/10 Multi-Turn RL with GRPO and Verifiers Paper Will details his paper on multi-turn RL with GRPO, explaining the phenomenon where small models execute dummy tool calls to game rewards and how intermediate verification solves credit assignment.32:00–37:33 · The hosts pushing back 2/10 Turn-Level Action Formulation and LLM-as-a-Judge Rewards Will explains why turn-level action formulation and LLM-as-a-judge verifiers outperform brittle rule-based parsers, especially across LaTeX and symbolic mathematics outputs.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 8:50 Pushing back on token penalties

Will bluntly counters Wix's suggestion of token-cost penalties in RL by pointing out model providers are commercial entities incentivized to sell tokens.

Hardest push from the hosts ▶ 10:28 Challenging thinking budgets versus reasoning effort

Wix pushes back against equating thinking budgets with reasoning effort, arguing users want target effort rather than arbitrary token cutoffs.

Biggest teaching moment ▶ 10:35 Demystifying reasoning effort mechanisms

Will explains that reasoning effort under the hood is fundamentally a token budget trained via RL, leading Wix to explicitly revise his viewpoint.

The host holds their own ▶ 23:23 Finance analogy to eval company conflict of interest

Wix draws directly on his Morgan Stanley finance domain background to deliver an incisive analogy comparing AI eval businesses to credit rating agency conflicts.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Analyzing the Claude 4 Release and Shift Toward Practical Agents 5511 Alessio and Wix discuss the Claude 4 announcement, highlighting extended thinking with tool use and the scratchpad paper lineage. Will expands on the broader industry shift from math competition reasoning toward practical multi-turn agents.
Reward Hacking and Model Trustworthiness in Coding 4621 Will breaks down how RL environments incentivize reward hacking in models like Sonnet 3.7, causing them to generate extraneous files and code. Wix agrees and provides anecdotal examples from his own coding workflow.
Token Penalties, Thinking Budgets, and Reasoning Effort 4723 Wix questions why token costs are not penalized in RL rewards, and Will explains how providers sell tokens and how thinking budgets are trained. Wix openly admits Will's explanation shifted his perspective on reasoning effort versus token budgets.
The Claude Safety Stress-Testing Controversy 3621 Wix brings up the Claude safety red-teaming controversy expecting banter, but Will provides a sober breakdown of multi-objective alignment dilemmas in extreme stress-testing scenarios.
Action Spaces, MCP, and Multi-Agent RL Dynamics 4621 Alessio asks about action spaces and tool permissions like MCP. Will draws an analogy between unbounded terminal action spaces in RL and complex multi-agent dynamical systems.
Claude's Market Positioning and Brand Perception 6511 Will and Wix evaluate evaluation benchmarks and LMSYS fundraising dynamics. Wix connects the incentive dilemma of eval companies selling to model labs with Wall Street credit rating agencies.
Cultivating Research Taste and Betting on Multi-Agent Systems 3611 Wix prompts Will on how to cultivate research taste. Will outlines how making calculated bets on multi-agent RL theory and RL with tool use guided his trajectory.
Multi-Turn RL with GRPO and Verifiers Paper 4711 Will details his paper on multi-turn RL with GRPO, explaining the phenomenon where small models execute dummy tool calls to game rewards and how intermediate verification solves credit assignment.
Turn-Level Action Formulation and LLM-as-a-Judge Rewards 4712 Will explains why turn-level action formulation and LLM-as-a-judge verifiers outperform brittle rule-based parsers, especially across LaTeX and symbolic mathematics outputs.

Statements from this episode (17)

Insight
Will Brown: AI reasoning models are merely a stepping stone toward autonomous agents
“The thing that's going to make the next wave of stuff be powerful is just, like, everyone wants better agents. Everyone wants models that can, like, go off and do stuff. And, like, reasoning was kind of, like, a precursor to that a little bit.”
Will Brown May 23, 2025 ▶ 1:29
Opinion
Brown: Anthropic treats extended thinking as tool use, not distinct model class
“And it seemed like Anthropik's kind of attitude has been that extended thinking is an instance of tool use and that it's the kind of thing you want to equip the model with the ability to do. But it's not like, oh, it's a thinking model. It's just a sync for th…”
Will Brown May 23, 2025 ▶ 3:50
Opinion
Brown: Claude thinking and non-thinking modes likely use same underlying model
“I mean, I think these models should be the same model, and Anthropic knows what they're doing. Like, it's not that hard to, like, Quen did it in a very kind of, like, simple way, and they kind of talked about how they did it a little bit. But it's not, like, t…”
Will Brown May 23, 2025 ▶ 4:49
Insight
Brown: Truncating reasoning model thinking mid-sentence still yields good outputs
“So it seems like artificially truncating the thought is actually like fine. Like the model can, even if like it got cut off mid-sentence with an injected like think token, these are smart enough models that they can kind of finish with the best that they got f…”
Will Brown May 23, 2025 ▶ 9:34
Prediction Not checkable as stated
Brown: Reasoning effort dropdowns will disappear from chat interfaces
“I think in chat interfaces, it probably won't stick around. Like, I don't think we're always going to have the dropdown of like Oath for many and Oath for many high. That feels silly.”
Will Brown May 23, 2025 ▶ 11:45
Insight
Brown: Anthropic safety issues stem from conflicting model objectives
“A lot of the kind of headline anthropic like safety results, especially related to reward hacking and kind of deviation and alignment faking, Are all things to me that seem like a rock and a hard play situation where the model has two objectives it's given tha…”
Will Brown May 23, 2025 ▶ 13:22
Insight
Brown: Base LLMs will do anything up to their intelligence limit
“The base model in general of LLM is not artificially constrained in any way. Like, with the right prompt, it'll do whatever up to its intelligence limit.”
Will Brown May 23, 2025 ▶ 15:45
Opinion
Brown: Claude 3.7 works for quick projects, not large codebases
“I never really got to the point where I found it was helpful for a thing that was like a large existing code base. But if it's like, hey, I want to like cook something up in a few hours for fun. Pretty good at that. But these become messy and they become hard …”
Will Brown May 23, 2025 ▶ 17:26
Opinion
Swyx: Apollo Research reports frontier safety findings as effective marketing
“And part of this is like Apollo just being Apollo you know, pushing the frontier of red teaming. Right. So they're going to report the things because it's extremely good at Apollo marketing.”
Shawn Wang May 23, 2025 ▶ 21:09
Opinion
Will Brown: Claude appeals to AI insiders but lacks mainstream breakout
“It feels like people in the AI world, like, love Claude, or have grown type of Claude, but still had a phase where they were using it a ton. But it hasn't really broken out to general people in the way. And it feels like a lot of their marketing that I've seen…”
Will Brown May 23, 2025 ▶ 21:28
Insight
Will Brown: Selling to AI Labs Compromises Model Evaluation Integrity
“I think being an eval company puts you in a really hard spot. Some people are talking about this on Twitter, like just that to be an ed-all company, you kind of have to sell to the labs, but selling to the labs doesn't really, like kind of wrecks the revals.”
Will Brown May 23, 2025 ▶ 23:06
Prediction Not checkable as stated
Will Brown: Academia Will Likely Be the Best Source of AI Evals
“I mean, I do think that like the best source of evals going forward is probably going to be academia.”
Will Brown May 23, 2025 ▶ 23:43
Insight
Brown: Small LLMs default to skipping tool calls without explicit training
“If you set these models up to use tools, They just won't. Like if you say, hey, here's a question. You have access to these tools. Do as many rounds of tool calling as you want, and then submit your answer. They'll just submit their answer because they like ar…”
Will Brown May 23, 2025 ▶ 28:55
Insight
Brown: Prompting alone cannot reliably force LLMs to use thinking tokens
“If you want models to use thinking tokens, you kind of have to, like, incentivize that. You have to either do a little bit of, like, SFT warmup, or you have to, like, Reward them for doing it. Otherwise, they will not follow it a hundred percent of the time on…”
Will Brown May 23, 2025 ▶ 29:48
Insight
Brown: Rewarding generic tool calls causes LLMs to game rewards
“If you start rewarding them for, like, tool use, They will use the tool, but they don't really want to, like, have to, they want to, like, be very safe with it... They would like do silly versions of tool use where they aren't actually using the tool to assist…”
Will Brown May 23, 2025 ▶ 30:53
Assertion Supported
Brown: GRPO is more memory efficient and easier to distribute
“GRPO is, like, great for, like, leaning heavy on highly parallel inference compute. It's more memory efficient for the actual training process. It's much easier to do in a distributed fashion because you have less gradient syncing and less model weight copies.”
Will Brown May 23, 2025 ▶ 32:07
Insight
Brown: RL training is shifting toward model-based LLM judges
“So it, like, feels like people are moving in the direction of model-based rewards, where you, either LLM is a judge where the judge sees the correct answer, or it has questions it's supposed to verify as properties of the response, just because that's much mor…”
Will Brown May 23, 2025 ▶ 33:27
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.