May 9, 2025 · 17m · latent-space

⚡️Open Questions in Agentic RL — Will Brown (Prime Intellect)

Will Brown · 15m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Will Brown from Prime Intellect presents a comprehensive technical roadmap for the open-source community to train frontier agentic reinforcement learning models capable of long-horizon autonomous task execution. He outlines key engineering unlocks including multi-turn tool use, reasoning-based reward modeling, asynchronous training pipelines, and decentralized model merging to achieve open-weight 'N-minute AGI'.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.0 Guest teaching 0.0 Guest disagreement 1.3 The hosts pushing back 0.0
05100:0010:001:18–4:44 · The hosts as informed peer 0/10 Tool Use and the Concept of Ten-Minute AGI Solo lecture segment where Will Brown lays out the technical vision for multi-turn RL, 10-minute AGI, and the engineering bottlenecks of credit assignment and context blowup. No host participation occurs.4:45–7:42 · The hosts as informed peer 0/10 Scaling Curves in Tool Calls and Search Agents Brown examines scaling curves for search agents and critiques naive end-to-end multimodal CoT and basic multi-agent group chats, arguing programmatic tool manipulation is far more effective. Monologue delivery without host pushback.7:43–10:24 · The hosts as informed peer 0/10 Sub-Agents and Modular Context Optimization Brown details using sub-agents as tool calls to manage context windows and explains how discarding thinking tokens breaks standard RL math. The monologue format features zero host interaction.10:26–12:52 · The hosts as informed peer 0/10 Intermediary Step Verification and Reasoning Reward Models Presentation on turn-level reward models and treating verifiers as reasoning models to enable inference-time scaling. Pure technical exposition with no host present.12:54–15:40 · The hosts as informed peer 0/10 Asynchronous Reinforcement Learning and Off-Policy Stability Brown covers asynchronous RL, off-policy stability, and decentralized capability merging across orthogonal domains. No host engagement occurs during the talk.15:41–17:00 · The hosts as informed peer 0/10 Synthesis and Roadmap for Decentralized N-Minute AGI Brown concludes his synthesis on open-source agentic RL roadmap before the event MC steps in to manage stage transitions.1:18–4:44 · Guest teaching 0/10 Tool Use and the Concept of Ten-Minute AGI Solo lecture segment where Will Brown lays out the technical vision for multi-turn RL, 10-minute AGI, and the engineering bottlenecks of credit assignment and context blowup. No host participation occurs.4:45–7:42 · Guest teaching 0/10 Scaling Curves in Tool Calls and Search Agents Brown examines scaling curves for search agents and critiques naive end-to-end multimodal CoT and basic multi-agent group chats, arguing programmatic tool manipulation is far more effective. Monologue delivery without host pushback.7:43–10:24 · Guest teaching 0/10 Sub-Agents and Modular Context Optimization Brown details using sub-agents as tool calls to manage context windows and explains how discarding thinking tokens breaks standard RL math. The monologue format features zero host interaction.10:26–12:52 · Guest teaching 0/10 Intermediary Step Verification and Reasoning Reward Models Presentation on turn-level reward models and treating verifiers as reasoning models to enable inference-time scaling. Pure technical exposition with no host present.12:54–15:40 · Guest teaching 0/10 Asynchronous Reinforcement Learning and Off-Policy Stability Brown covers asynchronous RL, off-policy stability, and decentralized capability merging across orthogonal domains. No host engagement occurs during the talk.15:41–17:00 · Guest teaching 0/10 Synthesis and Roadmap for Decentralized N-Minute AGI Brown concludes his synthesis on open-source agentic RL roadmap before the event MC steps in to manage stage transitions.1:18–4:44 · Guest disagreement 1/10 Tool Use and the Concept of Ten-Minute AGI Solo lecture segment where Will Brown lays out the technical vision for multi-turn RL, 10-minute AGI, and the engineering bottlenecks of credit assignment and context blowup. No host participation occurs.4:45–7:42 · Guest disagreement 3/10 Scaling Curves in Tool Calls and Search Agents Brown examines scaling curves for search agents and critiques naive end-to-end multimodal CoT and basic multi-agent group chats, arguing programmatic tool manipulation is far more effective. Monologue delivery without host pushback.7:43–10:24 · Guest disagreement 2/10 Sub-Agents and Modular Context Optimization Brown details using sub-agents as tool calls to manage context windows and explains how discarding thinking tokens breaks standard RL math. The monologue format features zero host interaction.10:26–12:52 · Guest disagreement 1/10 Intermediary Step Verification and Reasoning Reward Models Presentation on turn-level reward models and treating verifiers as reasoning models to enable inference-time scaling. Pure technical exposition with no host present.12:54–15:40 · Guest disagreement 1/10 Asynchronous Reinforcement Learning and Off-Policy Stability Brown covers asynchronous RL, off-policy stability, and decentralized capability merging across orthogonal domains. No host engagement occurs during the talk.15:41–17:00 · Guest disagreement 0/10 Synthesis and Roadmap for Decentralized N-Minute AGI Brown concludes his synthesis on open-source agentic RL roadmap before the event MC steps in to manage stage transitions.1:18–4:44 · The hosts pushing back 0/10 Tool Use and the Concept of Ten-Minute AGI Solo lecture segment where Will Brown lays out the technical vision for multi-turn RL, 10-minute AGI, and the engineering bottlenecks of credit assignment and context blowup. No host participation occurs.4:45–7:42 · The hosts pushing back 0/10 Scaling Curves in Tool Calls and Search Agents Brown examines scaling curves for search agents and critiques naive end-to-end multimodal CoT and basic multi-agent group chats, arguing programmatic tool manipulation is far more effective. Monologue delivery without host pushback.7:43–10:24 · The hosts pushing back 0/10 Sub-Agents and Modular Context Optimization Brown details using sub-agents as tool calls to manage context windows and explains how discarding thinking tokens breaks standard RL math. The monologue format features zero host interaction.10:26–12:52 · The hosts pushing back 0/10 Intermediary Step Verification and Reasoning Reward Models Presentation on turn-level reward models and treating verifiers as reasoning models to enable inference-time scaling. Pure technical exposition with no host present.12:54–15:40 · The hosts pushing back 0/10 Asynchronous Reinforcement Learning and Off-Policy Stability Brown covers asynchronous RL, off-policy stability, and decentralized capability merging across orthogonal domains. No host engagement occurs during the talk.15:41–17:00 · The hosts pushing back 0/10 Synthesis and Roadmap for Decentralized N-Minute AGI Brown concludes his synthesis on open-source agentic RL roadmap before the event MC steps in to manage stage transitions.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 6:55 Debunking multi-agent group chat systems

Brown bluntly dismisses popular multi-agent frameworks like crewAI as 'silly' and unworkable because the underlying models and algorithms were never trained for cooperative multi-agent dynamics.

Hardest push from the hosts ▶ 16:39 Event host transition cut

In a monologue talk lacking direct host challenges, the sole moderation event is the event host pausing the stream for a speaker changeover.

Biggest teaching moment ▶ 7:30 Explaining sub-agent tool call context offloading

Brown methodically breaks down why running hundred-website scrapes directly in primary model context fails and proves why delegating to smaller, frozen helper models as tool calls is the mathematically sound alternative.

The host holds their own ▶ 16:39 Stage manager handoff

The presentation is a solo keynote with no host technical rebuttals, leaving the logistical handoff as the only host intervention.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Tool Use and the Concept of Ten-Minute AGI 0010 Solo lecture segment where Will Brown lays out the technical vision for multi-turn RL, 10-minute AGI, and the engineering bottlenecks of credit assignment and context blowup. No host participation occurs.
Scaling Curves in Tool Calls and Search Agents 0030 Brown examines scaling curves for search agents and critiques naive end-to-end multimodal CoT and basic multi-agent group chats, arguing programmatic tool manipulation is far more effective. Monologue delivery without host pushback.
Sub-Agents and Modular Context Optimization 0020 Brown details using sub-agents as tool calls to manage context windows and explains how discarding thinking tokens breaks standard RL math. The monologue format features zero host interaction.
Intermediary Step Verification and Reasoning Reward Models 0010 Presentation on turn-level reward models and treating verifiers as reasoning models to enable inference-time scaling. Pure technical exposition with no host present.
Asynchronous Reinforcement Learning and Off-Policy Stability 0010 Brown covers asynchronous RL, off-policy stability, and decentralized capability merging across orthogonal domains. No host engagement occurs during the talk.
Synthesis and Roadmap for Decentralized N-Minute AGI 0000 Brown concludes his synthesis on open-source agentic RL roadmap before the event MC steps in to manage stage transitions.

Statements from this episode (15)

Assertion Not checkable as stated
No open-source model currently matches OpenAI's general agent capabilities
“There really isn't currently an open source model that behaves in the way that these models do. We have things like R-one, which are great at kind of the single-term math and code reasoning problems. But they are not in the general purpose agents world yet.”
Will Brown May 9, 2025 ▶ 0:38
Opinion
OpenAI's o3 acts as a '10-minute AGI' for human tasks
“Whether or not you want to call this, like, end minute AGI, I kind of like the phrase, 10 minute AGI, for just, like, how to think about O-three is that anything that you can do as a human in 10 minutes, O-three is usually going to be able to do reasonably wel…”
Will Brown May 9, 2025 ▶ 1:48
Insight
Scaling autonomous task duration is a plausible path to AGI
“And that, if you can crack that scaling direction of, like, pushing the boundary of how long these models can go out and do these things for, that is a plausible path towards things that become marvelous.”
Will Brown May 9, 2025 ▶ 2:02
Insight
Long-horizon agent RL requires intermediate turn-level verification
“You probably want some kind of intermediate verification where you're not just waiting for the final answer at the end, But you want something like turn-level reward, potentially, where you want to be able to ensure that the model is getting credit for the mov…”
Will Brown May 9, 2025 ▶ 3:15
Insight
Agent RL training infrastructure must be asynchronous to avoid compute bubbles
“You also ideally want to, whether you're doing this centralized or decentralized, move in the direction of, like, everything being async and overlaps because that is, like, otherwise you have these inefficiency bubbles that, like, pop up all over your compute …”
Will Brown May 9, 2025 ▶ 4:16
Insight
More tool calls and web browsing yield a clear scaling curve
“Using more tool calls, searching the web more, gives you a nice scaling curve where you get better answers by putting in more effort, by, like, spending more time browsing the internet, essentially.”
Will Brown May 9, 2025 ▶ 5:35
Insight
Programmatic tools beat end-to-end image generation for multimodal reasoning
“Where I think for a while some people were, like, speculating, like, oh, what if you have the model, like, generate images in its chain of thought reasoning where everything is, like end-to-end multimodal input and output, and it seems like you don't really ne…”
Will Brown May 9, 2025 ▶ 6:01
Opinion
Group-chat multi-agent systems like CrewAI do not work well
“I think people were, like, excited about multi-agent systems for a while, like, the crew AI sort of thing of, like, oh, I'm gonna put my coder agent, my finance agent in group chat, and, like, a lot of these are just kind of silly. They don't actually work ver…”
Will Brown May 9, 2025 ▶ 7:43
Assertion Not checkable as stated
Algorithms for practical multi-agent RL do not yet exist
“Multi-agent RL's hard. I did five years of it in grad school. It's like not easy. And to the algorithms don't really even exist for the things you would really want to do.”
Will Brown May 9, 2025 ▶ 8:04
Assertion Contradicted
No legitimate open-source million-token context models exist at scale
“Scaling to, like, million token contexts is, like, really, really hard. There, I don't think there are real, like, open source replications, open token context scaling, Beyond, like, tiny, like, academic model sizes.”
Will Brown May 9, 2025 ▶ 9:50
Insight
Current multi-turn RL research ignores discarded thinking tokens, breaking the math
“The existing paper people are writing about multi-turn RL are not actually incorporating this, and it kind of, like, breaks all the math.”
Will Brown May 9, 2025 ▶ 10:48
Insight
Reasoning models acting as reward models are key to agent RL
“And the most, one of the most promising ways, I think, towards doing this is having the reward models also be able to answer harder questions by themselves being reasoning models.”
Will Brown May 9, 2025 ▶ 11:54
Insight
Averaging weights of models trained on separate domains works effectively
“You can have a model trained on code, and a model trained on math, and a model trained on Spanish, and you can literally average the weights, and it works.”
Will Brown May 9, 2025 ▶ 14:33
Insight
Weight updates across specialized tasks are orthogonal enough to merge asynchronously
“The updates made to model weights are orthogonal enough for specialized tasks that this is actually like totally fine. Things are nice and linear in most cases, things are nice and orthogonal, and you can get away with a lot of async updates to models that are…”
Will Brown May 9, 2025 ▶ 14:53
Assertion Supported
Cohere used extensive asynchronous model merging to build its flagship models
“Coherent to this very extensively in their latest flagship model, command A”
Will Brown May 9, 2025 ▶ 15:11
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.