Opinion certainty 3/5 debate potential 3/5

Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes

Nathan Lambert · The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai) · Jul 31, 2025 · at 7:35

AI2 researcher Nathan Lambert discusses how reasoning and research agents like OpenAI Deep Research are trained using RL on modular tasks rather than full-document evaluation.

0:00 / 0:28exact quote · 28.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that when you put these pieces together or a couple of different fine tunes of a model. So it seems like deep research has some fine tune of O three in it. As you do that with some different domains of RL, it works rather than deep research being trained on the outcome”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Nathan Lambert

Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Nathan Lambert Jul 31, 2025 ▶ 16:42 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Nathan Lambert Jul 31, 2025 ▶ 19:28 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Not checkable as stated
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Nathan Lambert Jul 31, 2025 ▶ 20:52 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Nathan Lambert Jul 31, 2025 ▶ 24:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Nathan Lambert Jul 31, 2025 ▶ 1:03:38 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.