Nathan Lambert

Founder, Interconnects AI · 3 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistauthorfounderhostengineer@natolambert ↗LinkedIn ↗interconnects.ai ↗

Nathan Lambert is an AI researcher known for his work in post-training and reinforcement learning from human feedback (RLHF), having played a central role in open-source language model projects including Ai2’s OLMo and Tülu. He writes and hosts the technical publication Interconnects and authored the textbook Reinforcement Learning from Human Feedback.

109statements → 50claims → 23claims resolved → 87%fully supported → 3.56/5average certainty → 2.23/5average debate potential → ≈4.0/5argument clarity, estimated → 6said about them ↓

20 supported 2 partly supported 1 contradicted 4 not yet assessed 21 not checkable as stated how the 50 claims stand · each chip opens the sources

18 predictions · 32 assertions · 21 opinions · 35 insights · 3 disclosures · every statement was checked. The predictions and assertions are the 50 claims: statements the public record can support or contradict. 25 are resolved, 4 are not yet assessed, and 21 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Nathan argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

Their most notable contradicted claim

Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Nathan Lambert Jan 11, 2024 ▶ 1:26:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
88% certainty 3
93% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Nathan Lambert on measured tape to publish a rate. This says nothing about how they speak.

Everything Nathan Lambert said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Opinion
Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that whe…”
Nathan Lambert Jul 31, 2025 ▶ 7:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Nathan Lambert Jul 31, 2025 ▶ 16:42 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Nathan Lambert Jul 31, 2025 ▶ 19:28 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Not checkable as stated
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Nathan Lambert Jul 31, 2025 ▶ 20:52 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Nathan Lambert Jul 31, 2025 ▶ 24:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Nathan Lambert Jul 31, 2025 ▶ 1:03:38 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Top AI talent is dramatically cheaper than GPU clusters
“Talent is cheaper than GPUs by a dramatic margin, and At the end of the day, it's like, okay, if we're spending this much, they go to the room and they stare in the mirror and you're like, wait, it might not actually be that ridiculous to spend this money on t…”
Nathan Lambert Jul 31, 2025 ▶ 1:13:46 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Nathan Lambert Jul 31, 2025 ▶ 1:16:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Nathan Lambert Jan 2, 2025 ▶ 5:23 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Prediction Not checkable as stated
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Nathan Lambert Oct 13, 2024 ▶ 12:15 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Insight
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Nathan Lambert Jan 11, 2024 ▶ 29:37 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Nathan Lambert Jan 11, 2024 ▶ 45:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Opinion
Lambert: Chatbot Arena is the best available evaluation benchmark for LLMs
“I have, if we make it to evaluation, I'd pretty much say that Chat Arena is the best limited evaluation that people have to learn how to use language models, and like, It's very valuable data”
Nathan Lambert Jan 11, 2024 ▶ 52:38 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Nathan Lambert Jan 11, 2024 ▶ 58:36 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: RLVR is broader than ground truth because code is verifiable
“The verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.”
Nathan Lambert Jul 31, 2025 ▶ 5:13 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Inference scaling plots misleadingly suggest search is an easy control knob
“The core of that article is just, they're taking points from within training, or there's a natural variance, and then you line them up. And if you line them up, then you get this nice inference time-scaling behavior, which is, and now people, a lot of people h…”
Nathan Lambert Jul 31, 2025 ▶ 28:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Opinion
Lambert: AI benchmarks like ARC-AGI should prioritize testing without harnesses
“Harnesses are cool, but they're gonna, they're, They're a handicap that's changing the learning dynamics substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses.”
Nathan Lambert Jul 31, 2025 ▶ 34:02 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: AI academics must build datasets and evals rather than papers
“If you're trying to have impact in AI right now, it's as an academic, you have to like level up out of papers to artifacts, which is models, datasets, evals. Datasets and evals are easier for people to have impact on.”
Nathan Lambert Jul 31, 2025 ▶ 36:13 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Opinion
Lambert: Labs trade code usability for massive RL performance gains
“That's just like the labs are trading off massive gains in performance or small detriments in usability. And it's like, do you ship that model? Yeah. Like you just ship it and deal with it later, but I'm sure they could, I'm sure that's a fixable thing.”
Nathan Lambert Jul 31, 2025 ▶ 52:44 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Nathan Lambert Jul 31, 2025 ▶ 59:33 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Opinion
Lambert: The local model community is much smaller than assumed
“Like the local modeling community, I think is much smaller than people give it credit for, because most of the use for open models is still in APIs.”
Nathan Lambert Jul 31, 2025 ▶ 1:09:26 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Opinion
Lambert: Meta withholding its leading benchmark model is bad execution
“But to be a model that claims to be open and then not release the model that is your leading claim is just, like, that is, like, bad execution.”
Nathan Lambert Jul 31, 2025 ▶ 1:13:36 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

Show 24statements(85 left)

The other half of the tape: Nathan Lambert's own voice is left out of every number here. Other people bring the name up 6 times in 3 episodes on Latent Space. every mention, with the transcript →

Who brings them up most Alessio Fanelli 3Luca Soldani 2Shawn Wang 1

Every mention by year

tap a year for its mentions
00214220242025episodesmentions
01220242025episodes it came up in
00112220242025episodesmentions per episode

Appearances (3)

EpisodeDateSpeaking time
The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai) Jul 31, 2025 54m
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz Oct 13, 2024 4m
The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert Jan 11, 2024 1h 13m

Played on the show (1)

Episodes where a recording of Nathan Lambert was played rather than Nathan taking part, or where the tape carries an address with nobody putting questions to them. Listed because the words are on the record, kept out of every score on this page because they were not said on this show. We read this off the tape: who was spoken to, who was asked something, who answered whom.

EpisodeDateOn tapeWhat it is
The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024] Jan 2, 2025 14m aired address
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.