People, every show

Nathan Lambert

Founder, Interconnects AI. On 2 shows, 4 appearances, plus 1 compilation re-air not counted. The Shows tab opens the full record on each.

scientistauthorfounderhostengineer@natolambert ↗LinkedIn ↗interconnects.ai ↗

Nathan Lambert is an AI researcher known for his work in post-training and reinforcement learning from human feedback (RLHF), having played a central role in open-source language model projects including Ai2’s OLMo and Tülu. He writes and hosts the technical publication Interconnects and authored the textbook Reinforcement Learning from Human Feedback.

2shows
4appearances
132statements
28resolved
24supported
1contradicted
86%fully supported
6said about them ↓

Everything Nathan Lambert said on any show that made the record, most notable first. Each card names its show and opens the statement there.

LATENT SPACE Prediction Not checkable as stated
Lambert: Open source will learn to train models on arbitrary preference data
“I really think people in open source and academics are going to figure out how to use any preference data on any model just because they're scrappy.”
Nathan Lambert Jan 11, 2024 ▶ 47:48 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Nathan Lambert Jan 11, 2024 ▶ 48:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Prediction Not checkable as stated
Lambert: OpenAI Will Not Aggressively Ban Synthetic Training Scraping
“I don't expect OpenAI to go too crazy on this, because they're just gonna, there's gonna be so much backlash against them.”
Nathan Lambert Jan 11, 2024 ▶ 50:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”
Nathan Lambert Jan 11, 2024 ▶ 54:59 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Not checkable as stated
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Nathan Lambert Jan 11, 2024 ▶ 1:19:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Disclosure
AI2 Plans to Release Fully Open Pre-Trained LLMs With Data and Code
“The Allen Institute is training, pre-training language models, or pre-training, like, open language models, where we'll be able to share, like, data, code, everything, the kind of horn that everyone likes to get annoyed about these days, it's like, well, I'm n…”
Nathan Lambert Jan 11, 2024 ▶ 1:20:08 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Nathan Lambert Jan 11, 2024 ▶ 1:26:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Frontier Labs Lack Visibility into Cross-Model RLHF Sensitivity
“I think big labs are so over-indexed, are indexed on their own base models, so they don't know, like, what's swapping between CloudBase or GPT-IV-Base, how that would change any notion of preference or what you do with RLHF.”
Nathan Lambert Jan 11, 2024 ▶ 1:30:15 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Lambert: Scale AI has historically struggled to retain technical ML talent
“I think they've historically had trouble keeping, like, technical ML talent, but they've started a new research lab, so that should help.”
Nathan Lambert Jan 11, 2024 ▶ 1:34:33 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
MAD Assertion Supported
Lambert: OLMo 3 models are the best open models outside Qwen 3
“I would say in post training where The best models that don't start with Quinn three and we're like reasonable to say that they are comparable to Quinn three, like on some benchmarks would beat them on some benchmarks. They're way ahead.”
Nathan Lambert Nov 20, 2025 ▶ 8:59 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Supported
Lambert: Alibaba's Qwen 3 VL vision model is a superior text model
“They released these Quinn three VL, their vision models. And like on text only benchmarks, it's way better than the models they released in April. So it's like okay, like that's the new baseline. And most people don't know about it because they think it's just…”
Nathan Lambert Nov 20, 2025 ▶ 9:46 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Prediction Not checkable as stated
Lambert predicts more US labs will release open AI models
“If you look at this podcast in the coming months, I do think there's going to be, look like there's a lot more labs in the U S participating.”
Nathan Lambert Nov 20, 2025 ▶ 16:22 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Disclosure
Lambert: AI2 coined 'reinforcement learning with verifiable rewards' replicating Llama 3
“We spent a long time to try to replicate what we thought was close to Lama three post training with multiple stages and optimizers, which is the project that like came up with the name reinforcement learning with verifiable rewards with a bunch of people.”
Nathan Lambert Nov 20, 2025 ▶ 29:22 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Supported
Lambert: Hugging Face outcompeted AI2's AllenNLP library
“It was the main competitor to Hugging Face Transformers. And they ultimately outcompeted AI two as the thing that people use for that because they had very different model and amount of support.”
Nathan Lambert Nov 20, 2025 ▶ 32:47 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Long-context extension is essential for reasoning AI models
“Three is long context extension, which is absolutely essential for these reasoning models because they generate so many intermediate tokens before sharing an answer with you.”
Nathan Lambert Nov 20, 2025 ▶ 40:15 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD What-if
Lambert: Scaling AI 10x alters post-training, not pre-training methods
“If like, if we were to train a model that was 10 times as big, like all this post-training stuff would change. But the pre-training And mid training and long contacts, I think would actually become looking pretty similar.”
Nathan Lambert Nov 20, 2025 ▶ 40:55 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Larger pre-trained base models are easier to improve with RL
“A better base model and a bigger base model is much easier to improve with RL.”
Nathan Lambert Nov 20, 2025 ▶ 44:29 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Supported
Lambert: Kernel differences between vLLM and Hugging Face cause RL numerical instability
“VLLM and HuggingFace use different kernels to do the actual internal computation of the model. So these kernels are the things that make things like vLLM really fast. But these things, this then results in subtle numerical differences between the completions t…”
Nathan Lambert Nov 20, 2025 ▶ 1:15:29 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Most AI labs probably use evolved GRPO rather than PPO
“In reality, it seems like most people are using something like an evolved version of GRPO, which is a bit simpler than PPO.”
Nathan Lambert Nov 20, 2025 ▶ 1:16:39 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
LATENT SPACE Assertion Not checkable as stated
Lambert: Academia relied on UltraFeedback for open preference tuning for a year
“The academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular. And still a year later is like this state of the art data set for open preferen…”
Nathan Lambert Jul 31, 2025 ▶ 3:15 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Lambert: Context compression is crucial for long-horizon AI agents
“Compressing context, like that's not I don't think that's really a verifiable thing, but that being messed up, like that's a super crucial skill for long context actions and long longer tasks is just compressing well, and that's going to take some training nov…”
Nathan Lambert Jul 31, 2025 ▶ 9:50 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Assertion Supported
Lambert: Frontier AI labs still rely on human preference data
“Every time I check in with people at frontier labs, they're like, yeah, we still use human preference data.”
Nathan Lambert Jul 31, 2025 ▶ 12:12 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Open · timeframe Jul 2028
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Nathan Lambert Jul 31, 2025 ▶ 15:03 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

Show 24statements(60 left)

The other half of the tape: Nathan Lambert's own voice is left out of every number here. Other people bring the name up 6 times in 3 episodes across the shows. every mention, with the transcript →

Who brings them up most Alessio Fanelli 3Luca Soldani 2Shawn Wang 1

Every mention by year

tap a year for its mentions
00214220242025episodesmentions
01220242025episodes it came up in
00112220242025episodesmentions per episode

Latent Space 6

2025 2 mentions in 1 episode
2024 4 mentions in 2 episodes 2 per episode

One line per show, most statements first. The link opens Nathan's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER Founder, Interconnects AI 3 +1 109 87% 20/23 full record on Latent Space →
MADLEDGER Founder, Interconnects AI 1 23 80% 4/5 full record on the MAD Podcast →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.