Everything Nathan Lambert said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Lambert: As AI funding grows, fewer researchers speak in public
“There's so much money in AI and it only becomes increasingly so that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller.”
Lambert: Rich Sutton's RL theories are impractical for models like GPT-6
“Rich is a font of wonderful ideas, but Often not ones that are going to be immediately practical. This is how you get things like creating reinforcement learning, but not necessarily things that are going to impact what GPT six is.”
Ai2 fine-tuned OLMo 3 using Chinese teacher models DeepSeek-R1 and Qwen
“So in our case, we took a mix of existing data sets like Open Thoughts three and modified it, which is from Bespoke AI labs, a startup. And then we also generated a whole bunch of new data. So we ended up using a mix of teachers from like Deep Seek R one, oh f…”
Lambert: Ai2 generated billions of DeepSeek completions over a weekend
“We had a bunch of cloud credits and I, they were running out and we're behind and I just generated like as many completions as possible. So it was like a few billion completions from deep seek over the weekend.”
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Lambert: RLVR is broader than ground truth because code is verifiable
“The verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.”
Lambert: Inference scaling plots misleadingly suggest search is an easy control knob
“The core of that article is just, they're taking points from within training, or there's a natural variance, and then you line them up. And if you line them up, then you get this nice inference time-scaling behavior, which is, and now people, a lot of people h…”
Lambert: AI benchmarks like ARC-AGI should prioritize testing without harnesses
“Harnesses are cool, but they're gonna, they're,
They're a handicap that's changing the learning dynamics substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses.”
Lambert: AI academics must build datasets and evals rather than papers
“If you're trying to have impact in AI right now, it's as an academic, you have to like level up out of papers to artifacts, which is models, datasets, evals. Datasets and evals are easier for people to have impact on.”
Lambert: Labs trade code usability for massive RL performance gains
“That's just like the labs are trading off massive gains in performance or small detriments in usability. And it's like, do you ship that model? Yeah. Like you just ship it and deal with it later, but I'm sure they could, I'm sure that's a fixable thing.”
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Lambert: The local model community is much smaller than assumed
“Like the local modeling community, I think is much smaller than people give it credit for, because most of the use for open models is still in APIs.”
Lambert: Meta withholding its leading benchmark model is bad execution
“But to be a model that claims to be open and then not release the model that is your leading claim is just, like, that is, like, bad execution.”
Lambert: AI is developing new reasoning modes that look less human
“I think a big Trend of the year is that we're seeing new types of language model reasoning that look less human and that Can be good for kind of separating the discourse for expecting a really narrow type of behaviors.”
Lambert: OpenAI's o1 uses token streams as intermediate state compute
“Why oh, one is exciting is because it's a new type of language models that are going to maximize on this view of reasoning, which is that chain of thought in kind of a forward stream of tokens can actually do a lot to achieve better outcomes when you're doing …”
Molmo Reads Clocks but Fails to Generalize to Dials
“The model didn't work on clocks and then the lead was really on clocks and no models work on clocks. So they're like, we've got to make it work on clocks. One of the interesting things is that it doesn't work on dials, even though it works on clocks.”
Lambert: OpenAI's rumored Q* was likely just a moderate benchmark bump
“They probably just got like a moderate bump on one of their benchmarks, and then everyone lost their minds, so it doesn't really matter.”
Lambert: AI field hasn't learned much more, just articulates ignorance better
“I think it's, I try to do it every six or 12 months is my current, is my estimated cadence, just to refine the ways that I say things, and people will see That we don't know that much more, but we have a bit of better way of saying what we don't know.”
Lambert: Reinforcement learning in language models is contrived and not real RL
“And the view of RL in language models is pretty contrived already. So it's not like we're doing real RL.”
Lambert: RLHF reward models lack inductive biases for true preferences
“In RLHF calling things a preference model is a little annoying because There's no inductive bias of what a preference is.”
Lambert: Claude's constitution dictates output priorities, not model beliefs
“If you look at Claude's constitution, like, that doesn't mean the model believes these things. It's just trying Trained and to prioritize these things.”
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Lambert: Open-source claims of ChatGPT-level performance are overblown
“I think the claims of ChatGPT level are long overblown in most of the things in open source.”
Lambert: Annotator disagreement data in RLHF is not used for anything
“I don't really think this disagreement data is used for anything, but it's good to know, like, what the distribution of prompts is, who's doing it, how many samples you have, controlling the workforce.”