The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 35 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Insight
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Nathan Lambert Jul 31, 2025 ▶ 16:42 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Nathan Lambert Jul 31, 2025 ▶ 19:28 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Nathan Lambert Jul 31, 2025 ▶ 24:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Nathan Lambert Jul 31, 2025 ▶ 1:03:38 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Top AI talent is dramatically cheaper than GPU clusters
“Talent is cheaper than GPUs by a dramatic margin, and At the end of the day, it's like, okay, if we're spending this much, they go to the room and they stare in the mirror and you're like, wait, it might not actually be that ridiculous to spend this money on t…”
Nathan Lambert Jul 31, 2025 ▶ 1:13:46 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Nathan Lambert Jan 2, 2025 ▶ 5:23 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Insight
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Nathan Lambert Jan 11, 2024 ▶ 29:37 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Nathan Lambert Jan 11, 2024 ▶ 45:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Nathan Lambert Jan 11, 2024 ▶ 58:36 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: RLVR is broader than ground truth because code is verifiable
“The verifiable rewards is actually a more general notion because only like math questions have a ground truth where code is verifiable, precise instruction following is verifiable.”
Nathan Lambert Jul 31, 2025 ▶ 5:13 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Inference scaling plots misleadingly suggest search is an easy control knob
“The core of that article is just, they're taking points from within training, or there's a natural variance, and then you line them up. And if you line them up, then you get this nice inference time-scaling behavior, which is, and now people, a lot of people h…”
Nathan Lambert Jul 31, 2025 ▶ 28:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: AI academics must build datasets and evals rather than papers
“If you're trying to have impact in AI right now, it's as an academic, you have to like level up out of papers to artifacts, which is models, datasets, evals. Datasets and evals are easier for people to have impact on.”
Nathan Lambert Jul 31, 2025 ▶ 36:13 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: RLVR on math does not degrade knowledge benchmark performance
“I think part of the intuition of RLVR is that the model is good at knowing which prompt area it is, which is why the models don't get worse on knowledge benchmarks if you're trading on like just math or precise instruction following. So the model just kind of …”
Nathan Lambert Jul 31, 2025 ▶ 59:33 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: AI field hasn't learned much more, just articulates ignorance better
“I think it's, I try to do it every six or 12 months is my current, is my estimated cadence, just to refine the ways that I say things, and people will see That we don't know that much more, but we have a bit of better way of saying what we don't know.”
Nathan Lambert Jan 11, 2024 ▶ 3:58 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: RLHF reward models lack inductive biases for true preferences
“In RLHF calling things a preference model is a little annoying because There's no inductive bias of what a preference is.”
Nathan Lambert Jan 11, 2024 ▶ 12:41 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Claude's constitution dictates output priorities, not model beliefs
“If you look at Claude's constitution, like, that doesn't mean the model believes these things. It's just trying Trained and to prioritize these things.”
Nathan Lambert Jan 11, 2024 ▶ 13:06 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Context compression is crucial for long-horizon AI agents
“Compressing context, like that's not I don't think that's really a verifiable thing, but that being messed up, like that's a super crucial skill for long context actions and long longer tasks is just compressing well, and that's going to take some training nov…”
Nathan Lambert Jul 31, 2025 ▶ 9:50 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: North star of reasoning models is dynamic token budget calibration
“I think that has to be the north star for most people working on reasoning, which is the model will just Spend the right amount of tokens on it.”
Nathan Lambert Jul 31, 2025 ▶ 20:44 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Scaling RL long enough requires a curriculum of increasing difficulty
“If you scale RL long enough, You're going to need a curriculum of things getting harder. And like, that's pretty obvious.”
Nathan Lambert Jul 31, 2025 ▶ 32:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Reasoning models solved basic skills; planning is the next frontier
“So I came up with four and the foundational one was skills, which is What I would say that we have already done with O-one and R-one, which is you do a lot of RL, you show the inference time scaling works and you get really high benchmark numbers. And then the…”
Nathan Lambert Jul 31, 2025 ▶ 38:34 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Parallel compute provides robustness rather than low-probability search
“Well, I don't think we're using parallel compute in a way to search over like low probability tokens. We're using it to get robustness. If you use like O-one pro is, it was so nice because it Just had a very predictable depth to it, even on niche topics where …”
Nathan Lambert Jul 31, 2025 ▶ 47:45 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Code maintainability is a human preference problem in RL
“The software stuff is not easy because it's almost like maintainability almost feels like a human preference type issue again, where somebody could look at it and be like, yeah, that's not as good, but adding the heuristic and trading seems very messy.”
Nathan Lambert Jul 31, 2025 ▶ 53:24 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: RLVR is harder to over-optimize on math than code
“For math, it's a bit harder to over optimize, I think. Unless you have tools and the model learns to search and cheat instead of learning math, which I'm sure somebody could see that out in the world, which is like, oh, I'll just find the, you're training. It'…”
Nathan Lambert Jul 31, 2025 ▶ 57:25 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: Custom personality fine-tuning is open source AI's winning turf
“If open models are to win, part of it could be just, like, everybody can have exactly the model they want. We're serving GPT-IV. It's kind of its thing. You can prompt it, but if fine-tuning is more effective than prompting, everybody can have the model. That …”
Nathan Lambert Jul 31, 2025 ▶ 1:05:05 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Insight
Lambert: RLHF operates as a contextual bandit problem, not sequential RL
“What happens is the state is a prompt, and then you do a completion, and then you throw it away, and you grab a new prompt. Where really in like RL, you, as an RL researcher, you would think of this as being like, you take a state, you complete, Get some compl…”
Nathan Lambert Jan 11, 2024 ▶ 16:40 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: RLHF provides a richer signal per unit of compute than pre-training
“As reinforcement learning is so much less compute, like it, Is a richer signal in terms of its impact, because if they could do what RLHF is doing at pre-training, they would, but they don't know how to have that effect in like a stable manner. Otherwise, ever…”
Nathan Lambert Jan 11, 2024 ▶ 22:14 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Instruction tuning is more important than RLHF for most practitioners
“I think for most people, instruction tuning is probably still more important in their day-to-day life. I think instruction tuning works very well. You can write samples by hand that make sense. You can get the model to learn from them. You could do this with v…”
Nathan Lambert Jan 11, 2024 ▶ 28:11 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Scaling from 7B to 70B parameters fixes nuance and repetition
“I think the things that people see now is like the small models don't really handle nuance as well, and they could be more repetitive if, even if they have really good instruction tuning, but if you take that kind of seven to seventy billion parameter jump, li…”
Nathan Lambert Jan 11, 2024 ▶ 31:24 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: RLHF fails completely without a KL divergence constraint
“The KL constraint prevents that. There's not that much documented work on that, but there's a lot of people that know if you take that away, it just doesn't work at all.”
Nathan Lambert Jan 11, 2024 ▶ 35:51 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: RLHF avoids controversial preferences, focusing on correctness and style
“The reason this really is done on a deep level is that you're not actually trying to model any, like, contestable preference in this. Like, you're not trying to go into things that are controversial or anything. It's really, like, the notion of preference is t…”
Nathan Lambert Jan 11, 2024 ▶ 39:02 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Published AI training compute costs understate total experimentation budgets
“The compute costs listed in the paper always are way lower, because all they have to say is, like, what is one run cost, but they're running tens or hundreds of runs, so it's like, okay, like, it's kind of a meaningless number, yeah.”
Nathan Lambert Jan 11, 2024 ▶ 47:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: RLHF Performance Depends on Data and Systems Over RL Details
“It really ends up being kind of, like, gibberish that I think is less important now, because it's more about data and infrastructure than RL details, than, like, value functions and everything. A lot of the papers have different terms in the equations. I think…”
Nathan Lambert Jan 11, 2024 ▶ 57:58 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Anthropic Constitutional AI and OpenAI Superalignment share intellectual roots
“The constitutional AI and the super alignment is, like, very conceptually linked. It's like a group of people that has, like, a very similar intellectual upbringing, and they work together for a long time, like, coming to the same conclusions in different ways…”
Nathan Lambert Jan 11, 2024 ▶ 1:11:37 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think DPO is closer to RLHF than RLHF is to RL.”
Nathan Lambert Jan 11, 2024 ▶ 1:12:56 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Insight
Lambert: Autoregressive LLMs lack explicit internal structures for intermediate state
“Language models have no ability to do this. They are. Kind of per token computation devices where each token is outputted after doing this forward pass and within that there's no explicit structure to hold these intermediate states.”
Nathan Lambert Jan 2, 2025 ▶ 3:37 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.