The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 50 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 4 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Not checkable as stated
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Nathan Lambert Jul 31, 2025 ▶ 20:52 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Nathan Lambert Jul 31, 2025 ▶ 1:16:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Not checkable as stated
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Nathan Lambert Oct 13, 2024 ▶ 12:15 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Molmo Reads Clocks but Fails to Generalize to Dials
“The model didn't work on clocks and then the lead was really on clocks and no models work on clocks. So they're like, we've got to make it work on clocks. One of the interesting things is that it doesn't work on dials, even though it works on clocks.”
Nathan Lambert Oct 13, 2024 ▶ 28:34 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Not checkable as stated
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Nathan Lambert Jan 11, 2024 ▶ 14:06 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Not checkable as stated
Lambert: Open source will learn to train models on arbitrary preference data
“I really think people in open source and academics are going to figure out how to use any preference data on any model just because they're scrappy.”
Nathan Lambert Jan 11, 2024 ▶ 47:48 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Nathan Lambert Jan 11, 2024 ▶ 48:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Not checkable as stated
Lambert: OpenAI Will Not Aggressively Ban Synthetic Training Scraping
“I don't expect OpenAI to go too crazy on this, because they're just gonna, there's gonna be so much backlash against them.”
Nathan Lambert Jan 11, 2024 ▶ 50:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”
Nathan Lambert Jan 11, 2024 ▶ 54:59 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Not checkable as stated
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Nathan Lambert Jan 11, 2024 ▶ 1:19:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Nathan Lambert Jan 11, 2024 ▶ 1:26:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Not checkable as stated
Lambert: Academia relied on UltraFeedback for open preference tuning for a year
“The academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular. And still a year later is like this state of the art data set for open preferen…”
Nathan Lambert Jul 31, 2025 ▶ 3:15 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Supported
Lambert: Frontier AI labs still rely on human preference data
“Every time I check in with people at frontier labs, they're like, yeah, we still use human preference data.”
Nathan Lambert Jul 31, 2025 ▶ 12:12 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Open · timeframe Jul 2028
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Nathan Lambert Jul 31, 2025 ▶ 15:03 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Nathan Lambert Jul 31, 2025 ▶ 22:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Open · timeframe Jul 2030
Lambert: All major frontier AI labs will build their own search indexes
“I think they'll all do end up doing their own index and it should, it's one of those things that's like Google should have an advantage again, but who knows if they do.”
Nathan Lambert Jul 31, 2025 ▶ 24:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Open · timeframe Jul 2028
Lambert: Academics cannot match industry compute on Humanity's Last Exam
“I just think it's kind of unlikely that we're going to win as a academic and a state of the art number because they're going to start spending millions of tokens per query. And it's just a lot of, it's a lot of compute burn. Like the getting, beating that on t…”
Nathan Lambert Jul 31, 2025 ▶ 37:12 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Not checkable as stated
Lambert: Current Language Models Cannot Prioritize Experiments for Multi-Week Research Plans
“So it's like, how do you come up with a research plan in 10 weeks? Like there's a lot of, how do you prioritize which experiments to do? It's like, there's a lot of inductive biases that go into that, that I don't like a language model would not do well at tha…”
Nathan Lambert Jul 31, 2025 ▶ 46:31 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Not checkable as stated
Lambert: OpenAI's open model will be best-in-class in its size category
“I expected. It'll be best in class for some size Category in some subset of tasks. That's like, OpenAI only does things like that.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:07 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Open · timeframe Jul 2028
Lambert: Jony Ive and OpenAI hardware will run in the cloud
“I think that thing will run on the cloud. I don't think that'll run local anyways.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:54 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Prediction Not checkable as stated
Lambert: Open community will eventually match OpenAI's large-scale RL infrastructure
“And this is something that these early relative models are not going to be doing because we don't like, no one has this infrastructure like open AI does. It'll take a while to do that, but people will make it.”
Nathan Lambert Jan 2, 2025 ▶ 8:34 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Nathan Lambert Jan 2, 2025 ▶ 9:57 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Prediction Not checkable as stated
Lambert: Open judge models will become core open RL infrastructure
“We already have a bunch of open models that are doing like judge of models and Prometheus and other things that are designed specifically for LM as a judge. And I see that continuing to just become part of this kind of open RL infrastructure.”
Nathan Lambert Jan 2, 2025 ▶ 13:25 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Assertion Not checkable as stated
Lambert: Open-source and research communities lack large-scale RLHF capabilities
“At the same time, we don't have a lot of that in the like open and research communities at the same scale.”
Nathan Lambert Jan 11, 2024 ▶ 5:38 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Not checkable as stated
Nathan Lambert: RL will remain a distinct field from language modeling
“I think in the long run it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about, that it's just viewing the whole problem formulation different than predicting text, really,…”
Nathan Lambert Jan 11, 2024 ▶ 7:40 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Not checkable as stated
Lambert: AI community will clarify if chain-of-thought maps to RL within a year
“I think in the next year that'll probably get kind of made more concrete by the community on, like, if you can easily draw out, like, if chain of thought reasoning is more like RL.”
Nathan Lambert Jan 11, 2024 ▶ 18:00 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: DPO benchmark gains rely largely on the UltraFeedback dataset
“Everyone's using this ultra feedback data set and it boosts AlpacaVal, MTBench, TruthfulQA, and like the qualitative model a bit. We don't really know why.”
Nathan Lambert Jan 11, 2024 ▶ 29:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: Anthropic, ChatGPT, and Bard Use Post-Generation Moderation Classifiers
“Anthropic and ChatGPT and Bard almost surely have a classifier after, which is like, is this text good? Is this text bad?”
Nathan Lambert Jan 11, 2024 ▶ 45:13 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Not checkable as stated
Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data
“So I would say, still say, like, six to eight million is safe to say that they're spending, if not more, they're probably also buying other types of data and or throwing out data that they don't like.”
Nathan Lambert Jan 11, 2024 ▶ 46:51 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Held up
Lambert: Practitioners Will Adopt Constitutional AI for Preferences in 2024
“I think in twenty-twenty-four at some point people will start doing things like constitutional AI for preferences.”
Nathan Lambert Jan 11, 2024 ▶ 51:25 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Not checkable as stated
Lambert: Zephyr was the first open model to succeed with public RLHF
“I think Zephyr was the first model that showed success with RLHF in the public, but that's a long time from everyone knowing that it was something that people are interested in to having any, like, check mark.”
Nathan Lambert Jan 11, 2024 ▶ 51:42 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Nathan Lambert Jan 11, 2024 ▶ 1:03:01 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Not checkable as stated
Lambert: Offline RL could take off due to simpler training pipelines
“There's a few papers that people have published Not a lot of traction. I think it could take off. Some people that I know in the RLHF area really think a lot of people are doing this in industry just because it makes the kind of training process simpler and th…”
Nathan Lambert Jan 11, 2024 ▶ 1:04:08 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Prediction Held up
Lambert: More DPO models will emerge than any other method
“I expect to see more DPO models than anything else in the next six months.”
Nathan Lambert Jan 11, 2024 ▶ 1:15:07 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: DPO Has Become the Standard Release Expectation for Open-Source LLMs
“I think DPO releases are kind of becoming expected because Mistral released a DPO model as well. I think the slide after this is just like, there's a ton. It's like Intel releases DPO models, Stability releases DPO models. At some point, you just have to accep…”
Nathan Lambert Jan 11, 2024 ▶ 1:20:43 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: GPT-4 Turbo Showed a Noticeable Jump on LMSYS Chatbot Arena
“GPT-IV Turbo is also notably ahead of the other GPT-IVs, which it kind of showed up immediately once they added it to the leaderboard, or to the arena, and I was like, all the GPT-IV memes aside, it seems like this is effectively a bump in the model.”
Nathan Lambert Jan 11, 2024 ▶ 1:25:01 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: DeepSeek-R1 starts solving math questions immediately without explicit planning
“If you look at DeepSeq R-One and you ask it a hard math question, it's not like, here's my plan of attack. It just starts.”
Nathan Lambert Jul 31, 2025 ▶ 40:10 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Prediction Held up
Lambert: Labs will surely use parallel-compute models to generate synthetic data
“Well, I bet people, I mean, they surely will use these for synthetic data. It's just like the marginal gain on synthetic data is always very high.”
Nathan Lambert Jul 31, 2025 ▶ 50:22 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Not checkable as stated
Lambert: Long inference generations break RL infrastructure and require more GPUs
“The inference, high inference length generations definitely just, like, kind of breaks all infrastructure, because there's just so many tokens, there's more opportunity for out of memory or other things to go wrong. So it's like, just on a default, all of your…”
Nathan Lambert Jul 31, 2025 ▶ 1:00:39 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Assertion Supported
Lambert: Llama 3.1 math evals rely on SymPy and LLM judges
“Lama, 3.1 details their vows for math. They use both SIM by a Python process or Python package for extraction and it's a judge to extract their answers for math.”
Nathan Lambert Jan 2, 2025 ▶ 12:38 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Assertion Supported
Molmo Uses Base Model Without Instruction Tuning or Chat Template
“This is just, like, straight base model, no real instruction tuning. There's literally, like, no chat template for multi-turn. It just concatenates the messages together and, like, there's, like, go, look, good luck.”
Nathan Lambert Oct 13, 2024 ▶ 14:17 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Not checkable as stated
Lambert: Real-Time Human Feedback in LLM RL Loops Is Far From Feasible
“Setting up the infrastructure to take tens of thousands of prompts and generate them and then show them to a human and collect the human responses and then show that Shove that into your training architecture is very far away from working, so we don't really h…”
Nathan Lambert Jan 11, 2024 ▶ 16:13 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: Training LLM reward models on 0-to-10 ratings failed
“People tried that with language models, which is if you have a prompt and a completion and you just have someone rate it from zero to 10, could you then train a reward model on all of these completions and zero to 10 ratings and see if you could actually chang…”
Nathan Lambert Jan 11, 2024 ▶ 37:43 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Lambert: Anthropic and OpenAI reward model loss functions are mathematically identical
“Fun fact is that these loss functions Look different and anthropic in opening eyes papers, but they're just literally just log transform. So if you start like expantiating both sides and taking the log of both sides, you'll like converge on one of the two, the…”
Nathan Lambert Jan 11, 2024 ▶ 54:41 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.