The Ledger, every show

Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.

shows every show 44 of 44
every show
clear all ✕
MAD Assertion Not checkable as stated
Lambert: Chinese open AI models currently do not contain backdoors
“Like, you can't prove that the models aren't doing certain backdoors, where I'm fairly certain they definitely aren't now.”
Nathan Lambert Nov 20, 2025 ▶ 17:52 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Prediction Not checkable as stated
Lambert: AI progress will yield steady improvements rather than rapid singularity
“I think these researchers are going to grind out improvements for multiple years, but never in a way that results in this kind of accelerating well that we get drawn into.”
Nathan Lambert Nov 20, 2025 ▶ 1:21:34 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: OLMo 3 32B base model matches Qwen 2.5 32B quality
“This base model is similar in quality to the best available, which is like Quinn's 2.5, 32 B is, was still the best base model.”
Nathan Lambert Nov 20, 2025 ▶ 4:06 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: OLMo 3 7B outperforms Meta's Llama 3.1 8B in internal tests
“And I just think of this cause like Lama 3.1 AP is one of the most used models and hugging base of all time. And this should be better. We're, In our measurements, we see it as being better than Llama.”
Nathan Lambert Nov 20, 2025 ▶ 4:54 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Partly supported
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Nathan Lambert Nov 20, 2025 ▶ 17:01 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Chinese companies with $1B+ valuations routinely pirate SaaS software
“Mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software.”
Nathan Lambert Nov 20, 2025 ▶ 18:47 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Nathan Lambert Nov 20, 2025 ▶ 58:29 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Prediction Not checkable as stated
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Nathan Lambert Nov 20, 2025 ▶ 1:24:19 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
LATENT SPACE Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Nathan Lambert Jul 31, 2025 ▶ 2:20 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Not checkable as stated
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Nathan Lambert Jul 31, 2025 ▶ 20:52 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Assertion Partly supported
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Nathan Lambert Jul 31, 2025 ▶ 1:16:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Not checkable as stated
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Nathan Lambert Oct 13, 2024 ▶ 12:15 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
LATENT SPACE Assertion Supported
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Nathan Lambert Jan 11, 2024 ▶ 59:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
MAD Assertion Not checkable as stated
Lambert: As AI funding grows, fewer researchers speak in public
“There's so much money in AI and it only becomes increasingly so that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller.”
Nathan Lambert Nov 20, 2025 ▶ 29:52 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
LATENT SPACE Assertion Supported
Molmo Reads Clocks but Fails to Generalize to Dials
“The model didn't work on clocks and then the lead was really on clocks and no models work on clocks. So they're like, we've got to make it work on clocks. One of the interesting things is that it doesn't work on dials, even though it works on clocks.”
Nathan Lambert Oct 13, 2024 ▶ 28:34 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
LATENT SPACE Assertion Not checkable as stated
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Nathan Lambert Jan 11, 2024 ▶ 14:06 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Prediction Not checkable as stated
Lambert: Open source will learn to train models on arbitrary preference data
“I really think people in open source and academics are going to figure out how to use any preference data on any model just because they're scrappy.”
Nathan Lambert Jan 11, 2024 ▶ 47:48 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Nathan Lambert Jan 11, 2024 ▶ 48:53 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Prediction Not checkable as stated
Lambert: OpenAI Will Not Aggressively Ban Synthetic Training Scraping
“I don't expect OpenAI to go too crazy on this, because they're just gonna, there's gonna be so much backlash against them.”
Nathan Lambert Jan 11, 2024 ▶ 50:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”
Nathan Lambert Jan 11, 2024 ▶ 54:59 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Not checkable as stated
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Nathan Lambert Jan 11, 2024 ▶ 1:19:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Contradicted
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Nathan Lambert Jan 11, 2024 ▶ 1:26:04 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Nathan Lambert Jan 11, 2024 ▶ 1:32:27 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
MAD Assertion Supported
Lambert: OLMo 3 models are the best open models outside Qwen 3
“I would say in post training where The best models that don't start with Quinn three and we're like reasonable to say that they are comparable to Quinn three, like on some benchmarks would beat them on some benchmarks. They're way ahead.”
Nathan Lambert Nov 20, 2025 ▶ 8:59 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Supported
Lambert: Alibaba's Qwen 3 VL vision model is a superior text model
“They released these Quinn three VL, their vision models. And like on text only benchmarks, it's way better than the models they released in April. So it's like okay, like that's the new baseline. And most people don't know about it because they think it's just…”
Nathan Lambert Nov 20, 2025 ▶ 9:46 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Prediction Not checkable as stated
Lambert predicts more US labs will release open AI models
“If you look at this podcast in the coming months, I do think there's going to be, look like there's a lot more labs in the U S participating.”
Nathan Lambert Nov 20, 2025 ▶ 16:22 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Supported
Lambert: Hugging Face outcompeted AI2's AllenNLP library
“It was the main competitor to Hugging Face Transformers. And they ultimately outcompeted AI two as the thing that people use for that because they had very different model and amount of support.”
Nathan Lambert Nov 20, 2025 ▶ 32:47 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Long-context extension is essential for reasoning AI models
“Three is long context extension, which is absolutely essential for these reasoning models because they generate so many intermediate tokens before sharing an answer with you.”
Nathan Lambert Nov 20, 2025 ▶ 40:15 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Larger pre-trained base models are easier to improve with RL
“A better base model and a bigger base model is much easier to improve with RL.”
Nathan Lambert Nov 20, 2025 ▶ 44:29 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Supported
Lambert: Kernel differences between vLLM and Hugging Face cause RL numerical instability
“VLLM and HuggingFace use different kernels to do the actual internal computation of the model. So these kernels are the things that make things like vLLM really fast. But these things, this then results in subtle numerical differences between the completions t…”
Nathan Lambert Nov 20, 2025 ▶ 1:15:29 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
MAD Assertion Not checkable as stated
Lambert: Most AI labs probably use evolved GRPO rather than PPO
“In reality, it seems like most people are using something like an evolved version of GRPO, which is a bit simpler than PPO.”
Nathan Lambert Nov 20, 2025 ▶ 1:16:39 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
LATENT SPACE Assertion Not checkable as stated
Lambert: Academia relied on UltraFeedback for open preference tuning for a year
“The academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular. And still a year later is like this state of the art data set for open preferen…”
Nathan Lambert Jul 31, 2025 ▶ 3:15 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Assertion Supported
Lambert: Frontier AI labs still rely on human preference data
“Every time I check in with people at frontier labs, they're like, yeah, we still use human preference data.”
Nathan Lambert Jul 31, 2025 ▶ 12:12 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Open · timeframe Jul 2028
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Nathan Lambert Jul 31, 2025 ▶ 15:03 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Nathan Lambert Jul 31, 2025 ▶ 22:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Open · timeframe Jul 2030
Lambert: All major frontier AI labs will build their own search indexes
“I think they'll all do end up doing their own index and it should, it's one of those things that's like Google should have an advantage again, but who knows if they do.”
Nathan Lambert Jul 31, 2025 ▶ 24:23 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Open · timeframe Jul 2028
Lambert: Academics cannot match industry compute on Humanity's Last Exam
“I just think it's kind of unlikely that we're going to win as a academic and a state of the art number because they're going to start spending millions of tokens per query. And it's just a lot of, it's a lot of compute burn. Like the getting, beating that on t…”
Nathan Lambert Jul 31, 2025 ▶ 37:12 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Assertion Not checkable as stated
Lambert: Current Language Models Cannot Prioritize Experiments for Multi-Week Research Plans
“So it's like, how do you come up with a research plan in 10 weeks? Like there's a lot of, how do you prioritize which experiments to do? It's like, there's a lot of inductive biases that go into that, that I don't like a language model would not do well at tha…”
Nathan Lambert Jul 31, 2025 ▶ 46:31 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Not checkable as stated
Lambert: OpenAI's open model will be best-in-class in its size category
“I expected. It'll be best in class for some size Category in some subset of tasks. That's like, OpenAI only does things like that.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:07 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Prediction Open · timeframe Jul 2028
Lambert: Jony Ive and OpenAI hardware will run in the cloud
“I think that thing will run on the cloud. I don't think that'll run local anyways.”
Nathan Lambert Jul 31, 2025 ▶ 1:11:54 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Assertion Not checkable as stated
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Nathan Lambert Jan 2, 2025 ▶ 6:26 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
LATENT SPACE Prediction Not checkable as stated
Lambert: Open community will eventually match OpenAI's large-scale RL infrastructure
“And this is something that these early relative models are not going to be doing because we don't like, no one has this infrastructure like open AI does. It'll take a while to do that, but people will make it.”
Nathan Lambert Jan 2, 2025 ▶ 8:34 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
LATENT SPACE Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Nathan Lambert Jan 2, 2025 ▶ 9:57 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
LATENT SPACE Prediction Not checkable as stated
Lambert: Open judge models will become core open RL infrastructure
“We already have a bunch of open models that are doing like judge of models and Prometheus and other things that are designed specifically for LM as a judge. And I see that continuing to just become part of this kind of open RL infrastructure.”
Nathan Lambert Jan 2, 2025 ▶ 13:25 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
LATENT SPACE Assertion Not checkable as stated
Lambert: Open-source and research communities lack large-scale RLHF capabilities
“At the same time, we don't have a lot of that in the like open and research communities at the same scale.”
Nathan Lambert Jan 11, 2024 ▶ 5:38 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Prediction Not checkable as stated
Nathan Lambert: RL will remain a distinct field from language modeling
“I think in the long run it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about, that it's just viewing the whole problem formulation different than predicting text, really,…”
Nathan Lambert Jan 11, 2024 ▶ 7:40 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Prediction Not checkable as stated
Lambert: AI community will clarify if chain-of-thought maps to RL within a year
“I think in the next year that'll probably get kind of made more concrete by the community on, like, if you can easily draw out, like, if chain of thought reasoning is more like RL.”
Nathan Lambert Jan 11, 2024 ▶ 18:00 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: DPO benchmark gains rely largely on the UltraFeedback dataset
“Everyone's using this ultra feedback data set and it boosts AlpacaVal, MTBench, TruthfulQA, and like the qualitative model a bit. We don't really know why.”
Nathan Lambert Jan 11, 2024 ▶ 29:20 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.