The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Lambert: Chinese open AI models currently do not contain backdoors
“Like, you can't prove that the models aren't doing certain backdoors, where I'm fairly certain they definitely aren't now.”
Lambert: AI progress will yield steady improvements rather than rapid singularity
“I think these researchers are going to grind out improvements for multiple years, but never in a way that results in this kind of accelerating well that we get drawn into.”
Lambert: OLMo 3 32B base model matches Qwen 2.5 32B quality
“This base model is similar in quality to the best available, which is like Quinn's 2.5, 32 B is, was still the best base model.”
Lambert: OLMo 3 7B outperforms Meta's Llama 3.1 8B in internal tests
“And I just think of this cause like Lama 3.1 AP is one of the most used models and hugging base of all time. And this should be better. We're, In our measurements, we see it as being better than Llama.”
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Lambert: Chinese companies with $1B+ valuations routinely pirate SaaS software
“Mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software.”
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”
Lambert: As AI funding grows, fewer researchers speak in public
“There's so much money in AI and it only becomes increasingly so that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller.”
Molmo Reads Clocks but Fails to Generalize to Dials
“The model didn't work on clocks and then the lead was really on clocks and no models work on clocks. So they're like, we've got to make it work on clocks. One of the interesting things is that it doesn't work on dials, even though it works on clocks.”
Lambert: Frontier AI labs do not recruit economics or social choice academics
“The RLHF techniques that people use were built in, like, labs like OpenAI and DeepMind, where there are some of these people, they have, they, these places do a pretty good job of trying to get these people in the door when you compare them to, like, startups,…”
Lambert: Open source will learn to train models on arbitrary preference data
“I really think people in open source and academics are going to figure out how to use any preference data on any model just because they're scrappy.”
Lambert: GPT-4 achieves 80% preference labeling agreement versus 70% for humans
“Essentially, people also think that synthetic data is, like, GPT-IV is more accurate than humans at labeling preferences, so if you look at these diagrams, like, humans are about 60 to 70% agreement, or, like, that's what the models get to, and if humans are a…”
Lambert: OpenAI Will Not Aggressively Ban Synthetic Training Scraping
“I don't expect OpenAI to go too crazy on this, because they're just gonna, there's gonna be so much backlash against them.”
Lambert: RLHF reward models achieve only 65% to 75% validation agreement
“If you look at a test set, you'll have a chosen and rejected, and you can take the reward model you're training, pass in those completions, And you see if the chosen predicted reward, so the scalar number is higher than the rejected predicted reward, and this …”
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Lambert: GPT-4 Turbo Gap Over Original GPT-4 Exceeds TÜLU 2 to GPT-4 Gap
“So it's like the difference from these, the GPT-IV Turbo to like the GPT-IV that was first released is bigger than the difference from Tulu-II to GPT-IV.”
Lambert: OpenAI retrains reward models with curated and user prompt mixtures
“And this is like a sort of outer loop optimization that no one in the open is even remotely qualified to talk about, but OpenAI does monitor and they'll like rerun RLHF and train a new reward model with a mixture of their curated data and user prompts to try t…”
Lambert: OLMo 3 models are the best open models outside Qwen 3
“I would say in post training where The best models that don't start with Quinn three and we're like reasonable to say that they are comparable to Quinn three, like on some benchmarks would beat them on some benchmarks. They're way ahead.”
Lambert: Alibaba's Qwen 3 VL vision model is a superior text model
“They released these Quinn three VL, their vision models. And like on text only benchmarks, it's way better than the models they released in April. So it's like okay, like that's the new baseline. And most people don't know about it because they think it's just…”
Lambert predicts more US labs will release open AI models
“If you look at this podcast in the coming months, I do think there's going to be, look like there's a lot more labs in the U S participating.”
Lambert: Hugging Face outcompeted AI2's AllenNLP library
“It was the main competitor to Hugging Face Transformers. And they ultimately outcompeted AI two as the thing that people use for that because they had very different model and amount of support.”
Lambert: Long-context extension is essential for reasoning AI models
“Three is long context extension, which is absolutely essential for these reasoning models because they generate so many intermediate tokens before sharing an answer with you.”
Lambert: Larger pre-trained base models are easier to improve with RL
“A better base model and a bigger base model is much easier to improve with RL.”
Lambert: Kernel differences between vLLM and Hugging Face cause RL numerical instability
“VLLM and HuggingFace use different kernels to do the actual internal computation of the model. So these kernels are the things that make things like vLLM really fast. But these things, this then results in subtle numerical differences between the completions t…”
Lambert: Most AI labs probably use evolved GRPO rather than PPO
“In reality, it seems like most people are using something like an evolved version of GRPO, which is a bit simpler than PPO.”
Lambert: Academia relied on UltraFeedback for open preference tuning for a year
“The academic community had been using this one data set since like all the way back in the hugging face models of like Zephyr beta is when this ultra feedback data set got popular. And still a year later is like this state of the art data set for open preferen…”
Lambert: Frontier AI labs still rely on human preference data
“Every time I check in with people at frontier labs, they're like, yeah, we still use human preference data.”
Lambert: LMSYS is probably setting up a deep research arena
“I mean, they're probably setting up a deep research arena, because that's the data that, I mean, if I was open AI working on deep research, that's the data that I want, and there are competitors, and LMSYS is the entity that has the market placement to set it …”
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Lambert: All major frontier AI labs will build their own search indexes
“I think they'll all do end up doing their own index and it should, it's one of those things that's like Google should have an advantage again, but who knows if they do.”
Lambert: Academics cannot match industry compute on Humanity's Last Exam
“I just think it's kind of unlikely that we're going to win as a academic and a state of the art number because they're going to start spending millions of tokens per query. And it's just a lot of, it's a lot of compute burn. Like the getting, beating that on t…”
Lambert: Current Language Models Cannot Prioritize Experiments for Multi-Week Research Plans
“So it's like, how do you come up with a research plan in 10 weeks? Like there's a lot of, how do you prioritize which experiments to do? It's like, there's a lot of inductive biases that go into that, that I don't like a language model would not do well at tha…”
Lambert: OpenAI's open model will be best-in-class in its size category
“I expected. It'll be best in class for some size Category in some subset of tasks. That's like, OpenAI only does things like that.”
Lambert: Jony Ive and OpenAI hardware will run in the cloud
“I think that thing will run on the cloud. I don't think that'll run local anyways.”
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Lambert: Open community will eventually match OpenAI's large-scale RL infrastructure
“And this is something that these early relative models are not going to be doing because we don't like, no one has this infrastructure like open AI does. It'll take a while to do that, but people will make it.”
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Lambert: Open judge models will become core open RL infrastructure
“We already have a bunch of open models that are doing like judge of models and Prometheus and other things that are designed specifically for LM as a judge. And I see that continuing to just become part of this kind of open RL infrastructure.”
Lambert: Open-source and research communities lack large-scale RLHF capabilities
“At the same time, we don't have a lot of that in the like open and research communities at the same scale.”
Nathan Lambert: RL will remain a distinct field from language modeling
“I think in the long run it will still settle out, or RL will still be a field that people work on just because of these kind of fundamental things that I talked about, that it's just viewing the whole problem formulation different than predicting text, really,…”