Everything Nathan Lambert said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Lambert: Chinese open AI models currently do not contain backdoors
“Like, you can't prove that the models aren't doing certain backdoors, where I'm fairly certain they definitely aren't now.”
Lambert: AI progress will yield steady improvements rather than rapid singularity
“I think these researchers are going to grind out improvements for multiple years, but never in a way that results in this kind of accelerating well that we get drawn into.”
Lambert: OLMo 3 32B base model matches Qwen 2.5 32B quality
“This base model is similar in quality to the best available, which is like Quinn's 2.5, 32 B is, was still the best base model.”
Lambert: OLMo 3 7B outperforms Meta's Llama 3.1 8B in internal tests
“And I just think of this cause like Lama 3.1 AP is one of the most used models and hugging base of all time. And this should be better. We're, In our measurements, we see it as being better than Llama.”
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Lambert: Chinese companies with $1B+ valuations routinely pirate SaaS software
“Mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software.”
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Lambert: Deep Research relies on modular RL tasks rather than end-to-end outcomes
“I think the deep research blog post kind of hints that they do a bunch of small scale RL and then poof, the system works. Which I think is much more of what's happening is people train on a bunch of small things and they do some prompting and they see that whe…”
Lambert: RLHF can never be permanently solved
“In the same way that chatbot arena can never be saturated. RLHF can never be solved.”
Lambert: The RL algorithm is not the most important component in reasoning models
“I definitely don't think the algorithm tends to be the most important thing.”
Lambert: Hybrid reasoners may be phased out except for niche uses
“I think in plenty of ways, like hybrid reasoners might just be aged out except for niche applications because quality is so much more important than having a hundred X less inference tokens. It's like you just pay for it and compute and that'll get better.”
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Lambert: Top AI talent is dramatically cheaper than GPU clusters
“Talent is cheaper than GPUs by a dramatic margin, and At the end of the day, it's like, okay, if we're spending this much, they go to the room and they stare in the mirror and you're like, wait, it might not actually be that ridiculous to spend this money on t…”
Lambert: OLMo 32B roughly matches original GPT-4 level while fully open
“Like Olmo-Thirty-Tube is if you squint like original GPT-IV level and fully open.”
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Lambert: Startups should avoid RLHF unless it offers niche advantage
“I don't really recommend most startups to do it unless it's like going to provide them a clear competitive advantage in their kind of niche. Yeah. Because you're not going to make your model ChatGPT like better than OpenAI or anything like that.”
Lambert: A 100x smaller language model filters output better than RLHF
“You could use like a hundred times smaller language model and do much better at filtering than RLHF”
Lambert: Chatbot Arena is the best available evaluation benchmark for LLMs
“I have, if we make it to evaluation, I'd pretty much say that Chat Arena is the best limited evaluation that people have to learn how to use language models, and like, It's very valuable data”
Lambert: Reinforcement Learning Does Far Less for Alignment Than People Think
“It's like, now you're getting to the point where you don't even really need this to get a good model, so that's why it's like, okay, the RL is such a small part of the actual, like, doing RLHF. Like, RLHF is a metaphor for, like, all language model adaptation,…”
Lambert: RLHF has not been shown to improve underlying model benchmark capabilities
“RLHF is not that shown to improve capabilities yet. I think one of the fun ones is from the GPT-IV technical report. They essentially listed their kind of bogus evaluations, because it's a hilarious table, because it's like LSAT AP exams, and then like AMC-X a…”