Lambert predicts more US labs will release open AI models
“If you look at this podcast in the coming months, I do think there's going to be, look like there's a lot more labs in the U S participating.”
Lambert: Chinese open AI models currently do not contain backdoors
“Like, you can't prove that the models aren't doing certain backdoors, where I'm fairly certain they definitely aren't now.”
Lambert: As AI funding grows, fewer researchers speak in public
“There's so much money in AI and it only becomes increasingly so that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller.”
Lambert: Long-context extension is essential for reasoning AI models
“Three is long context extension, which is absolutely essential for these reasoning models because they generate so many intermediate tokens before sharing an answer with you.”
Lambert: Scaling AI 10x alters post-training, not pre-training methods
“If like, if we were to train a model that was 10 times as big, like all this post-training stuff would change. But the pre-training And mid training and long contacts, I think would actually become looking pretty similar.”
Lambert: Larger pre-trained base models are easier to improve with RL
“A better base model and a bigger base model is much easier to improve with RL.”
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Lambert: Ai2 generated billions of DeepSeek completions over a weekend
“We had a bunch of cloud credits and I, they were running out and we're behind and I just generated like as many completions as possible. So it was like a few billion completions from deep seek over the weekend.”
Lambert: AI progress will yield steady improvements rather than rapid singularity
“I think these researchers are going to grind out improvements for multiple years, but never in a way that results in this kind of accelerating well that we get drawn into.”
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Lambert: Scaling RL long enough requires a curriculum of increasing difficulty
“If you scale RL long enough,
You're going to need a curriculum of things getting harder. And like, that's pretty obvious.”
Lambert: AI benchmarks like ARC-AGI should prioritize testing without harnesses
“Harnesses are cool, but they're gonna, they're,
They're a handicap that's changing the learning dynamics substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses.”
Lambert: AI academics must build datasets and evals rather than papers
“If you're trying to have impact in AI right now, it's as an academic, you have to like level up out of papers to artifacts, which is models, datasets, evals. Datasets and evals are easier for people to have impact on.”
Lambert: Academics cannot match industry compute on Humanity's Last Exam
“I just think it's kind of unlikely that we're going to win as a academic and a state of the art number because they're going to start spending millions of tokens per query. And it's just a lot of, it's a lot of compute burn. Like the getting, beating that on t…”
Lambert: Reasoning models solved basic skills; planning is the next frontier
“So I came up with four and the foundational one was skills, which is What I would say that we have already done with O-one and R-one, which is you do a lot of RL, you show the inference time scaling works and you get really high benchmark numbers. And then the…”
Lambert: Long inference generations break RL infrastructure and require more GPUs
“The inference, high inference length generations definitely just, like, kind of breaks all infrastructure, because there's just so many tokens, there's more opportunity for out of memory or other things to go wrong. So it's like, just on a default, all of your…”
Lambert: AI2 received early industry tip on reinforcement fine-tuning
“We got a tip from a industry lab member to do this a few months early. So we got a head start”
AI2 Did Not Log Raw Audio for Pixmo Annotations
“Yeah, it was like Whisper, and I don't remember the details, and I had asked, and it's like, I don't think they actually logged the audio. I was like, oh, this would be super cool for, like, other types of multimodal, but I don't think it was in the terms of t…”