AI2

8 statements across 1 episodes · 1 bullish · 0 bearish · 2 people on the record · first statement Nov 20, 2025 by Nathan Lambert · across every show →

Everything said about AI2, oldest first

Nov 20, 2025
Assertion Supported
Lambert: Hugging Face outcompeted AI2's AllenNLP library
“It was the main competitor to Hugging Face Transformers. And they ultimately outcompeted AI two as the thing that people use for that because they had very different model and amount of support.”
Nathan Lambert Nov 20, 2025 ▶ 32:47 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 20, 2025
Disclosure
Ai2 samples 6T tokens from 10T pool for OLMo 3
“There's like a pool of about 10 trillion tokens from which we have like an algorithm also fully open source. To like sample about six trillion tokens that we use during training.”
Luca Soldaini Nov 20, 2025 ▶ 6:07 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 20, 2025 neutral
Assertion Not checkable as stated
Soldaini: Most open AI models are open weights, not open source
“Majority of models that get release I think the best term to describe them is open weights. Your Quinn, your Gemma, your Lama you know, Kimi it's what gets release is a set of weights that correspond either to the final state of model, that's the most common, …”
Luca Soldaini Nov 20, 2025 ▶ 10:52 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 20, 2025
Disclosure
Lambert: Ai2 generated billions of DeepSeek completions over a weekend
“We had a bunch of cloud credits and I, they were running out and we're behind and I just generated like as many completions as possible. So it was like a few billion completions from deep seek over the weekend.”
Nathan Lambert Nov 20, 2025 ▶ 1:08:18 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 20, 2025 positive
Assertion Supported
Lambert: OLMo 3 models are the best open models outside Qwen 3
“I would say in post training where The best models that don't start with Quinn three and we're like reasonable to say that they are comparable to Quinn three, like on some benchmarks would beat them on some benchmarks. They're way ahead.”
Nathan Lambert Nov 20, 2025 ▶ 8:59 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 20, 2025
Disclosure
Ai2 fine-tuned OLMo 3 using Chinese teacher models DeepSeek-R1 and Qwen
“So in our case, we took a mix of existing data sets like Open Thoughts three and modified it, which is from Bespoke AI labs, a startup. And then we also generated a whole bunch of new data. So we ended up using a mix of teachers from like Deep Seek R one, oh f…”
Nathan Lambert Nov 20, 2025 ▶ 57:15 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 20, 2025
Disclosure
Ai2 filtered OLMo 3's pre-training dataset from 300 trillion tokens
“Our initial pool was closer to 300 trillion tokens. You shrink it down till you reach your target number, and hopefully as you shrink, you only keep the best part of this.”
Luca Soldaini Nov 20, 2025 ▶ 48:44 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Nov 20, 2025 neutral
Disclosure
Lambert: AI2 coined 'reinforcement learning with verifiable rewards' replicating Llama 3
“We spent a long time to try to replicate what we thought was close to Lama three post training with multiple stages and optimizers, which is the project that like came up with the name reinforcement learning with verifiable rewards with a bunch of people.”
Nathan Lambert Nov 20, 2025 ▶ 29:22 Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.