Nov 20, 2025 · 1h 28m · mad
Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews AI2 researchers Nathan Lambert and Luca Soldaini about the launch of the fully open-source OLMo 3 model family. They explore the distinction between open weights and true open source, analyze the global AI ecosystem, and detail the complete six-stage architecture behind state-of-the-art reasoning models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 16.4% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Nathan forcefully rejects Rich Sutton's podcast claim that pre-training is flawed, dismissing it as theoretical nerd-sniping that ignores practical engineering realities.
Hardest push from Matt ▶ 53:10 Matt holding Luca accountable to his past claim about data in long contextMatt directly challenges Luca by citing his prior statement that 'data doesn't matter for long context,' forcing Luca to defend and clarify the architectural dependencies.
Biggest teaching moment ▶ 1:13:08 Nathan explaining why RLVR fundamentally differs from RLHFNathan cleanly educates the host and audience on why verifiable rewards in RLVR avoid the fragile proxy pitfalls and emoji-gaming inherent to RLHF reward models.
Matt holds his own ▶ 1:18:12 Matt referencing Nathan's specific writing on the 'complexity tax'Matt demonstrates deep familiarity with Nathan's written work by bringing up his 'Thought on the Curve' essay and 'complexity tax' argument to steer the AGI debate.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Overview of the OLMo 3 Model Family Checkpoints | 2 | 5 | 1 | 0 | Matt opens by inviting the guests to present their OLMo 3 model family release. Luca and Nathan explain the model sizes, base model importance, and thinking capabilities relative to existing models like Qwen 2.5 and Llama 3.1. | |
| Pre-training Data Strategy and the Dolma 3 Dataset | 2 | 6 | 0 | 0 | Matt asks Luca to detail the Dolma 3 dataset. Luca explains pre-training token sampling, intelligent token repetition, and PDF crawling for long-context science data. | |
| Evaluating Base Model Benchmarks and Post-Training Magic | 2 | 6 | 1 | 0 | Matt prompts the guests to speak about model performance and efficiency. Luca and Nathan outline benchmark comparisons with Qwen 3 and NVIDIA Nemotron while noting the difficulty of base model evaluation. | |
| Defining True Open Source vs. Open Weights | 3 | 5 | 1 | 0 | Matt prompts a definition of open source versus open weights in AI. Luca clarifies that OLMo releases intermediate states, training data, and recipes rather than just final weight checkpoints. | |
| The Global Open Source AI Race: US vs. China | 5 | 6 | 2 | 1 | Matt asks for a recap of the global open source AI competition in 2025. Nathan provides detailed commentary on Chinese open source dominance (DeepSeek, Qwen, Kimi) following Meta leadership changes, and Matt references Martin Casado's quote on Qwen adoption. | |
| Enterprise Adoption, Economics, and US Strategic Alternatives | 4 | 6 | 1 | 0 | Matt questions why US ecosystem incentives favor closed APIs while China favors open source. Nathan outlines software monetization dynamics, enterprise constraints, and public US responses like the Atom project. | |
| Demystifying Thinking Models and Extended Inference | 3 | 5 | 2 | 0 | Matt asks the guests to define thinking models and extended inference. Nathan explains inference-time scaling, while Luca playfully expresses his preference for fast instruct models over slow thinking models. | |
| Guest Background: Luca Soldaini and AI2's Open Science Roots | 1 | 4 | 0 | 0 | Matt transitions to guest background stories. Luca recounts his academic career from Italy to Alexa and Semantic Scholar, and how AI2 started building open models with AMD compute. | |
| Guest Background: Nathan Lambert and Open Post-Training Research | 3 | 4 | 0 | 0 | Matt highlights Nathan's multi-faceted background including his Interconnects newsletter. Nathan traces his journey from Berkeley robotics to Hugging Face and joining AI2 to fill the public research communication vacuum. | |
| Inside AI2: Non-Profit Structure, Staff, and Key Projects | 5 | 5 | 0 | 0 | Matt displays specific knowledge of AI2's history and recent funding milestones including a $152M grant. Luca details AI2's organizational structure and active project streams across robotics, scientific agents, and climate. | |
| Deconstructing the 6-Stage OLMo 3 Training Pipeline | 5 | 6 | 1 | 1 | Nathan outlines the six stages of model training and brings up data contamination research. Matt demonstrates good grasp by reframing 'spurious rewards' into 'teaching to the test versus enabling true thinking.' | |
| Overview of Pre-Training versus Reinforcement Learning | 4 | 6 | 0 | 0 | Matt prompts an overview comparing pre-training against RL progress. Luca describes pre-training as an expensive initialization step, and Nathan emphasizes that base model quality dictates RL upside. | |
| Debating Richard Sutton's Perspective on RL and Pre-Training | 6 | 5 | 4 | 3 | Matt brings up Rich Sutton's recent podcast comments claiming pre-training is a flawed premise. Nathan forcefully dismisses Sutton's thesis as impractical theoretical nerd-sniping that ignores real-world LLM engineering realities. | |
| Stage 1: Pre-Training Data Selection and Execution | 2 | 6 | 0 | 0 | Luca details the mechanics of Stage 1 pre-training, including hardware constraints, execution timelines, and down-sampling 300 trillion candidate tokens into 6 trillion tokens. | |
| Stage 2: Mid-Training and Tail Patching Techniques | 7 | 6 | 2 | 4 | Luca describes mid-training tail patching. Matt directly challenges Luca on his past statement that 'data doesn't matter for long context,' prompting Luca to explain that architectural choices like QK norm dictate long context capability. | |
| Stage 4: Supervised Fine-Tuning and Model Distillation | 5 | 6 | 1 | 2 | Nathan breaks down supervised fine-tuning and model distillation from larger teachers like DeepSeek R1 and Qwen. Matt asks clarifying questions to help listeners distinguish SFT next-token prediction from reinforcement learning. | |
| Stage 5: Direct Preference Optimization and Preference Tuning | 3 | 6 | 1 | 0 | Nathan explains Direct Preference Optimization (DPO) and the Delta Learning Hypothesis. Luca quotes Dario Amodei on how simple 50-to-100-line post-training code implementations yield major gains when thoroughly validated. | |
| Stage 6: Reinforcement Learning with Verifiable Rewards | 5 | 7 | 1 | 1 | Nathan details Reinforcement Learning with Verifiable Rewards (RLVR). Matt asks for a plain-English definition contrasting RLVR with RLHF, leading Nathan to explain exact environment correctness versus subjective reward model proxies. | |
| The Reality of AGI, Physical Constraints, and Progress | 7 | 5 | 2 | 3 | Matt cites Nathan's recent essay 'Thought on the Curve' and concept of 'complexity tax' to probe the gap between AGI hype and practical engineering constraints. Nathan and Luca elaborate on physical power limits and continuous co-evolution. | |
| Preparing for the AI Future and Open-Source Scaffolding | 6 | 4 | 0 | 1 | Matt synthesizes the guests' views on AGI timelines into a clear takeaway. Luca and Nathan emphasize that open-source scaffolding outside frontier labs will drive the real societal impact of AI. |