Nov 20, 2025 · 1h 28m · mad

Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"

Nathan Lambert · 42m spoken Luca Soldaini · 27m spoken Matt Turck · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews AI2 researchers Nathan Lambert and Luca Soldaini about the launch of the fully open-source OLMo 3 model family. They explore the distinction between open weights and true open source, analyze the global AI ecosystem, and detail the complete six-stage architecture behind state-of-the-art reasoning models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 16.4% of the talking time here. How this is scored →

Matt as informed peer 4.0 Guest teaching 5.5 Guest disagreement 1.0 Matt pushing back 0.8
05100:0020:0040:001:00:001:20:001:18–5:48 · Matt as informed peer 2/10 Overview of the OLMo 3 Model Family Checkpoints Matt opens by inviting the guests to present their OLMo 3 model family release. Luca and Nathan explain the model sizes, base model importance, and thinking capabilities relative to existing models like Qwen 2.5 and Llama 3.1.5:48–8:07 · Matt as informed peer 2/10 Pre-training Data Strategy and the Dolma 3 Dataset Matt asks Luca to detail the Dolma 3 dataset. Luca explains pre-training token sampling, intelligent token repetition, and PDF crawling for long-context science data.8:07–10:28 · Matt as informed peer 2/10 Evaluating Base Model Benchmarks and Post-Training Magic Matt prompts the guests to speak about model performance and efficiency. Luca and Nathan outline benchmark comparisons with Qwen 3 and NVIDIA Nemotron while noting the difficulty of base model evaluation.10:28–12:51 · Matt as informed peer 3/10 Defining True Open Source vs. Open Weights Matt prompts a definition of open source versus open weights in AI. Luca clarifies that OLMo releases intermediate states, training data, and recipes rather than just final weight checkpoints.12:51–18:32 · Matt as informed peer 5/10 The Global Open Source AI Race: US vs. China Matt asks for a recap of the global open source AI competition in 2025. Nathan provides detailed commentary on Chinese open source dominance (DeepSeek, Qwen, Kimi) following Meta leadership changes, and Matt references Martin Casado's quote on Qwen adoption.18:32–22:13 · Matt as informed peer 4/10 Enterprise Adoption, Economics, and US Strategic Alternatives Matt questions why US ecosystem incentives favor closed APIs while China favors open source. Nathan outlines software monetization dynamics, enterprise constraints, and public US responses like the Atom project.22:13–24:16 · Matt as informed peer 3/10 Demystifying Thinking Models and Extended Inference Matt asks the guests to define thinking models and extended inference. Nathan explains inference-time scaling, while Luca playfully expresses his preference for fast instruct models over slow thinking models.24:16–27:30 · Matt as informed peer 1/10 Guest Background: Luca Soldaini and AI2's Open Science Roots Matt transitions to guest background stories. Luca recounts his academic career from Italy to Alexa and Semantic Scholar, and how AI2 started building open models with AMD compute.27:30–30:21 · Matt as informed peer 3/10 Guest Background: Nathan Lambert and Open Post-Training Research Matt highlights Nathan's multi-faceted background including his Interconnects newsletter. Nathan traces his journey from Berkeley robotics to Hugging Face and joining AI2 to fill the public research communication vacuum.30:21–35:58 · Matt as informed peer 5/10 Inside AI2: Non-Profit Structure, Staff, and Key Projects Matt displays specific knowledge of AI2's history and recent funding milestones including a $152M grant. Luca details AI2's organizational structure and active project streams across robotics, scientific agents, and climate.35:58–41:24 · Matt as informed peer 5/10 Deconstructing the 6-Stage OLMo 3 Training Pipeline Nathan outlines the six stages of model training and brings up data contamination research. Matt demonstrates good grasp by reframing 'spurious rewards' into 'teaching to the test versus enabling true thinking.'41:26–44:44 · Matt as informed peer 4/10 Overview of Pre-Training versus Reinforcement Learning Matt prompts an overview comparing pre-training against RL progress. Luca describes pre-training as an expensive initialization step, and Nathan emphasizes that base model quality dictates RL upside.44:44–46:51 · Matt as informed peer 6/10 Debating Richard Sutton's Perspective on RL and Pre-Training Matt brings up Rich Sutton's recent podcast comments claiming pre-training is a flawed premise. Nathan forcefully dismisses Sutton's thesis as impractical theoretical nerd-sniping that ignores real-world LLM engineering realities.46:51–50:26 · Matt as informed peer 2/10 Stage 1: Pre-Training Data Selection and Execution Luca details the mechanics of Stage 1 pre-training, including hardware constraints, execution timelines, and down-sampling 300 trillion candidate tokens into 6 trillion tokens.50:26–55:34 · Matt as informed peer 7/10 Stage 2: Mid-Training and Tail Patching Techniques Luca describes mid-training tail patching. Matt directly challenges Luca on his past statement that 'data doesn't matter for long context,' prompting Luca to explain that architectural choices like QK norm dictate long context capability.55:34–1:04:53 · Matt as informed peer 5/10 Stage 4: Supervised Fine-Tuning and Model Distillation Nathan breaks down supervised fine-tuning and model distillation from larger teachers like DeepSeek R1 and Qwen. Matt asks clarifying questions to help listeners distinguish SFT next-token prediction from reinforcement learning.1:04:53–1:10:51 · Matt as informed peer 3/10 Stage 5: Direct Preference Optimization and Preference Tuning Nathan explains Direct Preference Optimization (DPO) and the Delta Learning Hypothesis. Luca quotes Dario Amodei on how simple 50-to-100-line post-training code implementations yield major gains when thoroughly validated.1:10:51–1:18:12 · Matt as informed peer 5/10 Stage 6: Reinforcement Learning with Verifiable Rewards Nathan details Reinforcement Learning with Verifiable Rewards (RLVR). Matt asks for a plain-English definition contrasting RLVR with RLHF, leading Nathan to explain exact environment correctness versus subjective reward model proxies.1:18:12–1:23:51 · Matt as informed peer 7/10 The Reality of AGI, Physical Constraints, and Progress Matt cites Nathan's recent essay 'Thought on the Curve' and concept of 'complexity tax' to probe the gap between AGI hype and practical engineering constraints. Nathan and Luca elaborate on physical power limits and continuous co-evolution.1:23:51–1:27:50 · Matt as informed peer 6/10 Preparing for the AI Future and Open-Source Scaffolding Matt synthesizes the guests' views on AGI timelines into a clear takeaway. Luca and Nathan emphasize that open-source scaffolding outside frontier labs will drive the real societal impact of AI.1:18–5:48 · Guest teaching 5/10 Overview of the OLMo 3 Model Family Checkpoints Matt opens by inviting the guests to present their OLMo 3 model family release. Luca and Nathan explain the model sizes, base model importance, and thinking capabilities relative to existing models like Qwen 2.5 and Llama 3.1.5:48–8:07 · Guest teaching 6/10 Pre-training Data Strategy and the Dolma 3 Dataset Matt asks Luca to detail the Dolma 3 dataset. Luca explains pre-training token sampling, intelligent token repetition, and PDF crawling for long-context science data.8:07–10:28 · Guest teaching 6/10 Evaluating Base Model Benchmarks and Post-Training Magic Matt prompts the guests to speak about model performance and efficiency. Luca and Nathan outline benchmark comparisons with Qwen 3 and NVIDIA Nemotron while noting the difficulty of base model evaluation.10:28–12:51 · Guest teaching 5/10 Defining True Open Source vs. Open Weights Matt prompts a definition of open source versus open weights in AI. Luca clarifies that OLMo releases intermediate states, training data, and recipes rather than just final weight checkpoints.12:51–18:32 · Guest teaching 6/10 The Global Open Source AI Race: US vs. China Matt asks for a recap of the global open source AI competition in 2025. Nathan provides detailed commentary on Chinese open source dominance (DeepSeek, Qwen, Kimi) following Meta leadership changes, and Matt references Martin Casado's quote on Qwen adoption.18:32–22:13 · Guest teaching 6/10 Enterprise Adoption, Economics, and US Strategic Alternatives Matt questions why US ecosystem incentives favor closed APIs while China favors open source. Nathan outlines software monetization dynamics, enterprise constraints, and public US responses like the Atom project.22:13–24:16 · Guest teaching 5/10 Demystifying Thinking Models and Extended Inference Matt asks the guests to define thinking models and extended inference. Nathan explains inference-time scaling, while Luca playfully expresses his preference for fast instruct models over slow thinking models.24:16–27:30 · Guest teaching 4/10 Guest Background: Luca Soldaini and AI2's Open Science Roots Matt transitions to guest background stories. Luca recounts his academic career from Italy to Alexa and Semantic Scholar, and how AI2 started building open models with AMD compute.27:30–30:21 · Guest teaching 4/10 Guest Background: Nathan Lambert and Open Post-Training Research Matt highlights Nathan's multi-faceted background including his Interconnects newsletter. Nathan traces his journey from Berkeley robotics to Hugging Face and joining AI2 to fill the public research communication vacuum.30:21–35:58 · Guest teaching 5/10 Inside AI2: Non-Profit Structure, Staff, and Key Projects Matt displays specific knowledge of AI2's history and recent funding milestones including a $152M grant. Luca details AI2's organizational structure and active project streams across robotics, scientific agents, and climate.35:58–41:24 · Guest teaching 6/10 Deconstructing the 6-Stage OLMo 3 Training Pipeline Nathan outlines the six stages of model training and brings up data contamination research. Matt demonstrates good grasp by reframing 'spurious rewards' into 'teaching to the test versus enabling true thinking.'41:26–44:44 · Guest teaching 6/10 Overview of Pre-Training versus Reinforcement Learning Matt prompts an overview comparing pre-training against RL progress. Luca describes pre-training as an expensive initialization step, and Nathan emphasizes that base model quality dictates RL upside.44:44–46:51 · Guest teaching 5/10 Debating Richard Sutton's Perspective on RL and Pre-Training Matt brings up Rich Sutton's recent podcast comments claiming pre-training is a flawed premise. Nathan forcefully dismisses Sutton's thesis as impractical theoretical nerd-sniping that ignores real-world LLM engineering realities.46:51–50:26 · Guest teaching 6/10 Stage 1: Pre-Training Data Selection and Execution Luca details the mechanics of Stage 1 pre-training, including hardware constraints, execution timelines, and down-sampling 300 trillion candidate tokens into 6 trillion tokens.50:26–55:34 · Guest teaching 6/10 Stage 2: Mid-Training and Tail Patching Techniques Luca describes mid-training tail patching. Matt directly challenges Luca on his past statement that 'data doesn't matter for long context,' prompting Luca to explain that architectural choices like QK norm dictate long context capability.55:34–1:04:53 · Guest teaching 6/10 Stage 4: Supervised Fine-Tuning and Model Distillation Nathan breaks down supervised fine-tuning and model distillation from larger teachers like DeepSeek R1 and Qwen. Matt asks clarifying questions to help listeners distinguish SFT next-token prediction from reinforcement learning.1:04:53–1:10:51 · Guest teaching 6/10 Stage 5: Direct Preference Optimization and Preference Tuning Nathan explains Direct Preference Optimization (DPO) and the Delta Learning Hypothesis. Luca quotes Dario Amodei on how simple 50-to-100-line post-training code implementations yield major gains when thoroughly validated.1:10:51–1:18:12 · Guest teaching 7/10 Stage 6: Reinforcement Learning with Verifiable Rewards Nathan details Reinforcement Learning with Verifiable Rewards (RLVR). Matt asks for a plain-English definition contrasting RLVR with RLHF, leading Nathan to explain exact environment correctness versus subjective reward model proxies.1:18:12–1:23:51 · Guest teaching 5/10 The Reality of AGI, Physical Constraints, and Progress Matt cites Nathan's recent essay 'Thought on the Curve' and concept of 'complexity tax' to probe the gap between AGI hype and practical engineering constraints. Nathan and Luca elaborate on physical power limits and continuous co-evolution.1:23:51–1:27:50 · Guest teaching 4/10 Preparing for the AI Future and Open-Source Scaffolding Matt synthesizes the guests' views on AGI timelines into a clear takeaway. Luca and Nathan emphasize that open-source scaffolding outside frontier labs will drive the real societal impact of AI.1:18–5:48 · Guest disagreement 1/10 Overview of the OLMo 3 Model Family Checkpoints Matt opens by inviting the guests to present their OLMo 3 model family release. Luca and Nathan explain the model sizes, base model importance, and thinking capabilities relative to existing models like Qwen 2.5 and Llama 3.1.5:48–8:07 · Guest disagreement 0/10 Pre-training Data Strategy and the Dolma 3 Dataset Matt asks Luca to detail the Dolma 3 dataset. Luca explains pre-training token sampling, intelligent token repetition, and PDF crawling for long-context science data.8:07–10:28 · Guest disagreement 1/10 Evaluating Base Model Benchmarks and Post-Training Magic Matt prompts the guests to speak about model performance and efficiency. Luca and Nathan outline benchmark comparisons with Qwen 3 and NVIDIA Nemotron while noting the difficulty of base model evaluation.10:28–12:51 · Guest disagreement 1/10 Defining True Open Source vs. Open Weights Matt prompts a definition of open source versus open weights in AI. Luca clarifies that OLMo releases intermediate states, training data, and recipes rather than just final weight checkpoints.12:51–18:32 · Guest disagreement 2/10 The Global Open Source AI Race: US vs. China Matt asks for a recap of the global open source AI competition in 2025. Nathan provides detailed commentary on Chinese open source dominance (DeepSeek, Qwen, Kimi) following Meta leadership changes, and Matt references Martin Casado's quote on Qwen adoption.18:32–22:13 · Guest disagreement 1/10 Enterprise Adoption, Economics, and US Strategic Alternatives Matt questions why US ecosystem incentives favor closed APIs while China favors open source. Nathan outlines software monetization dynamics, enterprise constraints, and public US responses like the Atom project.22:13–24:16 · Guest disagreement 2/10 Demystifying Thinking Models and Extended Inference Matt asks the guests to define thinking models and extended inference. Nathan explains inference-time scaling, while Luca playfully expresses his preference for fast instruct models over slow thinking models.24:16–27:30 · Guest disagreement 0/10 Guest Background: Luca Soldaini and AI2's Open Science Roots Matt transitions to guest background stories. Luca recounts his academic career from Italy to Alexa and Semantic Scholar, and how AI2 started building open models with AMD compute.27:30–30:21 · Guest disagreement 0/10 Guest Background: Nathan Lambert and Open Post-Training Research Matt highlights Nathan's multi-faceted background including his Interconnects newsletter. Nathan traces his journey from Berkeley robotics to Hugging Face and joining AI2 to fill the public research communication vacuum.30:21–35:58 · Guest disagreement 0/10 Inside AI2: Non-Profit Structure, Staff, and Key Projects Matt displays specific knowledge of AI2's history and recent funding milestones including a $152M grant. Luca details AI2's organizational structure and active project streams across robotics, scientific agents, and climate.35:58–41:24 · Guest disagreement 1/10 Deconstructing the 6-Stage OLMo 3 Training Pipeline Nathan outlines the six stages of model training and brings up data contamination research. Matt demonstrates good grasp by reframing 'spurious rewards' into 'teaching to the test versus enabling true thinking.'41:26–44:44 · Guest disagreement 0/10 Overview of Pre-Training versus Reinforcement Learning Matt prompts an overview comparing pre-training against RL progress. Luca describes pre-training as an expensive initialization step, and Nathan emphasizes that base model quality dictates RL upside.44:44–46:51 · Guest disagreement 4/10 Debating Richard Sutton's Perspective on RL and Pre-Training Matt brings up Rich Sutton's recent podcast comments claiming pre-training is a flawed premise. Nathan forcefully dismisses Sutton's thesis as impractical theoretical nerd-sniping that ignores real-world LLM engineering realities.46:51–50:26 · Guest disagreement 0/10 Stage 1: Pre-Training Data Selection and Execution Luca details the mechanics of Stage 1 pre-training, including hardware constraints, execution timelines, and down-sampling 300 trillion candidate tokens into 6 trillion tokens.50:26–55:34 · Guest disagreement 2/10 Stage 2: Mid-Training and Tail Patching Techniques Luca describes mid-training tail patching. Matt directly challenges Luca on his past statement that 'data doesn't matter for long context,' prompting Luca to explain that architectural choices like QK norm dictate long context capability.55:34–1:04:53 · Guest disagreement 1/10 Stage 4: Supervised Fine-Tuning and Model Distillation Nathan breaks down supervised fine-tuning and model distillation from larger teachers like DeepSeek R1 and Qwen. Matt asks clarifying questions to help listeners distinguish SFT next-token prediction from reinforcement learning.1:04:53–1:10:51 · Guest disagreement 1/10 Stage 5: Direct Preference Optimization and Preference Tuning Nathan explains Direct Preference Optimization (DPO) and the Delta Learning Hypothesis. Luca quotes Dario Amodei on how simple 50-to-100-line post-training code implementations yield major gains when thoroughly validated.1:10:51–1:18:12 · Guest disagreement 1/10 Stage 6: Reinforcement Learning with Verifiable Rewards Nathan details Reinforcement Learning with Verifiable Rewards (RLVR). Matt asks for a plain-English definition contrasting RLVR with RLHF, leading Nathan to explain exact environment correctness versus subjective reward model proxies.1:18:12–1:23:51 · Guest disagreement 2/10 The Reality of AGI, Physical Constraints, and Progress Matt cites Nathan's recent essay 'Thought on the Curve' and concept of 'complexity tax' to probe the gap between AGI hype and practical engineering constraints. Nathan and Luca elaborate on physical power limits and continuous co-evolution.1:23:51–1:27:50 · Guest disagreement 0/10 Preparing for the AI Future and Open-Source Scaffolding Matt synthesizes the guests' views on AGI timelines into a clear takeaway. Luca and Nathan emphasize that open-source scaffolding outside frontier labs will drive the real societal impact of AI.1:18–5:48 · Matt pushing back 0/10 Overview of the OLMo 3 Model Family Checkpoints Matt opens by inviting the guests to present their OLMo 3 model family release. Luca and Nathan explain the model sizes, base model importance, and thinking capabilities relative to existing models like Qwen 2.5 and Llama 3.1.5:48–8:07 · Matt pushing back 0/10 Pre-training Data Strategy and the Dolma 3 Dataset Matt asks Luca to detail the Dolma 3 dataset. Luca explains pre-training token sampling, intelligent token repetition, and PDF crawling for long-context science data.8:07–10:28 · Matt pushing back 0/10 Evaluating Base Model Benchmarks and Post-Training Magic Matt prompts the guests to speak about model performance and efficiency. Luca and Nathan outline benchmark comparisons with Qwen 3 and NVIDIA Nemotron while noting the difficulty of base model evaluation.10:28–12:51 · Matt pushing back 0/10 Defining True Open Source vs. Open Weights Matt prompts a definition of open source versus open weights in AI. Luca clarifies that OLMo releases intermediate states, training data, and recipes rather than just final weight checkpoints.12:51–18:32 · Matt pushing back 1/10 The Global Open Source AI Race: US vs. China Matt asks for a recap of the global open source AI competition in 2025. Nathan provides detailed commentary on Chinese open source dominance (DeepSeek, Qwen, Kimi) following Meta leadership changes, and Matt references Martin Casado's quote on Qwen adoption.18:32–22:13 · Matt pushing back 0/10 Enterprise Adoption, Economics, and US Strategic Alternatives Matt questions why US ecosystem incentives favor closed APIs while China favors open source. Nathan outlines software monetization dynamics, enterprise constraints, and public US responses like the Atom project.22:13–24:16 · Matt pushing back 0/10 Demystifying Thinking Models and Extended Inference Matt asks the guests to define thinking models and extended inference. Nathan explains inference-time scaling, while Luca playfully expresses his preference for fast instruct models over slow thinking models.24:16–27:30 · Matt pushing back 0/10 Guest Background: Luca Soldaini and AI2's Open Science Roots Matt transitions to guest background stories. Luca recounts his academic career from Italy to Alexa and Semantic Scholar, and how AI2 started building open models with AMD compute.27:30–30:21 · Matt pushing back 0/10 Guest Background: Nathan Lambert and Open Post-Training Research Matt highlights Nathan's multi-faceted background including his Interconnects newsletter. Nathan traces his journey from Berkeley robotics to Hugging Face and joining AI2 to fill the public research communication vacuum.30:21–35:58 · Matt pushing back 0/10 Inside AI2: Non-Profit Structure, Staff, and Key Projects Matt displays specific knowledge of AI2's history and recent funding milestones including a $152M grant. Luca details AI2's organizational structure and active project streams across robotics, scientific agents, and climate.35:58–41:24 · Matt pushing back 1/10 Deconstructing the 6-Stage OLMo 3 Training Pipeline Nathan outlines the six stages of model training and brings up data contamination research. Matt demonstrates good grasp by reframing 'spurious rewards' into 'teaching to the test versus enabling true thinking.'41:26–44:44 · Matt pushing back 0/10 Overview of Pre-Training versus Reinforcement Learning Matt prompts an overview comparing pre-training against RL progress. Luca describes pre-training as an expensive initialization step, and Nathan emphasizes that base model quality dictates RL upside.44:44–46:51 · Matt pushing back 3/10 Debating Richard Sutton's Perspective on RL and Pre-Training Matt brings up Rich Sutton's recent podcast comments claiming pre-training is a flawed premise. Nathan forcefully dismisses Sutton's thesis as impractical theoretical nerd-sniping that ignores real-world LLM engineering realities.46:51–50:26 · Matt pushing back 0/10 Stage 1: Pre-Training Data Selection and Execution Luca details the mechanics of Stage 1 pre-training, including hardware constraints, execution timelines, and down-sampling 300 trillion candidate tokens into 6 trillion tokens.50:26–55:34 · Matt pushing back 4/10 Stage 2: Mid-Training and Tail Patching Techniques Luca describes mid-training tail patching. Matt directly challenges Luca on his past statement that 'data doesn't matter for long context,' prompting Luca to explain that architectural choices like QK norm dictate long context capability.55:34–1:04:53 · Matt pushing back 2/10 Stage 4: Supervised Fine-Tuning and Model Distillation Nathan breaks down supervised fine-tuning and model distillation from larger teachers like DeepSeek R1 and Qwen. Matt asks clarifying questions to help listeners distinguish SFT next-token prediction from reinforcement learning.1:04:53–1:10:51 · Matt pushing back 0/10 Stage 5: Direct Preference Optimization and Preference Tuning Nathan explains Direct Preference Optimization (DPO) and the Delta Learning Hypothesis. Luca quotes Dario Amodei on how simple 50-to-100-line post-training code implementations yield major gains when thoroughly validated.1:10:51–1:18:12 · Matt pushing back 1/10 Stage 6: Reinforcement Learning with Verifiable Rewards Nathan details Reinforcement Learning with Verifiable Rewards (RLVR). Matt asks for a plain-English definition contrasting RLVR with RLHF, leading Nathan to explain exact environment correctness versus subjective reward model proxies.1:18:12–1:23:51 · Matt pushing back 3/10 The Reality of AGI, Physical Constraints, and Progress Matt cites Nathan's recent essay 'Thought on the Curve' and concept of 'complexity tax' to probe the gap between AGI hype and practical engineering constraints. Nathan and Luca elaborate on physical power limits and continuous co-evolution.1:23:51–1:27:50 · Matt pushing back 1/10 Preparing for the AI Future and Open-Source Scaffolding Matt synthesizes the guests' views on AGI timelines into a clear takeaway. Luca and Nathan emphasize that open-source scaffolding outside frontier labs will drive the real societal impact of AI.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 30.7% · guest 69.3%0:00 · Matt 30.7% · guest 69.3%3:00 · Matt 2.7% · guest 97.3%3:00 · Matt 2.7% · guest 97.3%6:00 · Matt 3.4% · guest 96.6%6:00 · Matt 3.4% · guest 96.6%9:00 · Matt 5.6% · guest 94.4%9:00 · Matt 5.6% · guest 94.4%12:00 · Matt 15.3% · guest 84.7%12:00 · Matt 15.3% · guest 84.7%15:00 · Matt 9.8% · guest 90.2%15:00 · Matt 9.8% · guest 90.2%18:00 · Matt 12.2% · guest 87.8%18:00 · Matt 12.2% · guest 87.8%21:00 · Matt 17% · guest 83%21:00 · Matt 17% · guest 83%24:00 · Matt 19% · guest 81%24:00 · Matt 19% · guest 81%27:00 · Matt 10.8% · guest 89.2%27:00 · Matt 10.8% · guest 89.2%30:00 · Matt 28.9% · guest 71.1%30:00 · Matt 28.9% · guest 71.1%33:00 · Matt 13.3% · guest 86.7%33:00 · Matt 13.3% · guest 86.7%36:00 · Matt 45.2% · guest 54.8%36:00 · Matt 45.2% · guest 54.8%39:00 · Matt 22.9% · guest 77.1%39:00 · Matt 22.9% · guest 77.1%42:00 · Matt 13.8% · guest 86.2%42:00 · Matt 13.8% · guest 86.2%45:00 · Matt 10.5% · guest 89.5%45:00 · Matt 10.5% · guest 89.5%48:00 · Matt 6.7% · guest 93.3%48:00 · Matt 6.7% · guest 93.3%51:00 · Matt 13.5% · guest 86.5%51:00 · Matt 13.5% · guest 86.5%54:00 · Matt 21.8% · guest 78.2%54:00 · Matt 21.8% · guest 78.2%57:00 · Matt 7% · guest 93%57:00 · Matt 7% · guest 93%1:00:00 · Matt 15.2% · guest 84.8%1:00:00 · Matt 15.2% · guest 84.8%1:03:00 · Matt 12.2% · guest 87.8%1:03:00 · Matt 12.2% · guest 87.8%1:06:00 · Matt 0% · guest 100%1:06:00 · Matt 0% · guest 100%1:09:00 · Matt 15% · guest 85%1:09:00 · Matt 15% · guest 85%1:12:00 · Matt 11.2% · guest 88.8%1:12:00 · Matt 11.2% · guest 88.8%1:15:00 · Matt 0% · guest 100%1:15:00 · Matt 0% · guest 100%1:18:00 · Matt 60.7% · guest 39.3%1:18:00 · Matt 60.7% · guest 39.3%1:21:00 · Matt 16.9% · guest 83.1%1:21:00 · Matt 16.9% · guest 83.1%1:24:00 · Matt 15.5% · guest 84.5%1:24:00 · Matt 15.5% · guest 84.5%1:27:00 · Matt 63% · guest 37%1:27:00 · Matt 63% · guest 37%
Sharpest disagreement ▶ 45:10 Nathan forcefully rejecting Rich Sutton's RL thesis

Nathan forcefully rejects Rich Sutton's podcast claim that pre-training is flawed, dismissing it as theoretical nerd-sniping that ignores practical engineering realities.

Hardest push from Matt ▶ 53:10 Matt holding Luca accountable to his past claim about data in long context

Matt directly challenges Luca by citing his prior statement that 'data doesn't matter for long context,' forcing Luca to defend and clarify the architectural dependencies.

Biggest teaching moment ▶ 1:13:08 Nathan explaining why RLVR fundamentally differs from RLHF

Nathan cleanly educates the host and audience on why verifiable rewards in RLVR avoid the fragile proxy pitfalls and emoji-gaming inherent to RLHF reward models.

Matt holds his own ▶ 1:18:12 Matt referencing Nathan's specific writing on the 'complexity tax'

Matt demonstrates deep familiarity with Nathan's written work by bringing up his 'Thought on the Curve' essay and 'complexity tax' argument to steer the AGI debate.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Overview of the OLMo 3 Model Family Checkpoints 2510 Matt opens by inviting the guests to present their OLMo 3 model family release. Luca and Nathan explain the model sizes, base model importance, and thinking capabilities relative to existing models like Qwen 2.5 and Llama 3.1.
Pre-training Data Strategy and the Dolma 3 Dataset 2600 Matt asks Luca to detail the Dolma 3 dataset. Luca explains pre-training token sampling, intelligent token repetition, and PDF crawling for long-context science data.
Evaluating Base Model Benchmarks and Post-Training Magic 2610 Matt prompts the guests to speak about model performance and efficiency. Luca and Nathan outline benchmark comparisons with Qwen 3 and NVIDIA Nemotron while noting the difficulty of base model evaluation.
Defining True Open Source vs. Open Weights 3510 Matt prompts a definition of open source versus open weights in AI. Luca clarifies that OLMo releases intermediate states, training data, and recipes rather than just final weight checkpoints.
The Global Open Source AI Race: US vs. China 5621 Matt asks for a recap of the global open source AI competition in 2025. Nathan provides detailed commentary on Chinese open source dominance (DeepSeek, Qwen, Kimi) following Meta leadership changes, and Matt references Martin Casado's quote on Qwen adoption.
Enterprise Adoption, Economics, and US Strategic Alternatives 4610 Matt questions why US ecosystem incentives favor closed APIs while China favors open source. Nathan outlines software monetization dynamics, enterprise constraints, and public US responses like the Atom project.
Demystifying Thinking Models and Extended Inference 3520 Matt asks the guests to define thinking models and extended inference. Nathan explains inference-time scaling, while Luca playfully expresses his preference for fast instruct models over slow thinking models.
Guest Background: Luca Soldaini and AI2's Open Science Roots 1400 Matt transitions to guest background stories. Luca recounts his academic career from Italy to Alexa and Semantic Scholar, and how AI2 started building open models with AMD compute.
Guest Background: Nathan Lambert and Open Post-Training Research 3400 Matt highlights Nathan's multi-faceted background including his Interconnects newsletter. Nathan traces his journey from Berkeley robotics to Hugging Face and joining AI2 to fill the public research communication vacuum.
Inside AI2: Non-Profit Structure, Staff, and Key Projects 5500 Matt displays specific knowledge of AI2's history and recent funding milestones including a $152M grant. Luca details AI2's organizational structure and active project streams across robotics, scientific agents, and climate.
Deconstructing the 6-Stage OLMo 3 Training Pipeline 5611 Nathan outlines the six stages of model training and brings up data contamination research. Matt demonstrates good grasp by reframing 'spurious rewards' into 'teaching to the test versus enabling true thinking.'
Overview of Pre-Training versus Reinforcement Learning 4600 Matt prompts an overview comparing pre-training against RL progress. Luca describes pre-training as an expensive initialization step, and Nathan emphasizes that base model quality dictates RL upside.
Debating Richard Sutton's Perspective on RL and Pre-Training 6543 Matt brings up Rich Sutton's recent podcast comments claiming pre-training is a flawed premise. Nathan forcefully dismisses Sutton's thesis as impractical theoretical nerd-sniping that ignores real-world LLM engineering realities.
Stage 1: Pre-Training Data Selection and Execution 2600 Luca details the mechanics of Stage 1 pre-training, including hardware constraints, execution timelines, and down-sampling 300 trillion candidate tokens into 6 trillion tokens.
Stage 2: Mid-Training and Tail Patching Techniques 7624 Luca describes mid-training tail patching. Matt directly challenges Luca on his past statement that 'data doesn't matter for long context,' prompting Luca to explain that architectural choices like QK norm dictate long context capability.
Stage 4: Supervised Fine-Tuning and Model Distillation 5612 Nathan breaks down supervised fine-tuning and model distillation from larger teachers like DeepSeek R1 and Qwen. Matt asks clarifying questions to help listeners distinguish SFT next-token prediction from reinforcement learning.
Stage 5: Direct Preference Optimization and Preference Tuning 3610 Nathan explains Direct Preference Optimization (DPO) and the Delta Learning Hypothesis. Luca quotes Dario Amodei on how simple 50-to-100-line post-training code implementations yield major gains when thoroughly validated.
Stage 6: Reinforcement Learning with Verifiable Rewards 5711 Nathan details Reinforcement Learning with Verifiable Rewards (RLVR). Matt asks for a plain-English definition contrasting RLVR with RLHF, leading Nathan to explain exact environment correctness versus subjective reward model proxies.
The Reality of AGI, Physical Constraints, and Progress 7523 Matt cites Nathan's recent essay 'Thought on the Curve' and concept of 'complexity tax' to probe the gap between AGI hype and practical engineering constraints. Nathan and Luca elaborate on physical power limits and continuous co-evolution.
Preparing for the AI Future and Open-Source Scaffolding 6401 Matt synthesizes the guests' views on AGI timelines into a clear takeaway. Luca and Nathan emphasize that open-source scaffolding outside frontier labs will drive the real societal impact of AI.

Statements from this episode (34)

Disclosure
Ai2 releases OLMo 3 with full training recipes, data, and intermediate checkpoints
“We're not just releasing the final models. We're releasing, you know, the entire recipe we followed to get this model. So the data, the intermediate states, the evaluation frameworks, all the details, all the bits that people need to know to make models like O…”
Luca Soldaini Nov 20, 2025 ▶ 1:46
Assertion Not checkable as stated
Lambert: OLMo 3 32B base model matches Qwen 2.5 32B quality
“This base model is similar in quality to the best available, which is like Quinn's 2.5, 32 B is, was still the best base model.”
Nathan Lambert Nov 20, 2025 ▶ 4:06
Assertion Not checkable as stated
Lambert: OLMo 3 7B outperforms Meta's Llama 3.1 8B in internal tests
“And I just think of this cause like Lama 3.1 AP is one of the most used models and hugging base of all time. And this should be better. We're, In our measurements, we see it as being better than Llama.”
Nathan Lambert Nov 20, 2025 ▶ 4:54
Disclosure
Ai2 samples 6T tokens from 10T pool for OLMo 3
“There's like a pool of about 10 trillion tokens from which we have like an algorithm also fully open source. To like sample about six trillion tokens that we use during training.”
Luca Soldaini Nov 20, 2025 ▶ 6:07
Assertion Not checkable as stated
Soldaini: 95% of web pages are under 3,000 tokens
“Like 95% web pages are below 3000 tokens.”
Luca Soldaini Nov 20, 2025 ▶ 7:31
Assertion Supported
Lambert: OLMo 3 models are the best open models outside Qwen 3
“I would say in post training where The best models that don't start with Quinn three and we're like reasonable to say that they are comparable to Quinn three, like on some benchmarks would beat them on some benchmarks. They're way ahead.”
Nathan Lambert Nov 20, 2025 ▶ 8:59
Assertion Supported
Lambert: Alibaba's Qwen 3 VL vision model is a superior text model
“They released these Quinn three VL, their vision models. And like on text only benchmarks, it's way better than the models they released in April. So it's like okay, like that's the new baseline. And most people don't know about it because they think it's just…”
Nathan Lambert Nov 20, 2025 ▶ 9:46
Assertion Not checkable as stated
Soldaini: Most open AI models are open weights, not open source
“Majority of models that get release I think the best term to describe them is open weights. Your Quinn, your Gemma, your Lama you know, Kimi it's what gets release is a set of weights that correspond either to the final state of model, that's the most common, …”
Luca Soldaini Nov 20, 2025 ▶ 10:52
Prediction Not checkable as stated
Lambert predicts more US labs will release open AI models
“If you look at this podcast in the coming months, I do think there's going to be, look like there's a lot more labs in the U S participating.”
Nathan Lambert Nov 20, 2025 ▶ 16:22
Assertion Partly supported
Lambert: 80% of a16z's open-model portfolio startups use Alibaba's Qwen
“80% of companies building with open models are using Quinn, which is like 16 to 24% of his portfolio, which is still a lot.”
Nathan Lambert Nov 20, 2025 ▶ 17:01
Assertion Not checkable as stated
Lambert: Chinese open AI models currently do not contain backdoors
“Like, you can't prove that the models aren't doing certain backdoors, where I'm fairly certain they definitely aren't now.”
Nathan Lambert Nov 20, 2025 ▶ 17:52
Assertion Not checkable as stated
Lambert: Chinese companies with $1B+ valuations routinely pirate SaaS software
“Mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software.”
Nathan Lambert Nov 20, 2025 ▶ 18:47
Assertion Open · timeframe Nov 2028
Ai2 received an initial grant of two million GPU hours from AMD
“We got an initial grant from AMD at the time. There was about two million GPU hours.”
Luca Soldaini Nov 20, 2025 ▶ 26:58
Disclosure
Lambert: AI2 coined 'reinforcement learning with verifiable rewards' replicating Llama 3
“We spent a long time to try to replicate what we thought was close to Lama three post training with multiple stages and optimizers, which is the project that like came up with the name reinforcement learning with verifiable rewards with a bunch of people.”
Nathan Lambert Nov 20, 2025 ▶ 29:22
Assertion Not checkable as stated
Lambert: As AI funding grows, fewer researchers speak in public
“There's so much money in AI and it only becomes increasingly so that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller.”
Nathan Lambert Nov 20, 2025 ▶ 29:52
Assertion Supported
Lambert: Hugging Face outcompeted AI2's AllenNLP library
“It was the main competitor to Hugging Face Transformers. And they ultimately outcompeted AI two as the thing that people use for that because they had very different model and amount of support.”
Nathan Lambert Nov 20, 2025 ▶ 32:47
Assertion Not checkable as stated
Lambert: Long-context extension is essential for reasoning AI models
“Three is long context extension, which is absolutely essential for these reasoning models because they generate so many intermediate tokens before sharing an answer with you.”
Nathan Lambert Nov 20, 2025 ▶ 40:15
What-if
Lambert: Scaling AI 10x alters post-training, not pre-training methods
“If like, if we were to train a model that was 10 times as big, like all this post-training stuff would change. But the pre-training And mid training and long contacts, I think would actually become looking pretty similar.”
Nathan Lambert Nov 20, 2025 ▶ 40:55
Assertion Not checkable as stated
Lambert: Larger pre-trained base models are easier to improve with RL
“A better base model and a bigger base model is much easier to improve with RL.”
Nathan Lambert Nov 20, 2025 ▶ 44:29
Opinion
Lambert: Rich Sutton's RL theories are impractical for models like GPT-6
“Rich is a font of wonderful ideas, but Often not ones that are going to be immediately practical. This is how you get things like creating reinforcement learning, but not necessarily things that are going to impact what GPT six is.”
Nathan Lambert Nov 20, 2025 ▶ 45:11
Assertion Not checkable as stated
Soldaini: Frontier AI labs limit final pre-training runs to two months
“I think it's standard practice among the frontier labs to try to cap your big final pre-training run to two months not more than that.”
Luca Soldaini Nov 20, 2025 ▶ 47:21
Disclosure
Ai2 filtered OLMo 3's pre-training dataset from 300 trillion tokens
“Our initial pool was closer to 300 trillion tokens. You shrink it down till you reach your target number, and hopefully as you shrink, you only keep the best part of this.”
Luca Soldaini Nov 20, 2025 ▶ 48:44
Insight
Soldaini: Mid-training requires re-mixing pre-training data to avoid model forgetting
“When you do that, you also need to make sure that The model doesn't forget stuff that I've seen during pre-training, so that's why, like, you mix some of the best data from pre-training, you do carry over.”
Luca Soldaini Nov 20, 2025 ▶ 50:58
Assertion Supported
Soldaini: Training LLMs on longer sequences causes quadratic compute slowdown
“It's because the longer the input that a model is trained on, the slower it is. The rate at which it gets slower, it's higher than the length of a context. It's a quadratic slowdown.”
Luca Soldaini Nov 20, 2025 ▶ 52:41
Insight
Soldaini: Flawed long-context model architecture cannot be saved by good data
“But they're like technical decisions in how you set up your model that you can have the best data in the world. And your model will not be able to reason over many, many tokens. So it doesn't matter in the sense that you can't train the model on bad data, but …”
Luca Soldaini Nov 20, 2025 ▶ 53:40
Disclosure
Ai2 fine-tuned OLMo 3 using Chinese teacher models DeepSeek-R1 and Qwen
“So in our case, we took a mix of existing data sets like Open Thoughts three and modified it, which is from Bespoke AI labs, a startup. And then we also generated a whole bunch of new data. So we ended up using a mix of teachers from like Deep Seek R one, oh f…”
Nathan Lambert Nov 20, 2025 ▶ 57:15
Assertion Not checkable as stated
Lambert: Best open-license AI models near the frontier in 2025 were Chinese
“The models that are from closest to the frontier in performance with good license all happened to be Chinese models throughout the year for this case.”
Nathan Lambert Nov 20, 2025 ▶ 58:29
Disclosure
Lambert: Ai2 generated billions of DeepSeek completions over a weekend
“We had a bunch of cloud credits and I, they were running out and we're behind and I just generated like as many completions as possible. So it was like a few billion completions from deep seek over the weekend.”
Nathan Lambert Nov 20, 2025 ▶ 1:08:18
Insight
Lambert: RLVR targets performance characteristics better than traditional RLHF reward models
“These reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about where RLVR is much better matc…”
Nathan Lambert Nov 20, 2025 ▶ 1:13:35
Assertion Supported
Lambert: Kernel differences between vLLM and Hugging Face cause RL numerical instability
“VLLM and HuggingFace use different kernels to do the actual internal computation of the model. So these kernels are the things that make things like vLLM really fast. But these things, this then results in subtle numerical differences between the completions t…”
Nathan Lambert Nov 20, 2025 ▶ 1:15:29
Assertion Not checkable as stated
Lambert: Most AI labs probably use evolved GRPO rather than PPO
“In reality, it seems like most people are using something like an evolved version of GRPO, which is a bit simpler than PPO.”
Nathan Lambert Nov 20, 2025 ▶ 1:16:39
Prediction Not checkable as stated
Lambert: AI progress will yield steady improvements rather than rapid singularity
“I think these researchers are going to grind out improvements for multiple years, but never in a way that results in this kind of accelerating well that we get drawn into.”
Nathan Lambert Nov 20, 2025 ▶ 1:21:34
Prediction Not checkable as stated
Lambert: Big tech will realize 95-98% of LLM potential by 2030
“I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech E…”
Nathan Lambert Nov 20, 2025 ▶ 1:24:19
Insight
Soldaini: AI scaffolding allows people outside frontier labs to drive capabilities
“If the scaffolding is what really moves a lot of like from, you know, broad capability model to like something that actually has meaningful impact, that scaffolding is not just like, oh, only the labs of people are trained models can do it. Like the number of …”
Luca Soldaini Nov 20, 2025 ▶ 1:26:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.