Apr 3, 2025 · 1h 0m · mad
Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
On The MAD Podcast, host Matt Turck interviews François Chollet and Mike Knoop to explore the limits of LLM scaling, the release of the ARC-AGI-2 benchmark, breakthroughs in test-time reasoning like OpenAI's o3, and the founding of their new AGI research lab, NDEA.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 18.3% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
When Matt expresses slight skepticism with a brief 'Perhaps' regarding human performance on ARC tasks, Francois forcefully counters with 'No, definitely' and insists the eval set is straightforward for humans.
Hardest push from Matt ▶ 11:06 Host Challenges Guest's High Human Baseline ClaimMatt interjects with 'Perhaps' to challenge Francois's assertion that humans naturally score over 95% on ARC-1, compelling the guest to defend his statement.
Biggest teaching moment ▶ 17:30 Francois Corrects Baseline Benchmark FiguresFrancois gently corrects Matt's cited baseline performance stats, explaining that base LLMs actually scored around 10% on the semi-private evaluation set rather than 20-25%.
Matt holds his own ▶ 16:23 Host Demonstrates Knowledge of o3 Scores and Outreach DetailsMatt demonstrates high technical awareness by citing specific model scores across Claude 3.5 Sonnet and o3 while accurately recounting OpenAI's post-competition contact with the ARC team.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Episode Preview and Highlights | 1 | 3 | 1 | 0 | Matt sets up the episode with clips and asks a broad framing question about what ARC-AGI is and why Francois created it. Francois provides a detailed overview of fluid intelligence versus skill memorization. | |
| The Limitations of Pre-Training Scaling and Brute Force | 3 | 5 | 2 | 2 | Matt cites specific benchmark figures and asks about the limitations of brute-force LLMs. When Matt voices slight doubt ('Perhaps') about human ARC performance, Francois firmly corrects him ('No, definitely'), explaining task simplicity for humans. | |
| Intelligence as Skill Acquisition Efficiency | 2 | 4 | 0 | 0 | Matt asks a conceptual question regarding the definition of intelligence in light of scaling limitations. Francois responds by framing intelligence as skill acquisition efficiency rather than raw compute application. | |
| OpenAI o3 Breakthroughs and Test-Time Search Dynamics | 4 | 5 | 1 | 1 | Matt shows strong industry context by naming Claude 3.5 Sonnet scores and recounting OpenAI's outreach after ARC Prize 2024. Francois refines Matt's cited baseline numbers, clarifying that base LLMs scored at most around 10% on the private set. | |
| ARC Fine-Tuning and Program Synthesis Approaches | 4 | 4 | 0 | 2 | Matt presses Francois on whether OpenAI fine-tuned o3 directly on ARC training tasks. Francois breaks down the technical distinction between GitHub pre-training data contamination and targeted RL fine-tuning. | |
| Deep Learning-Guided Program Search Mechanics | 4 | 4 | 0 | 1 | Matt demonstrates technical literacy by introducing concepts like transductive models and program synthesis, then asks a sharp question clarifying search versus synthesis. Francois explains how human intuition restricts search space. | |
| Mike Knoop's AI Journey and ARC Prize Origins | 2 | 2 | 0 | 0 | Matt introduces Mike Knoop and asks about his background. Mike details his journey at Zapier, customer feedback on AI unreliability, and his initial meeting with Francois. | |
| ARC Prize Structure, Open Science, and Paper Awards | 3 | 3 | 1 | 1 | Matt asks about the competition structure, referring to specific participants like MindAI/Jack Cole. Mike and Francois explain open science rules and why top teams were disqualified for withholding open-source code. | |
| Open Source Imperatives and What is New in ARC Prize 2025 | 3 | 2 | 0 | 0 | Matt contrasts last year's closed-source OpenAI narrative with current open-weight breakthroughs like DeepSeek. Mike highlights how annual contest baselines help reset community research. | |
| Human Testing Baselines and Core Capabilities in ARC-AGI-2 | 4 | 3 | 0 | 0 | Matt quotes specific core capabilities from the ARC write-up such as compositional reasoning and contextual rule application. Francois details the human evaluation study in San Diego and explains panel voting math. | |
| NDEA: Building an Engine for Autonomous Innovation | 3 | 2 | 0 | 0 | Matt asks about NDEA, quoting its mission statement as a factory for rapid scientific advancement. Mike and Francois clarify that NDEA targets verifiable symbolic domains rather than everyday assistant tasks like email generation. | |
| Global Recruitment for NDEA and Episode Conclusion | 2 | 1 | 0 | 0 | Matt wraps up the interview with lighthearted praise for NDEA's recruitment pitch. Mike outlines their global remote strategy to recruit rare program synthesis engineers. |