Mar 27, 2026 · 57m · y-combinator
François Chollet: Why Scaling Alone Isn’t Enough for AGI · Y Combinator
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Y Combinator podcast episode, François Chollet discusses why scaling traditional large language models alone is insufficient for reaching AGI, detailing his work on symbolic program synthesis at NDEA and the evolution of the ARC-AGI benchmark. He explains how moving beyond parametric deep learning toward verifiable reasoning, interactive environments, and sample-efficient architectures will pave the way for true fluid intelligence.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The partners hold 16.2% of the talking time here. How this is scored →
speaking balance: gold is the partners, purple is the guest (3 minute bins)
Chollet firmly dismisses the widely accepted definition of AGI as economic task automation, arguing it reflects task automation rather than genuine general intelligence.
Hardest push from the partners ▶ 7:11 Diana Hu pushes back on non-verifiable domainsDiana Hu challenges the universality of verifiable reward loops by asking how fuzzy, subjective tasks like essay writing can realistically be transformed into formal verification functions.
Biggest teaching moment ▶ 14:05 Explaining the fundamental limit of gradient descentChollet walks through his Google Brain research showing that gradient descent is mathematically incapable of discovering generalizable symbolic programs and defaults to overfit pattern matching.
The partners hold their own ▶ 23:57 Diana Hu shares YC batch benchmark resultsDiana Hu demonstrates YC's direct front-row exposure to cutting-edge AI developments by detailing how Confluence Labs saturated ARC-AGI-2 to 97% using custom harnesses during the W26 batch.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The partners as informed peer | Guest teaching | Guest disagreement | The partners pushing back | Why |
|---|---|---|---|---|---|---|
| Introducing NDEA and Symbolic Program Synthesis | 4 | 5 | 2 | 1 | Garry Tan brings up his open-source experience with G-Stack, prompting François Chollet to explain why symbolic program synthesis is fundamentally different from coding agents and gradient descent. | |
| The Case for Exploring Non-LLM AI Paradigms | 6 | 6 | 3 | 4 | Diana Hu provides thoughtful technical commentary regarding the boundary between formally verifiable domains and nebulous natural language tasks like writing essays, prompting Chollet to elaborate on post-training execution models. | |
| Redefining General Intelligence vs. Automation | 4 | 7 | 4 | 2 | Chollet explicitly rejects the industry's dominant definition of AGI as economic task automation, reframing intelligence purely around sample efficiency and skill acquisition. | |
| The Origins of ARC-AGI & Limits of Gradient Descent | 5 | 7 | 3 | 1 | The hosts demonstrate knowledge of Chollet's background with Keras and early deep learning, while Chollet explains the mathematical bottleneck of gradient descent converging on overfit pattern matching instead of programmatic reasoning. | |
| Evolution from ARC-AGI-1 to ARC-AGI-2 | 5 | 6 | 2 | 2 | Hosts and guest track the historical evolution of ARC benchmarks from base LLMs to OpenAI reasoning models, with Chollet breaking down how brute-force RL loops saturated ARC-2 without truly raising fluid intelligence. | |
| Agentic Harnesses & Confluence Labs Success | 7 | 4 | 1 | 2 | Diana Hu demonstrates deep domain knowledge by citing Confluence Labs from YC W26 batch achieving a 97% score on ARC-AGI-2, prompting Chollet to contextualize external harnesses as engineering scaffolding rather than intrinsic AGI. | |
| Introducing ARC-AGI-3: Measuring Interactive Intelligence | 3 | 6 | 1 | 1 | The co-host asks clarifying questions about ARC-AGI-3, and Chollet explains the shift from static pattern recognition to interactive, agentic video game environments evaluated strictly on exploration efficiency. | |
| Inside the ARC Game Studio and Core Knowledge Priors | 4 | 5 | 1 | 1 | Garry Tan connects interactive testing to DeepMind and early OpenAI work, while Chollet highlights how ARC-3 prevents memorization by grounding games strictly in core knowledge priors rather than seen cultural or visual tropes. | |
| Model Scale vs. Symbolic Compression | 5 | 6 | 2 | 2 | Garry Tan references Douglas Lenat's Cyc project to explore knowledge bases versus reasoning codebases, leading Chollet to frame intelligence as symbolic compression and algorithmic science. | |
| First Principles of Intelligence & NDEA's Research Stack | 4 | 5 | 1 | 1 | The co-host queries whether human biology provides a third path for learning, and Chollet explains why NDEA prioritizes first-principles symbolic program search guided by deep learning rather than replicating biological quirks. | |
| Future ARC Benchmarks and the AGI Timeline | 3 | 5 | 1 | 1 | Chollet outlines the progression roadmap towards ARC-4 through ARC-7 and provides his expected timeline for measurable AGI arrival. | |
| Exploring Alternative AI Research & Scaling Without Human Bottlenecks | 5 | 6 | 2 | 2 | Diana Hu and Garry Tan explore historical AI parallels such as SVMs and genetic algorithms, while Chollet emphasizes the need for research paradigms that eliminate human bottlenecks from the improvement loop. | |
| Open Source Management & Lessons from Keras | 4 | 6 | 1 | 1 | Garry Tan asks for tactical advice on managing his open-source project G-Stack, and Chollet shares leadership lessons and user empathy principles from creating Keras. |