Mar 27, 2026 · 57m · y-combinator

François Chollet: Why Scaling Alone Isn’t Enough for AGI · Y Combinator

François Chollet · 41m spoken Garry Tan · 5m spoken Diana Hu · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Y Combinator podcast episode, François Chollet discusses why scaling traditional large language models alone is insufficient for reaching AGI, detailing his work on symbolic program synthesis at NDEA and the evolution of the ARC-AGI benchmark. He explains how moving beyond parametric deep learning toward verifiable reasoning, interactive environments, and sample-efficient architectures will pave the way for true fluid intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The partners hold 16.2% of the talking time here. How this is scored →

The partners as informed peer 4.5 Guest teaching 5.7 Guest disagreement 1.9 The partners pushing back 1.6
05100:0015:0030:0045:001:02–4:24 · The partners as informed peer 4/10 Introducing NDEA and Symbolic Program Synthesis Garry Tan brings up his open-source experience with G-Stack, prompting François Chollet to explain why symbolic program synthesis is fundamentally different from coding agents and gradient descent.4:24–9:41 · The partners as informed peer 6/10 The Case for Exploring Non-LLM AI Paradigms Diana Hu provides thoughtful technical commentary regarding the boundary between formally verifiable domains and nebulous natural language tasks like writing essays, prompting Chollet to elaborate on post-training execution models.9:41–12:35 · The partners as informed peer 4/10 Redefining General Intelligence vs. Automation Chollet explicitly rejects the industry's dominant definition of AGI as economic task automation, reframing intelligence purely around sample efficiency and skill acquisition.12:35–17:14 · The partners as informed peer 5/10 The Origins of ARC-AGI & Limits of Gradient Descent The hosts demonstrate knowledge of Chollet's background with Keras and early deep learning, while Chollet explains the mathematical bottleneck of gradient descent converging on overfit pattern matching instead of programmatic reasoning.17:14–23:57 · The partners as informed peer 5/10 Evolution from ARC-AGI-1 to ARC-AGI-2 Hosts and guest track the historical evolution of ARC benchmarks from base LLMs to OpenAI reasoning models, with Chollet breaking down how brute-force RL loops saturated ARC-2 without truly raising fluid intelligence.23:57–26:02 · The partners as informed peer 7/10 Agentic Harnesses & Confluence Labs Success Diana Hu demonstrates deep domain knowledge by citing Confluence Labs from YC W26 batch achieving a 97% score on ARC-AGI-2, prompting Chollet to contextualize external harnesses as engineering scaffolding rather than intrinsic AGI.26:02–29:35 · The partners as informed peer 3/10 Introducing ARC-AGI-3: Measuring Interactive Intelligence The co-host asks clarifying questions about ARC-AGI-3, and Chollet explains the shift from static pattern recognition to interactive, agentic video game environments evaluated strictly on exploration efficiency.29:35–34:35 · The partners as informed peer 4/10 Inside the ARC Game Studio and Core Knowledge Priors Garry Tan connects interactive testing to DeepMind and early OpenAI work, while Chollet highlights how ARC-3 prevents memorization by grounding games strictly in core knowledge priors rather than seen cultural or visual tropes.34:35–39:51 · The partners as informed peer 5/10 Model Scale vs. Symbolic Compression Garry Tan references Douglas Lenat's Cyc project to explore knowledge bases versus reasoning codebases, leading Chollet to frame intelligence as symbolic compression and algorithmic science.39:51–43:59 · The partners as informed peer 4/10 First Principles of Intelligence & NDEA's Research Stack The co-host queries whether human biology provides a third path for learning, and Chollet explains why NDEA prioritizes first-principles symbolic program search guided by deep learning rather than replicating biological quirks.43:59–46:25 · The partners as informed peer 3/10 Future ARC Benchmarks and the AGI Timeline Chollet outlines the progression roadmap towards ARC-4 through ARC-7 and provides his expected timeline for measurable AGI arrival.46:25–52:21 · The partners as informed peer 5/10 Exploring Alternative AI Research & Scaling Without Human Bottlenecks Diana Hu and Garry Tan explore historical AI parallels such as SVMs and genetic algorithms, while Chollet emphasizes the need for research paradigms that eliminate human bottlenecks from the improvement loop.52:21–55:18 · The partners as informed peer 4/10 Open Source Management & Lessons from Keras Garry Tan asks for tactical advice on managing his open-source project G-Stack, and Chollet shares leadership lessons and user empathy principles from creating Keras.1:02–4:24 · Guest teaching 5/10 Introducing NDEA and Symbolic Program Synthesis Garry Tan brings up his open-source experience with G-Stack, prompting François Chollet to explain why symbolic program synthesis is fundamentally different from coding agents and gradient descent.4:24–9:41 · Guest teaching 6/10 The Case for Exploring Non-LLM AI Paradigms Diana Hu provides thoughtful technical commentary regarding the boundary between formally verifiable domains and nebulous natural language tasks like writing essays, prompting Chollet to elaborate on post-training execution models.9:41–12:35 · Guest teaching 7/10 Redefining General Intelligence vs. Automation Chollet explicitly rejects the industry's dominant definition of AGI as economic task automation, reframing intelligence purely around sample efficiency and skill acquisition.12:35–17:14 · Guest teaching 7/10 The Origins of ARC-AGI & Limits of Gradient Descent The hosts demonstrate knowledge of Chollet's background with Keras and early deep learning, while Chollet explains the mathematical bottleneck of gradient descent converging on overfit pattern matching instead of programmatic reasoning.17:14–23:57 · Guest teaching 6/10 Evolution from ARC-AGI-1 to ARC-AGI-2 Hosts and guest track the historical evolution of ARC benchmarks from base LLMs to OpenAI reasoning models, with Chollet breaking down how brute-force RL loops saturated ARC-2 without truly raising fluid intelligence.23:57–26:02 · Guest teaching 4/10 Agentic Harnesses & Confluence Labs Success Diana Hu demonstrates deep domain knowledge by citing Confluence Labs from YC W26 batch achieving a 97% score on ARC-AGI-2, prompting Chollet to contextualize external harnesses as engineering scaffolding rather than intrinsic AGI.26:02–29:35 · Guest teaching 6/10 Introducing ARC-AGI-3: Measuring Interactive Intelligence The co-host asks clarifying questions about ARC-AGI-3, and Chollet explains the shift from static pattern recognition to interactive, agentic video game environments evaluated strictly on exploration efficiency.29:35–34:35 · Guest teaching 5/10 Inside the ARC Game Studio and Core Knowledge Priors Garry Tan connects interactive testing to DeepMind and early OpenAI work, while Chollet highlights how ARC-3 prevents memorization by grounding games strictly in core knowledge priors rather than seen cultural or visual tropes.34:35–39:51 · Guest teaching 6/10 Model Scale vs. Symbolic Compression Garry Tan references Douglas Lenat's Cyc project to explore knowledge bases versus reasoning codebases, leading Chollet to frame intelligence as symbolic compression and algorithmic science.39:51–43:59 · Guest teaching 5/10 First Principles of Intelligence & NDEA's Research Stack The co-host queries whether human biology provides a third path for learning, and Chollet explains why NDEA prioritizes first-principles symbolic program search guided by deep learning rather than replicating biological quirks.43:59–46:25 · Guest teaching 5/10 Future ARC Benchmarks and the AGI Timeline Chollet outlines the progression roadmap towards ARC-4 through ARC-7 and provides his expected timeline for measurable AGI arrival.46:25–52:21 · Guest teaching 6/10 Exploring Alternative AI Research & Scaling Without Human Bottlenecks Diana Hu and Garry Tan explore historical AI parallels such as SVMs and genetic algorithms, while Chollet emphasizes the need for research paradigms that eliminate human bottlenecks from the improvement loop.52:21–55:18 · Guest teaching 6/10 Open Source Management & Lessons from Keras Garry Tan asks for tactical advice on managing his open-source project G-Stack, and Chollet shares leadership lessons and user empathy principles from creating Keras.1:02–4:24 · Guest disagreement 2/10 Introducing NDEA and Symbolic Program Synthesis Garry Tan brings up his open-source experience with G-Stack, prompting François Chollet to explain why symbolic program synthesis is fundamentally different from coding agents and gradient descent.4:24–9:41 · Guest disagreement 3/10 The Case for Exploring Non-LLM AI Paradigms Diana Hu provides thoughtful technical commentary regarding the boundary between formally verifiable domains and nebulous natural language tasks like writing essays, prompting Chollet to elaborate on post-training execution models.9:41–12:35 · Guest disagreement 4/10 Redefining General Intelligence vs. Automation Chollet explicitly rejects the industry's dominant definition of AGI as economic task automation, reframing intelligence purely around sample efficiency and skill acquisition.12:35–17:14 · Guest disagreement 3/10 The Origins of ARC-AGI & Limits of Gradient Descent The hosts demonstrate knowledge of Chollet's background with Keras and early deep learning, while Chollet explains the mathematical bottleneck of gradient descent converging on overfit pattern matching instead of programmatic reasoning.17:14–23:57 · Guest disagreement 2/10 Evolution from ARC-AGI-1 to ARC-AGI-2 Hosts and guest track the historical evolution of ARC benchmarks from base LLMs to OpenAI reasoning models, with Chollet breaking down how brute-force RL loops saturated ARC-2 without truly raising fluid intelligence.23:57–26:02 · Guest disagreement 1/10 Agentic Harnesses & Confluence Labs Success Diana Hu demonstrates deep domain knowledge by citing Confluence Labs from YC W26 batch achieving a 97% score on ARC-AGI-2, prompting Chollet to contextualize external harnesses as engineering scaffolding rather than intrinsic AGI.26:02–29:35 · Guest disagreement 1/10 Introducing ARC-AGI-3: Measuring Interactive Intelligence The co-host asks clarifying questions about ARC-AGI-3, and Chollet explains the shift from static pattern recognition to interactive, agentic video game environments evaluated strictly on exploration efficiency.29:35–34:35 · Guest disagreement 1/10 Inside the ARC Game Studio and Core Knowledge Priors Garry Tan connects interactive testing to DeepMind and early OpenAI work, while Chollet highlights how ARC-3 prevents memorization by grounding games strictly in core knowledge priors rather than seen cultural or visual tropes.34:35–39:51 · Guest disagreement 2/10 Model Scale vs. Symbolic Compression Garry Tan references Douglas Lenat's Cyc project to explore knowledge bases versus reasoning codebases, leading Chollet to frame intelligence as symbolic compression and algorithmic science.39:51–43:59 · Guest disagreement 1/10 First Principles of Intelligence & NDEA's Research Stack The co-host queries whether human biology provides a third path for learning, and Chollet explains why NDEA prioritizes first-principles symbolic program search guided by deep learning rather than replicating biological quirks.43:59–46:25 · Guest disagreement 1/10 Future ARC Benchmarks and the AGI Timeline Chollet outlines the progression roadmap towards ARC-4 through ARC-7 and provides his expected timeline for measurable AGI arrival.46:25–52:21 · Guest disagreement 2/10 Exploring Alternative AI Research & Scaling Without Human Bottlenecks Diana Hu and Garry Tan explore historical AI parallels such as SVMs and genetic algorithms, while Chollet emphasizes the need for research paradigms that eliminate human bottlenecks from the improvement loop.52:21–55:18 · Guest disagreement 1/10 Open Source Management & Lessons from Keras Garry Tan asks for tactical advice on managing his open-source project G-Stack, and Chollet shares leadership lessons and user empathy principles from creating Keras.1:02–4:24 · The partners pushing back 1/10 Introducing NDEA and Symbolic Program Synthesis Garry Tan brings up his open-source experience with G-Stack, prompting François Chollet to explain why symbolic program synthesis is fundamentally different from coding agents and gradient descent.4:24–9:41 · The partners pushing back 4/10 The Case for Exploring Non-LLM AI Paradigms Diana Hu provides thoughtful technical commentary regarding the boundary between formally verifiable domains and nebulous natural language tasks like writing essays, prompting Chollet to elaborate on post-training execution models.9:41–12:35 · The partners pushing back 2/10 Redefining General Intelligence vs. Automation Chollet explicitly rejects the industry's dominant definition of AGI as economic task automation, reframing intelligence purely around sample efficiency and skill acquisition.12:35–17:14 · The partners pushing back 1/10 The Origins of ARC-AGI & Limits of Gradient Descent The hosts demonstrate knowledge of Chollet's background with Keras and early deep learning, while Chollet explains the mathematical bottleneck of gradient descent converging on overfit pattern matching instead of programmatic reasoning.17:14–23:57 · The partners pushing back 2/10 Evolution from ARC-AGI-1 to ARC-AGI-2 Hosts and guest track the historical evolution of ARC benchmarks from base LLMs to OpenAI reasoning models, with Chollet breaking down how brute-force RL loops saturated ARC-2 without truly raising fluid intelligence.23:57–26:02 · The partners pushing back 2/10 Agentic Harnesses & Confluence Labs Success Diana Hu demonstrates deep domain knowledge by citing Confluence Labs from YC W26 batch achieving a 97% score on ARC-AGI-2, prompting Chollet to contextualize external harnesses as engineering scaffolding rather than intrinsic AGI.26:02–29:35 · The partners pushing back 1/10 Introducing ARC-AGI-3: Measuring Interactive Intelligence The co-host asks clarifying questions about ARC-AGI-3, and Chollet explains the shift from static pattern recognition to interactive, agentic video game environments evaluated strictly on exploration efficiency.29:35–34:35 · The partners pushing back 1/10 Inside the ARC Game Studio and Core Knowledge Priors Garry Tan connects interactive testing to DeepMind and early OpenAI work, while Chollet highlights how ARC-3 prevents memorization by grounding games strictly in core knowledge priors rather than seen cultural or visual tropes.34:35–39:51 · The partners pushing back 2/10 Model Scale vs. Symbolic Compression Garry Tan references Douglas Lenat's Cyc project to explore knowledge bases versus reasoning codebases, leading Chollet to frame intelligence as symbolic compression and algorithmic science.39:51–43:59 · The partners pushing back 1/10 First Principles of Intelligence & NDEA's Research Stack The co-host queries whether human biology provides a third path for learning, and Chollet explains why NDEA prioritizes first-principles symbolic program search guided by deep learning rather than replicating biological quirks.43:59–46:25 · The partners pushing back 1/10 Future ARC Benchmarks and the AGI Timeline Chollet outlines the progression roadmap towards ARC-4 through ARC-7 and provides his expected timeline for measurable AGI arrival.46:25–52:21 · The partners pushing back 2/10 Exploring Alternative AI Research & Scaling Without Human Bottlenecks Diana Hu and Garry Tan explore historical AI parallels such as SVMs and genetic algorithms, while Chollet emphasizes the need for research paradigms that eliminate human bottlenecks from the improvement loop.52:21–55:18 · The partners pushing back 1/10 Open Source Management & Lessons from Keras Garry Tan asks for tactical advice on managing his open-source project G-Stack, and Chollet shares leadership lessons and user empathy principles from creating Keras.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 31.8% · guest 68.2%0:00 · the partners 31.8% · guest 68.2%3:00 · the partners 0.4% · guest 99.6%3:00 · the partners 0.4% · guest 99.6%6:00 · the partners 24.1% · guest 75.9%6:00 · the partners 24.1% · guest 75.9%9:00 · the partners 5.9% · guest 94.1%9:00 · the partners 5.9% · guest 94.1%12:00 · the partners 5.4% · guest 94.6%12:00 · the partners 5.4% · guest 94.6%15:00 · the partners 34.7% · guest 65.3%15:00 · the partners 34.7% · guest 65.3%18:00 · the partners 8.2% · guest 91.8%18:00 · the partners 8.2% · guest 91.8%21:00 · the partners 16.5% · guest 83.5%21:00 · the partners 16.5% · guest 83.5%24:00 · the partners 48% · guest 52%24:00 · the partners 48% · guest 52%27:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%30:00 · the partners 13.2% · guest 86.8%30:00 · the partners 13.2% · guest 86.8%33:00 · the partners 32.5% · guest 67.5%33:00 · the partners 32.5% · guest 67.5%36:00 · the partners 20.1% · guest 79.9%36:00 · the partners 20.1% · guest 79.9%39:00 · the partners 0% · guest 100%39:00 · the partners 0% · guest 100%42:00 · the partners 0% · guest 100%42:00 · the partners 0% · guest 100%45:00 · the partners 6.2% · guest 93.8%45:00 · the partners 6.2% · guest 93.8%48:00 · the partners 18.2% · guest 81.8%48:00 · the partners 18.2% · guest 81.8%51:00 · the partners 23% · guest 77%51:00 · the partners 23% · guest 77%54:00 · the partners 18.4% · guest 81.6%54:00 · the partners 18.4% · guest 81.6%57:00 · the partners 57.5% · guest 42.5%57:00 · the partners 57.5% · guest 42.5%
Sharpest disagreement ▶ 9:52 Rejecting the industry's economic definition of AGI

Chollet firmly dismisses the widely accepted definition of AGI as economic task automation, arguing it reflects task automation rather than genuine general intelligence.

Hardest push from the partners ▶ 7:11 Diana Hu pushes back on non-verifiable domains

Diana Hu challenges the universality of verifiable reward loops by asking how fuzzy, subjective tasks like essay writing can realistically be transformed into formal verification functions.

Biggest teaching moment ▶ 14:05 Explaining the fundamental limit of gradient descent

Chollet walks through his Google Brain research showing that gradient descent is mathematically incapable of discovering generalizable symbolic programs and defaults to overfit pattern matching.

The partners hold their own ▶ 23:57 Diana Hu shares YC batch benchmark results

Diana Hu demonstrates YC's direct front-row exposure to cutting-edge AI developments by detailing how Confluence Labs saturated ARC-AGI-2 to 97% using custom harnesses during the W26 batch.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Introducing NDEA and Symbolic Program Synthesis 4521 Garry Tan brings up his open-source experience with G-Stack, prompting François Chollet to explain why symbolic program synthesis is fundamentally different from coding agents and gradient descent.
The Case for Exploring Non-LLM AI Paradigms 6634 Diana Hu provides thoughtful technical commentary regarding the boundary between formally verifiable domains and nebulous natural language tasks like writing essays, prompting Chollet to elaborate on post-training execution models.
Redefining General Intelligence vs. Automation 4742 Chollet explicitly rejects the industry's dominant definition of AGI as economic task automation, reframing intelligence purely around sample efficiency and skill acquisition.
The Origins of ARC-AGI & Limits of Gradient Descent 5731 The hosts demonstrate knowledge of Chollet's background with Keras and early deep learning, while Chollet explains the mathematical bottleneck of gradient descent converging on overfit pattern matching instead of programmatic reasoning.
Evolution from ARC-AGI-1 to ARC-AGI-2 5622 Hosts and guest track the historical evolution of ARC benchmarks from base LLMs to OpenAI reasoning models, with Chollet breaking down how brute-force RL loops saturated ARC-2 without truly raising fluid intelligence.
Agentic Harnesses & Confluence Labs Success 7412 Diana Hu demonstrates deep domain knowledge by citing Confluence Labs from YC W26 batch achieving a 97% score on ARC-AGI-2, prompting Chollet to contextualize external harnesses as engineering scaffolding rather than intrinsic AGI.
Introducing ARC-AGI-3: Measuring Interactive Intelligence 3611 The co-host asks clarifying questions about ARC-AGI-3, and Chollet explains the shift from static pattern recognition to interactive, agentic video game environments evaluated strictly on exploration efficiency.
Inside the ARC Game Studio and Core Knowledge Priors 4511 Garry Tan connects interactive testing to DeepMind and early OpenAI work, while Chollet highlights how ARC-3 prevents memorization by grounding games strictly in core knowledge priors rather than seen cultural or visual tropes.
Model Scale vs. Symbolic Compression 5622 Garry Tan references Douglas Lenat's Cyc project to explore knowledge bases versus reasoning codebases, leading Chollet to frame intelligence as symbolic compression and algorithmic science.
First Principles of Intelligence & NDEA's Research Stack 4511 The co-host queries whether human biology provides a third path for learning, and Chollet explains why NDEA prioritizes first-principles symbolic program search guided by deep learning rather than replicating biological quirks.
Future ARC Benchmarks and the AGI Timeline 3511 Chollet outlines the progression roadmap towards ARC-4 through ARC-7 and provides his expected timeline for measurable AGI arrival.
Exploring Alternative AI Research & Scaling Without Human Bottlenecks 5622 Diana Hu and Garry Tan explore historical AI parallels such as SVMs and genetic algorithms, while Chollet emphasizes the need for research paradigms that eliminate human bottlenecks from the improvement loop.
Open Source Management & Lessons from Keras 4611 Garry Tan asks for tactical advice on managing his open-source project G-Stack, and Chollet shares leadership lessons and user empathy principles from creating Keras.

Statements from this episode (31)

Assertion Supported
Tan: gstack Reached 40,000 Stars and 100+ Contributor PRs
“I have sort of this viral moment right now where I got to 40,000 stars this morning on a G stack. So it's like, oh, this is an open source project that now is one of the biggest ones. And I have more than a hundred PRs from contributors to deal with.”
Garry Tan Mar 27, 2026 ▶ 1:28
Prediction Not checkable as stated
Chollet: Symbolic Models Will Eventually Replicate and Outperform Deep Learning
“And so everything you're doing with machine learning today, with parametric curves, we should be able to do it. With symbolic models in the future in a way that will be much, much closer to optimality. Much closer to optimality in the sense that you're going t…”
François Chollet Mar 27, 2026 ▶ 3:42
Insight
Chollet: Parametric Learning Cannot Find Minimum Description Length Models
“You know, the minimum description length principle that the model of the data that is most likely to generalize Is the shortest. And I think you cannot find a model like this. If you're doing parametric learning, you need to try symbolic learning.”
François Chollet Mar 27, 2026 ▶ 4:10
Prediction Open · timeframe Mar 2076
Chollet: AI in 50 Years Will Not Use Today's LLM Stack
“I personally don't think that machine learning or AI in 50 years is still going to be built on this stack.”
François Chollet Mar 27, 2026 ▶ 5:01
Disclosure
Chollet: NDEA Has Only a 10% to 15% Chance of Success
“Like we have maybe a 10 or 15% chance of success. But that is enough that it's worth trying, right?”
François Chollet Mar 27, 2026 ▶ 5:35
Insight
Chollet: Current LLM Stack Can Fully Automate Any Formally Verifiable Domain
“And I think right now we're in this situation where any problem where the solutions you've proposed can be formally verified, and you can actually trust the reward signal. It's not just some guess made by a model. Any domain like this can be fully automated wi…”
François Chollet Mar 27, 2026 ▶ 6:40
Prediction Not checkable as stated
Chollet: Mathematics AI Revolution Is Coming in the Next Few Years
“I think mathematics is also, it's also primed to see a revolution in the next few years for the same reasons, again, because The domain just gives you verifiable rewards.”
François Chollet Mar 27, 2026 ▶ 7:01
Prediction Not checkable as stated
Chollet: LLM Progress in Non-Verifiable Domains Will Slow or Stall
“Progress of reasoning models and base LLMs on this type of domain is, is, you know, it's going to be very slow because the stack we're using, like the LLM stack is very, very reliant on its trained data. It's basically just operationalizing the trained data. A…”
François Chollet Mar 27, 2026 ▶ 8:02
Insight
Chollet: General Intelligence Is Human-Level Skill Acquisition Efficiency
“General intelligence is human level. Skill acquisition efficiency on the same scope of tasks that humans could potentially learn to do.”
François Chollet Mar 27, 2026 ▶ 10:42
Prediction Open · timeframe Mar 2056
Chollet: Future AI Won't Just Be Layered Harnesses on Base LLMs
“Future AI in a few decades it's not going to be this harness on top of a reasoning model on top of a base LLM.”
François Chollet Mar 27, 2026 ▶ 12:24
Insight
Chollet: Gradient Descent Fails at Reasoning by Defaulting to Pattern Matching
“You could not really get Gradient descent to encode sort of like reasoning style algorithms. It was not because the models could not represent these algorithms. It was because gradient descent could not find them, right? So the problem was that it wasn't about…”
François Chollet Mar 27, 2026 ▶ 14:11
Assertion Supported
Chollet: Base LLMs Score Under 10% on ARC-AGI-1
“So basal alarms were scoring extremely low on V-one, like sub-ten percent, basically. And, I mean, it was true of the original, like, GPT-III actually scoring zero, but that's even true of the latest basal alarms today, you know, as of March.”
François Chollet Mar 27, 2026 ▶ 17:57
Insight
Chollet: Competency Balances a Trade-Off Between Intelligence and Knowledge
“When it comes to Competency. There's always a trade-off between intelligence and knowledge. If you have more knowledge, if you have better training, you need less intelligence to be competent.”
François Chollet Mar 27, 2026 ▶ 23:01
Assertion Supported
Hu: Confluence Labs Saturated the ARC-2 Benchmark at 97% Accuracy
“I actually worked with a company in the winter 26 batch not too long ago called Confluence Lab, which actually ended up saturating the V-two results with 97%, and I think their task cost was a lot more efficient too.”
Diana Hu Mar 27, 2026 ▶ 24:27
Insight
Chollet: Human-Engineered Agent Harnesses Prove AI Is Far From AGI
“I mean, to me, the fact that you need humans to engineer these harnesses is also a sign that we're short of AGI today, because if we had AGI, you know, AGI would just make its own harness. It would not need to be told how to solve a problem.”
François Chollet Mar 27, 2026 ▶ 25:24
Assertion Supported
Chollet: All ARC-3 environments are solvable by untrained humans
“All of these test environments in Arc three Are solvable by humans with no prior training because we actually tested them on, on regular people.”
François Chollet Mar 27, 2026 ▶ 27:52
Disclosure
Chollet: ARC-3 private test set differs substantially from public set
“We've deliberately tried to create a private set of environments that is significantly different from the public set. Like you can look at the public set. It's not actually giving you that much information about what's in the private set. In the private set, y…”
François Chollet Mar 27, 2026 ▶ 28:57
Disclosure
Chollet: ARC Prize Built a Video Game Studio Generating 250+ Games
“We set up an entire video game studio, right, to create them. So we got over 250 games.”
François Chollet Mar 27, 2026 ▶ 29:39
Insight
Chollet: RL Benchmarks Like Dota and Atari Test Memorization, Not Intelligence
“If you look at Atari games, for instance, or even Dota, you're training on, on the same environment as what you use for testing. So effectively, you're just trying to memorize the best strategies. You're trying to at training time, explore the full space of po…”
François Chollet Mar 27, 2026 ▶ 33:08
Prediction Open · timeframe Mar 2031
Chollet: AGI Codebase Will Be Under 10,000 Lines on 1980s Compute
“I do believe that, you know, when you create a GI retrospectively, it will turn out that it's a code base that's less than 10,000 lines of code. And that if you had known about it back in the 19 eighties, you could have done a GI back then using the computer r…”
François Chollet Mar 27, 2026 ▶ 36:09
Insight
Chollet: Science Is Symbolic Compression, Not Merely Curve Fitting
“Science is not about curve fitting. Science is about finding the equation, finding the most compressive symbolic model of your pile of observation.”
François Chollet Mar 27, 2026 ▶ 39:30
Prediction Not checkable as stated
Chollet: AI built from first principles will be more efficient than human brains
“It's an implementation of fundamental principles, the fundamental principles of intelligence, which, you know, I think we can identify these principles and re-implement intelligence from scratch, from first principles, in a way it will be much more efficient t…”
François Chollet Mar 27, 2026 ▶ 40:27
Opinion
Chollet: Replicating Biologically Plausible Brain Mechanics for AI Is Counter-Productive
“I think it would be counter-productive to just try to, you know, observe it and re-implement it, like and make it biologically plausible.”
François Chollet Mar 27, 2026 ▶ 40:50
Insight
Chollet: Human generalization stems from symbolic, causal program synthesis
“I do believe the human mind does at the highest level something that looks a lot like programs in this, like we're currently building Causal models of our surroundings. Like we are describing our surroundings in our mind as, you know, a set of objects and agen…”
François Chollet Mar 27, 2026 ▶ 41:10
Insight
Chollet: Deep learning guidance is necessary to break combinatorial program search
“You have to break the combinatorial wall, and the way to do it is to add deep learning guidance. It's actually very similar to the principles that analyze something like AlphaGo or AlphaZero.”
François Chollet Mar 27, 2026 ▶ 42:54
Disclosure
Chollet: ARC-4 will focus on continual learning and compounding levels
“So there will be ARC four, which will be in the spirit of ARC three, but more focused on continual learning and curriculum learning at longer timescales. So you're gonna have fewer games but they're gonna have way more levels. And the levels are going to be co…”
François Chollet Mar 27, 2026 ▶ 44:44
Prediction Not checkable as stated
Chollet: Artificial General Intelligence Will Likely Arrive Around ARC-6 or ARC-7
“My timeline to AGI, you know, if you just try to extrapolate from the current rate of progress and the amount of investment that's going into, not just the LLM stack, but also like side ideas, side bets that might work out, like, you know, India, for instance,…”
François Chollet Mar 27, 2026 ▶ 45:50
What-if
Chollet: Equal investment in genetic algorithms would have yielded exciting results
“If you had thrown the same amount of investment into almost anything else, you would also have seen extremely exciting results, like genetic algorithms, for instance.”
François Chollet Mar 27, 2026 ▶ 46:45
Insight
Chollet: Scaling genetic algorithms could automate scientific discovery
“If you try to scale up genetic algorithms, I mean, I'm sure you can do incredible things with that. You could, in fact, probably do new science. Because that's based on search, and search is the best fit for automating the scientific method.”
François Chollet Mar 27, 2026 ▶ 47:04
Insight
Chollet: AI Architectures Requiring Human Engineers to Scale Will Fail
“If you're working on something, but the only way to increase the capabilities of the system Is to have human engineers and researchers spend time on it. It will not work because even if the idea is very clever and very elegant and works really well, capabiliti…”
François Chollet Mar 27, 2026 ▶ 50:27
Insight
Chollet: AI Empowers Workers With Deep Expertise Rather Than Replacing Them
“The more, you know, the more expertise you have, but things like programming, for instance, the better you're able to Use and leverage these tools for your own benefit. And with the right kind of expertise all this AI progress is actually empowerment. Like, it…”
François Chollet Mar 27, 2026 ▶ 56:01
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.