Jul 3, 2025 · 34m · y-combinator

François Chollet: How We Get To AGI · Y Combinator

François Chollet · 31m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In his Y Combinator AI Startup School talk, François Chollet challenges the dominant pretraining scaling dogma and outlines a roadmap toward AGI centered on fluid intelligence, test-time adaptation, and hybrid cognitive architectures. He presents the ARC-AGI benchmark suite and demonstrates how combining continuous neural perception with discrete symbolic program search enables true, on-the-fly problem solving.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The partners as informed peer 0.0 Guest teaching 3.2 Guest disagreement 2.0 The partners pushing back 0.0
05100:0010:0020:0030:001:38–4:51 · The partners as informed peer 0/10 Fluid Intelligence vs Pretraining Scaling and Test Time Adaptation This is a presentation segment where François Chollet explains how the AI community moved from the pretraining scaling dogma to test-time adaptation. The host only provides a brief bridging prompt ('So what happened?'), resulting in zero host pushback and expertise scores.4:51–11:44 · The partners as informed peer 0/10 Defining Intelligence: Process, Efficiency, and the Shortcut Rule Chollet delivers a monologue contrasting the Minsky task-focused view with McCarthy's definition of intelligence as dealing with novel situations. He rejects human exam-style benchmarks and explains the shortcut rule without host interaction.11:44–20:17 · The partners as informed peer 0/10 ARC-AGI Benchmark Evolution: ARC-1, ARC-2, and ARC-3 Chollet walks through the technical progression from ARC-1 to ARC-2 and ARC-3, detailing why base LLMs score near zero while everyday humans solve them easily. The host does not speak in this segment.20:17–29:32 · The partners as informed peer 0/10 Building AGI: The Kaleidoscope Hypothesis and Two Poles of Abstraction In an uninterrupted technical lecture, Chollet articulates the Kaleidoscope Hypothesis and distinguishes between Type 1 continuous perception and Type 2 discrete program search.29:32–34:47 · The partners as informed peer 0/10 The Road Ahead: Hybrid Architectures and NDEA Chollet concludes his keynote by outlining the hybrid architecture being built at NDEA to unify deep learning intuition and program synthesis for scientific discovery.1:38–4:51 · Guest teaching 2/10 Fluid Intelligence vs Pretraining Scaling and Test Time Adaptation This is a presentation segment where François Chollet explains how the AI community moved from the pretraining scaling dogma to test-time adaptation. The host only provides a brief bridging prompt ('So what happened?'), resulting in zero host pushback and expertise scores.4:51–11:44 · Guest teaching 4/10 Defining Intelligence: Process, Efficiency, and the Shortcut Rule Chollet delivers a monologue contrasting the Minsky task-focused view with McCarthy's definition of intelligence as dealing with novel situations. He rejects human exam-style benchmarks and explains the shortcut rule without host interaction.11:44–20:17 · Guest teaching 3/10 ARC-AGI Benchmark Evolution: ARC-1, ARC-2, and ARC-3 Chollet walks through the technical progression from ARC-1 to ARC-2 and ARC-3, detailing why base LLMs score near zero while everyday humans solve them easily. The host does not speak in this segment.20:17–29:32 · Guest teaching 4/10 Building AGI: The Kaleidoscope Hypothesis and Two Poles of Abstraction In an uninterrupted technical lecture, Chollet articulates the Kaleidoscope Hypothesis and distinguishes between Type 1 continuous perception and Type 2 discrete program search.29:32–34:47 · Guest teaching 3/10 The Road Ahead: Hybrid Architectures and NDEA Chollet concludes his keynote by outlining the hybrid architecture being built at NDEA to unify deep learning intuition and program synthesis for scientific discovery.1:38–4:51 · Guest disagreement 2/10 Fluid Intelligence vs Pretraining Scaling and Test Time Adaptation This is a presentation segment where François Chollet explains how the AI community moved from the pretraining scaling dogma to test-time adaptation. The host only provides a brief bridging prompt ('So what happened?'), resulting in zero host pushback and expertise scores.4:51–11:44 · Guest disagreement 3/10 Defining Intelligence: Process, Efficiency, and the Shortcut Rule Chollet delivers a monologue contrasting the Minsky task-focused view with McCarthy's definition of intelligence as dealing with novel situations. He rejects human exam-style benchmarks and explains the shortcut rule without host interaction.11:44–20:17 · Guest disagreement 2/10 ARC-AGI Benchmark Evolution: ARC-1, ARC-2, and ARC-3 Chollet walks through the technical progression from ARC-1 to ARC-2 and ARC-3, detailing why base LLMs score near zero while everyday humans solve them easily. The host does not speak in this segment.20:17–29:32 · Guest disagreement 2/10 Building AGI: The Kaleidoscope Hypothesis and Two Poles of Abstraction In an uninterrupted technical lecture, Chollet articulates the Kaleidoscope Hypothesis and distinguishes between Type 1 continuous perception and Type 2 discrete program search.29:32–34:47 · Guest disagreement 1/10 The Road Ahead: Hybrid Architectures and NDEA Chollet concludes his keynote by outlining the hybrid architecture being built at NDEA to unify deep learning intuition and program synthesis for scientific discovery.1:38–4:51 · The partners pushing back 0/10 Fluid Intelligence vs Pretraining Scaling and Test Time Adaptation This is a presentation segment where François Chollet explains how the AI community moved from the pretraining scaling dogma to test-time adaptation. The host only provides a brief bridging prompt ('So what happened?'), resulting in zero host pushback and expertise scores.4:51–11:44 · The partners pushing back 0/10 Defining Intelligence: Process, Efficiency, and the Shortcut Rule Chollet delivers a monologue contrasting the Minsky task-focused view with McCarthy's definition of intelligence as dealing with novel situations. He rejects human exam-style benchmarks and explains the shortcut rule without host interaction.11:44–20:17 · The partners pushing back 0/10 ARC-AGI Benchmark Evolution: ARC-1, ARC-2, and ARC-3 Chollet walks through the technical progression from ARC-1 to ARC-2 and ARC-3, detailing why base LLMs score near zero while everyday humans solve them easily. The host does not speak in this segment.20:17–29:32 · The partners pushing back 0/10 Building AGI: The Kaleidoscope Hypothesis and Two Poles of Abstraction In an uninterrupted technical lecture, Chollet articulates the Kaleidoscope Hypothesis and distinguishes between Type 1 continuous perception and Type 2 discrete program search.29:32–34:47 · The partners pushing back 0/10 The Road Ahead: Hybrid Architectures and NDEA Chollet concludes his keynote by outlining the hybrid architecture being built at NDEA to unify deep learning intuition and program synthesis for scientific discovery.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 0% · guest 100%0:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%3:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%6:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%9:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%12:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%15:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%18:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%21:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%24:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%30:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%33:00 · the partners 0% · guest 100%
Sharpest disagreement ▶ 4:25 Dismantling the pretraining scaling dogma

Chollet directly challenges the AI field's dominant assumption from recent years, noting that scaling models was dogmatic and almost no serious researcher believes it will reach AGI anymore.

Hardest push from the partners ▶ 4:35 Prompting the breakdown of scaling

In a talk format with virtually no host intervention, the host asks 'So what happened?' to prompt Chollet to explain why pure scaling failed.

Biggest teaching moment ▶ 5:45 Category error of skill versus intelligence

Chollet clarifies that measuring task performance is a fundamental category error, illustrating the difference using a road network versus a road-building company.

The partners hold their own ▶ 4:35 Guiding the talk flow

Because this episode is a solo presentation with almost no host presence, the brief interjection guiding the topic is the only host contribution.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Fluid Intelligence vs Pretraining Scaling and Test Time Adaptation 0220 This is a presentation segment where François Chollet explains how the AI community moved from the pretraining scaling dogma to test-time adaptation. The host only provides a brief bridging prompt ('So what happened?'), resulting in zero host pushback and expertise scores.
Defining Intelligence: Process, Efficiency, and the Shortcut Rule 0430 Chollet delivers a monologue contrasting the Minsky task-focused view with McCarthy's definition of intelligence as dealing with novel situations. He rejects human exam-style benchmarks and explains the shortcut rule without host interaction.
ARC-AGI Benchmark Evolution: ARC-1, ARC-2, and ARC-3 0320 Chollet walks through the technical progression from ARC-1 to ARC-2 and ARC-3, detailing why base LLMs score near zero while everyday humans solve them easily. The host does not speak in this segment.
Building AGI: The Kaleidoscope Hypothesis and Two Poles of Abstraction 0420 In an uninterrupted technical lecture, Chollet articulates the Kaleidoscope Hypothesis and distinguishes between Type 1 continuous perception and Type 2 discrete program search.
The Road Ahead: Hybrid Architectures and NDEA 0310 Chollet concludes his keynote by outlining the hybrid architecture being built at NDEA to unify deep learning intuition and program synthesis for scientific discovery.

Statements from this episode (14)

Assertion Partly supported
Chollet: Compute costs have fallen two orders of magnitude per decade since 1940
“The cost of compute has been consistently falling by two orders of magnitude every decade since 1940.”
François Chollet Jul 3, 2025 ▶ 0:13
Opinion
Chollet: AI field is obsessed with idea that scale alone yields AGI
“Our field became obsessed with the idea that general intelligence would spontaneously emerge by cramming more and more data into bigger and bigger models.”
François Chollet Jul 3, 2025 ▶ 1:27
Assertion Supported
Chollet: Scaling LLMs 50,000x Only Lifted ARC Accuracy to 10%
“And from at the time, back in 2019 to now, with a model like GPT 4.5, for instance, there's been a roughly 50,000 X scale up of basal alarms. And we went from zero percent accuracy on that benchmark to roughly 10%, which is not a lot.”
François Chollet Jul 3, 2025 ▶ 2:09
Assertion Supported
Chollet: Fine-Tuned OpenAI o3 Reached Human-Level Performance on ARC
“So in particular, in December last year, OpenAI previewed its, ah, all three model, and they used a version of it that was, ah, fine-tuned specifically on Arc, and that showed human-level performance on that benchmark versus time.”
François Chollet Jul 3, 2025 ▶ 3:29
Assertion Supported
Chollet: Every High-Performing ARC AI Method Uses Test-Time Adaptation
“And today, every single AI approach that performs well on Arc is using one of these techniques.”
François Chollet Jul 3, 2025 ▶ 4:13
Insight
Chollet: Displaying skill across tasks does not demonstrate intelligence
“Intelligence is a process, and skill is the output of that process. The skill itself is not intelligence, and displaying skill at any number of tasks does not show intelligence.”
François Chollet Jul 3, 2025 ▶ 5:52
Opinion
Chollet: Human exam benchmarks cannot measure progress toward AGI
“And that's the reason why using exam-like benchmarks with AI models is a bad idea. They're not going to tell you how close we are to AI. Because human exams weren't designed to measure intelligence. They were designed to measure task-specific skill and knowled…”
François Chollet Jul 3, 2025 ▶ 7:21
Assertion Open · timeframe Jul 2026
Chollet: Ten random people with majority voting score 100% on ARC-2
“And all tasks in Arc-II were sold by at least two other people that saw it. And each task was seen on average by about seven people. And so what that tells you is that a group of 10 random people with majority voting would score 100% on Arc-II.”
François Chollet Jul 3, 2025 ▶ 17:10
Assertion Supported
Chollet: Base LLMs score 0% and static reasoning scores 1-2% on ARC-2
“Well, if you take Bazel Alums, model Slack, GPT-IV-IV-V, LAMA-IV, it's simple, they get zero percent. There is simply no way to do these tasks simply via memorization. Next, if you look at static reasoning systems, so systems that use a single chain of tasks t…”
François Chollet Jul 3, 2025 ▶ 17:36
Assertion Not checkable as stated
Chollet: Gradient descent requires 3 to 4 orders of magnitude more data than humans
“Gradient descent requires vast amounts of data to distill simple abstractions. Many orders of magnitude more data than what humans need. Roughly three to four orders of magnitude more.”
François Chollet Jul 3, 2025 ▶ 24:13
Assertion Supported
Chollet: SOTA test-time adaptation takes thousands in compute to solve ARC-1
“Even the latest set of the art CTA techniques they still need thousands of dollars of compute to solve arc one at human level. And that doesn't even scale to arc two.”
François Chollet Jul 3, 2025 ▶ 24:28
Assertion Not checkable as stated
Chollet: All inventive AI systems rely on discrete search
“All known AI systems today that are capable of some kind of invention, some kind of creativity, they rely on discrete search.”
François Chollet Jul 3, 2025 ▶ 27:37
Opinion
Chollet: AI cannot reach high capability without combining Type 1 and Type 2 systems
“I really don't think that you're going to go very far if you go all in on just one of them, like all in on type one or all in on type two. I think that if you want to really unlock their potential, you have to combine them together, and that's what human intel…”
François Chollet Jul 3, 2025 ▶ 29:36
Prediction Not checkable as stated
Chollet: AI will evolve into meta-learners synthesizing software on the fly
“AI is going to move towards systems that are more like programmers that approach a new task by writing software for it. And when faced with a new task, your programmer like MetaLearner will synthesize on the fly a program or model that is adapted to the task.”
François Chollet Jul 3, 2025 ▶ 31:59
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.