Apr 6, 2025 · 33m · tbpn

Mike Knoop (Arc Prize) on Why Scaling AI Won’t Get Us to AGI

Mike Knoop · 19m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Zapier co-founder Mike Knoop explains why scaling pre-training and test-time compute is insufficient for reaching AGI, making the case for the ARC Prize to benchmark genuine fluid reasoning and stimulate decentralized algorithmic innovation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.0 Guest teaching 4.8 Guest disagreement 1.9 The hosts pushing back 1.1
05100:0010:0020:0030:000:00–2:14 · The hosts as informed peer 3/10 Mike Knoop on Zapier, LLM Reliability, and ARC Prize The host warmly opens the episode and asks Mike Knoop to introduce his background at Zapier and the genesis of ARC Prize. Mike explains his realization that enterprise customers found LLMs too unreliable for unsupervised automation.2:14–6:55 · The hosts as informed peer 4/10 ARC-AGI Evolution and the Limits of Test-Time Compute Mike explains the mechanics of ARC-AGI 1 and ARC-AGI 2, detailing how recent models like o3 and test-time compute fail to generalize efficiently on novel benchmarks. Jordy and the host interject with questions about capital constraints versus algorithmic bottlenecks.6:55–11:50 · The hosts as informed peer 6/10 Decentralizing Frontier AI Research and the Vesuvius Challenge The host draws a strong parallel between ARC Prize and the Vesuvius Challenge to decentralized problem solving. The host then asks technical questions about Python execution and overfitting guardrails, which Mike clarifies by explaining Kaggle compute constraints and private test splits.11:50–15:08 · The hosts as informed peer 4/10 Overcoming Reliability Bottlenecks to Build Consumer AI Agents Jordy asks why consumer agent experiences like flight booking remain underwhelming across major tech companies. Mike lays out his framework of concentric rings of reliability risk and explains how fluid intelligence on ARC translates directly to robust agentic execution.15:08–18:19 · The hosts as informed peer 5/10 Defining AGI Through Human-AI Task Gaps Versus Economics The host proposes measuring AGI by macroeconomic contribution thresholds such as percentage of global GDP. Mike firmly pushes back against this definition, arguing that economic metrics obscure underlying capability mechanisms like memorization versus true adaptation.18:19–26:56 · The hosts as informed peer 7/10 Viral Ghibli AI Models and Novel Architectural Exploration The host demonstrates sharp domain expertise by sharing an A/B social media test comparing a real Ghibli still with an AI-generated scene, explaining the uncanny valley dynamics of artistic style transfer. Mike connects this to non-diffusion architectural exploration and warns startups against brute-force model pre-training.26:56–29:30 · The hosts as informed peer 7/10 Reevaluating Gary Marcus, Scaling Limits, and the Bitter Lesson The host articulates an informed synthesis between Gary Marcus's critiques of pure deep learning and Rich Sutton's Bitter Lesson. Mike agrees Marcus was largely vindicated on generalization limits and reinterprets Sutton's paper by highlighting that scalable architectures still originate from human ingenuity.29:30–33:32 · The hosts as informed peer 4/10 Program Synthesis Pioneers and the Enduring Value of Coding Mike highlights key researchers in program synthesis before the host asks if someone could create an ARC-like challenge tailored specifically for human programmers. Mike gently educates the host by revealing that ARC is fundamentally already a program synthesis challenge operating on discrete numeric matrices.0:00–2:14 · Guest teaching 2/10 Mike Knoop on Zapier, LLM Reliability, and ARC Prize The host warmly opens the episode and asks Mike Knoop to introduce his background at Zapier and the genesis of ARC Prize. Mike explains his realization that enterprise customers found LLMs too unreliable for unsupervised automation.2:14–6:55 · Guest teaching 6/10 ARC-AGI Evolution and the Limits of Test-Time Compute Mike explains the mechanics of ARC-AGI 1 and ARC-AGI 2, detailing how recent models like o3 and test-time compute fail to generalize efficiently on novel benchmarks. Jordy and the host interject with questions about capital constraints versus algorithmic bottlenecks.6:55–11:50 · Guest teaching 4/10 Decentralizing Frontier AI Research and the Vesuvius Challenge The host draws a strong parallel between ARC Prize and the Vesuvius Challenge to decentralized problem solving. The host then asks technical questions about Python execution and overfitting guardrails, which Mike clarifies by explaining Kaggle compute constraints and private test splits.11:50–15:08 · Guest teaching 5/10 Overcoming Reliability Bottlenecks to Build Consumer AI Agents Jordy asks why consumer agent experiences like flight booking remain underwhelming across major tech companies. Mike lays out his framework of concentric rings of reliability risk and explains how fluid intelligence on ARC translates directly to robust agentic execution.15:08–18:19 · Guest teaching 6/10 Defining AGI Through Human-AI Task Gaps Versus Economics The host proposes measuring AGI by macroeconomic contribution thresholds such as percentage of global GDP. Mike firmly pushes back against this definition, arguing that economic metrics obscure underlying capability mechanisms like memorization versus true adaptation.18:19–26:56 · Guest teaching 3/10 Viral Ghibli AI Models and Novel Architectural Exploration The host demonstrates sharp domain expertise by sharing an A/B social media test comparing a real Ghibli still with an AI-generated scene, explaining the uncanny valley dynamics of artistic style transfer. Mike connects this to non-diffusion architectural exploration and warns startups against brute-force model pre-training.26:56–29:30 · Guest teaching 5/10 Reevaluating Gary Marcus, Scaling Limits, and the Bitter Lesson The host articulates an informed synthesis between Gary Marcus's critiques of pure deep learning and Rich Sutton's Bitter Lesson. Mike agrees Marcus was largely vindicated on generalization limits and reinterprets Sutton's paper by highlighting that scalable architectures still originate from human ingenuity.29:30–33:32 · Guest teaching 7/10 Program Synthesis Pioneers and the Enduring Value of Coding Mike highlights key researchers in program synthesis before the host asks if someone could create an ARC-like challenge tailored specifically for human programmers. Mike gently educates the host by revealing that ARC is fundamentally already a program synthesis challenge operating on discrete numeric matrices.0:00–2:14 · Guest disagreement 1/10 Mike Knoop on Zapier, LLM Reliability, and ARC Prize The host warmly opens the episode and asks Mike Knoop to introduce his background at Zapier and the genesis of ARC Prize. Mike explains his realization that enterprise customers found LLMs too unreliable for unsupervised automation.2:14–6:55 · Guest disagreement 2/10 ARC-AGI Evolution and the Limits of Test-Time Compute Mike explains the mechanics of ARC-AGI 1 and ARC-AGI 2, detailing how recent models like o3 and test-time compute fail to generalize efficiently on novel benchmarks. Jordy and the host interject with questions about capital constraints versus algorithmic bottlenecks.6:55–11:50 · Guest disagreement 1/10 Decentralizing Frontier AI Research and the Vesuvius Challenge The host draws a strong parallel between ARC Prize and the Vesuvius Challenge to decentralized problem solving. The host then asks technical questions about Python execution and overfitting guardrails, which Mike clarifies by explaining Kaggle compute constraints and private test splits.11:50–15:08 · Guest disagreement 1/10 Overcoming Reliability Bottlenecks to Build Consumer AI Agents Jordy asks why consumer agent experiences like flight booking remain underwhelming across major tech companies. Mike lays out his framework of concentric rings of reliability risk and explains how fluid intelligence on ARC translates directly to robust agentic execution.15:08–18:19 · Guest disagreement 4/10 Defining AGI Through Human-AI Task Gaps Versus Economics The host proposes measuring AGI by macroeconomic contribution thresholds such as percentage of global GDP. Mike firmly pushes back against this definition, arguing that economic metrics obscure underlying capability mechanisms like memorization versus true adaptation.18:19–26:56 · Guest disagreement 2/10 Viral Ghibli AI Models and Novel Architectural Exploration The host demonstrates sharp domain expertise by sharing an A/B social media test comparing a real Ghibli still with an AI-generated scene, explaining the uncanny valley dynamics of artistic style transfer. Mike connects this to non-diffusion architectural exploration and warns startups against brute-force model pre-training.26:56–29:30 · Guest disagreement 2/10 Reevaluating Gary Marcus, Scaling Limits, and the Bitter Lesson The host articulates an informed synthesis between Gary Marcus's critiques of pure deep learning and Rich Sutton's Bitter Lesson. Mike agrees Marcus was largely vindicated on generalization limits and reinterprets Sutton's paper by highlighting that scalable architectures still originate from human ingenuity.29:30–33:32 · Guest disagreement 2/10 Program Synthesis Pioneers and the Enduring Value of Coding Mike highlights key researchers in program synthesis before the host asks if someone could create an ARC-like challenge tailored specifically for human programmers. Mike gently educates the host by revealing that ARC is fundamentally already a program synthesis challenge operating on discrete numeric matrices.0:00–2:14 · The hosts pushing back 0/10 Mike Knoop on Zapier, LLM Reliability, and ARC Prize The host warmly opens the episode and asks Mike Knoop to introduce his background at Zapier and the genesis of ARC Prize. Mike explains his realization that enterprise customers found LLMs too unreliable for unsupervised automation.2:14–6:55 · The hosts pushing back 1/10 ARC-AGI Evolution and the Limits of Test-Time Compute Mike explains the mechanics of ARC-AGI 1 and ARC-AGI 2, detailing how recent models like o3 and test-time compute fail to generalize efficiently on novel benchmarks. Jordy and the host interject with questions about capital constraints versus algorithmic bottlenecks.6:55–11:50 · The hosts pushing back 1/10 Decentralizing Frontier AI Research and the Vesuvius Challenge The host draws a strong parallel between ARC Prize and the Vesuvius Challenge to decentralized problem solving. The host then asks technical questions about Python execution and overfitting guardrails, which Mike clarifies by explaining Kaggle compute constraints and private test splits.11:50–15:08 · The hosts pushing back 1/10 Overcoming Reliability Bottlenecks to Build Consumer AI Agents Jordy asks why consumer agent experiences like flight booking remain underwhelming across major tech companies. Mike lays out his framework of concentric rings of reliability risk and explains how fluid intelligence on ARC translates directly to robust agentic execution.15:08–18:19 · The hosts pushing back 2/10 Defining AGI Through Human-AI Task Gaps Versus Economics The host proposes measuring AGI by macroeconomic contribution thresholds such as percentage of global GDP. Mike firmly pushes back against this definition, arguing that economic metrics obscure underlying capability mechanisms like memorization versus true adaptation.18:19–26:56 · The hosts pushing back 1/10 Viral Ghibli AI Models and Novel Architectural Exploration The host demonstrates sharp domain expertise by sharing an A/B social media test comparing a real Ghibli still with an AI-generated scene, explaining the uncanny valley dynamics of artistic style transfer. Mike connects this to non-diffusion architectural exploration and warns startups against brute-force model pre-training.26:56–29:30 · The hosts pushing back 2/10 Reevaluating Gary Marcus, Scaling Limits, and the Bitter Lesson The host articulates an informed synthesis between Gary Marcus's critiques of pure deep learning and Rich Sutton's Bitter Lesson. Mike agrees Marcus was largely vindicated on generalization limits and reinterprets Sutton's paper by highlighting that scalable architectures still originate from human ingenuity.29:30–33:32 · The hosts pushing back 1/10 Program Synthesis Pioneers and the Enduring Value of Coding Mike highlights key researchers in program synthesis before the host asks if someone could create an ARC-like challenge tailored specifically for human programmers. Mike gently educates the host by revealing that ARC is fundamentally already a program synthesis challenge operating on discrete numeric matrices.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 0:16 Mike dismantles economic definitions of AGI

Mike directly rejects the host's proposed GDP threshold metric, asserting that economic output conflates pattern memorization with novel problem-solving capability.

Hardest push from the hosts ▶ 0:26 Host presses on the tension between Marcus and Sutton

The host challenges the prevailing industry consensus by demanding a reconciliation between Gary Marcus's symbolic AI skepticism and the Bitter Lesson's scale-maximalist doctrine.

Biggest teaching moment ▶ 0:31 Mike clarifies ARC's underlying programmatic nature

When the host proposes creating a coding-specific ARC benchmark, Mike corrects the premise by explaining that ARC tasks are already pure program synthesis problems evaluated as 2D numeric matrices.

The host holds their own ▶ 0:19 Host presents viral empirical data on visual style transfer

The host commands the dialogue by citing his own controlled viral posting experiment to break down why Ghibli-style generation succeeds aesthetically where VFX photo-realism fails.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Mike Knoop on Zapier, LLM Reliability, and ARC Prize 3210 The host warmly opens the episode and asks Mike Knoop to introduce his background at Zapier and the genesis of ARC Prize. Mike explains his realization that enterprise customers found LLMs too unreliable for unsupervised automation.
ARC-AGI Evolution and the Limits of Test-Time Compute 4621 Mike explains the mechanics of ARC-AGI 1 and ARC-AGI 2, detailing how recent models like o3 and test-time compute fail to generalize efficiently on novel benchmarks. Jordy and the host interject with questions about capital constraints versus algorithmic bottlenecks.
Decentralizing Frontier AI Research and the Vesuvius Challenge 6411 The host draws a strong parallel between ARC Prize and the Vesuvius Challenge to decentralized problem solving. The host then asks technical questions about Python execution and overfitting guardrails, which Mike clarifies by explaining Kaggle compute constraints and private test splits.
Overcoming Reliability Bottlenecks to Build Consumer AI Agents 4511 Jordy asks why consumer agent experiences like flight booking remain underwhelming across major tech companies. Mike lays out his framework of concentric rings of reliability risk and explains how fluid intelligence on ARC translates directly to robust agentic execution.
Defining AGI Through Human-AI Task Gaps Versus Economics 5642 The host proposes measuring AGI by macroeconomic contribution thresholds such as percentage of global GDP. Mike firmly pushes back against this definition, arguing that economic metrics obscure underlying capability mechanisms like memorization versus true adaptation.
Viral Ghibli AI Models and Novel Architectural Exploration 7321 The host demonstrates sharp domain expertise by sharing an A/B social media test comparing a real Ghibli still with an AI-generated scene, explaining the uncanny valley dynamics of artistic style transfer. Mike connects this to non-diffusion architectural exploration and warns startups against brute-force model pre-training.
Reevaluating Gary Marcus, Scaling Limits, and the Bitter Lesson 7522 The host articulates an informed synthesis between Gary Marcus's critiques of pure deep learning and Rich Sutton's Bitter Lesson. Mike agrees Marcus was largely vindicated on generalization limits and reinterprets Sutton's paper by highlighting that scalable architectures still originate from human ingenuity.
Program Synthesis Pioneers and the Enduring Value of Coding 4721 Mike highlights key researchers in program synthesis before the host asks if someone could create an ARC-like challenge tailored specifically for human programmers. Mike gently educates the host by revealing that ARC is fundamentally already a program synthesis challenge operating on discrete numeric matrices.

Statements from this episode (14)

Insight
Knoop: AI Agent Failure Rates Make Unsupervised Automation Unviable
“Like, hey, I get the hype, but like, they just, they're not reliable enough yet. You know, they don't work two out of 10 times, and that just doesn't work for these unsupervised automation products.”
Mike Knoop Apr 6, 2025 ▶ 1:02
Assertion Supported
Knoop: ARC Saw No Progress Despite 50,000x Model Scaling
“Surprise that it basically hadn't, and not only hadn't been beaten, there'd basically been no progress in it which I thought was really fascinating given the fact that we've like scaled up these language model systems by almost like 50,000 times over the last,…”
Mike Knoop Apr 6, 2025 ▶ 1:31
Assertion Supported
Knoop: Pure LLMs Score 0% and o1 Scores 1% on ARC-AGI-2
“Pure LLM systems are scoring like zero percent now again on on arc B two single COT systems like R one and O one score like one percent.”
Mike Knoop Apr 6, 2025 ▶ 6:01
Prediction Not checkable as stated
Knoop: Scaling Test-Time Compute Will Not Get Us to AGI
“And then there's a new story that's emerged over the last like five months, which is, oh, we're going to scale up this test time compute and that's going to get us to AGI. And I think what V two shows is that that's not quite either. We still need some structu…”
Mike Knoop Apr 6, 2025 ▶ 6:40
Insight
Knoop: ARC proves individuals can still advance frontier AI
“And I think arc shows that like, yeah, there actually are frontier problems. That are unsolved, that individual people and individual teams can actually make a difference on today.”
Mike Knoop Apr 6, 2025 ▶ 8:13
Insight
Knoop: AI intelligence benchmarks must measure compute efficiency, not brute force
“We do think that efficiency is actually a really, really important aspect of intelligence. You know, you can brute force your way up to intelligence, but we really do want to be shooting for like human targets and efficiency for this stuff.”
Mike Knoop Apr 6, 2025 ▶ 9:55
Assertion Supported
Knoop: OpenAI asked ARC Prize team to verify results on semi-private dataset
“Opening eye situation was a little different because they had reached out to us and said, Hey, we think we've got a really impressive result on the public eval set. And we'd like your help to verify. On the semi-private set, which is what we created that data …”
Mike Knoop Apr 6, 2025 ▶ 11:12
Prediction Not checkable as stated
Knoop: AI agents will start working in 2025 due to ARC progress
“I actually think we're going to start to see agents start to work this year specifically because of progress on arc.”
Mike Knoop Apr 6, 2025 ▶ 14:16
Opinion
Knoop: Language models operate by memorization rather than solving novel patterns
“Language models. Generally working like a memorization style regime where they're right. Learning lots of data. They're able to apply it to very similar types of patterns that they've seen before, but not novel patterns. That's what RKGI shows.”
Mike Knoop Apr 6, 2025 ▶ 17:32
Insight
Knoop: Startups just doing model training are lighting money on fire
“I think anyone who's like Just doing model training at this point is like lighting money on fire. If you really want to make a unique difference, especially if you're a small startup, like a founder, like you gotta go take an orthogonal approach. You gotta try…”
Mike Knoop Apr 6, 2025 ▶ 25:37
Assertion Supported
Knoop: Zapier never spent its $1M seed round from 2012
“It was about a million bucks back in 2012. We never spent the money. By the time we actually got the round closed, figured out who we wanted to hire, got the payroll started, like revenue had caught up. And so literally I think you could trace every dollar we …”
Mike Knoop Apr 6, 2025 ▶ 26:22
Opinion
Knoop: Gary Marcus has been more right than wrong on deep learning
“I generally think he's been more right than wrong. I think if you like just take a limited five year view on this from 20, 20 up until 20, 20, end of 20, 24, you know, I think it was a generally right. Like he was making the right ideas.”
Mike Knoop Apr 6, 2025 ▶ 27:38
Opinion
Knoop: Achieving AGI requires merging deep learning with program synthesis
“I actually don't think either is sufficient. I think some merger of the two is what's necessary to get to AGI.”
Mike Knoop Apr 6, 2025 ▶ 30:11
Opinion
Knoop: People should still learn to code for technological leverage
“My hot take is, like, I guess, yes, you should still learn to code. Primarily because it's been, it gives you like leverage over technology today. And yeah, like I don't see that leverage over technology going away anytime soon, particularly if you want to wor…”
Mike Knoop Apr 6, 2025 ▶ 32:37
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 500 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.