ARC AGI

product on 9 shows · 17 statements across 15 episodes · said 232 times in 48 episodes since 2024

TBPN 94 the MAD Podcast 40 the Y Combinator Startup Podcast 38 Latent Space 29 Big Technology 15 No Priors 10 All-In 4 Invest Like the Best 1 20VC 1

Mentions by year, every show

tap a year for its mentions
001002020040202420252026episodesmentions
02040202420252026episodes it came up in
00320640202420252026episodesmentions per episode

TBPN 94the MAD Podcast 40the Y Combinator Startup Podcast 38Latent Space 29Big Technology 15No Priors 10All-In 420VC 11 more shows

2026 44 mentions in 12 episodes 4 per episode
2025 170 mentions in 32 episodes 5 per episode
2024 18 mentions in 4 episodes 5 per episode

every mention on every show, scene by scene, with the transcript →

17 statements about ARC AGI, every show

Kantrowitz: AI benchmark success without economic impact reveals spiky intelligence
“If it can solve arc AGI and it's not necessarily crushing on these economic factors and these just kind of general rote work things that we would like it to do it shows that instead of being general, it's very spiky intelligence and hence much less useful.”
Alex Kantrowitz Sep 7, 2026 ▶ 18:58 GPT-6 & OpenAI’s Comeback, Hugging Face Attack Debate, Ballmer’s Scandalous Legacy
Y COMBINATOR Assertion Supported
Hu: Anthropic Opus 5 achieved 30% on ARC-AGI
“You guys got, took Arc AGI three to 30%, which is incredible.”
Diana Hu Jul 27, 2026 ▶ 0:33 Boris Cherny: We Cut 80% of Claude Code’s Prompt · Y Combinator
Y COMBINATOR Assertion Supported
Chollet: Base LLMs Score Under 10% on ARC-AGI-1
“So basal alarms were scoring extremely low on V-one, like sub-ten percent, basically. And, I mean, it was true of the original, like, GPT-III actually scoring zero, but that's even true of the latest basal alarms today, you know, as of March.”
François Chollet Mar 27, 2026 ▶ 17:57 François Chollet: Why Scaling Alone Isn’t Enough for AGI · Y Combinator
Y COMBINATOR Assertion Not checkable as stated
Poetiq's autonomous prompt generation system produced unexpected, non-human prompt structures
“It was pretty interesting to look at the prompt outputs in particular, I'd say, for ArcGi in that you know, I think you can read those and say, well, that's not what a human would have written. Pretty clearly. And it's, you know, there's some unexpected stuff …”
Ian Fisher Feb 27, 2026 ▶ 12:20 The Powerful Alternative To Fine-Tuning · Y Combinator
BIG TECHNOLOGY Assertion Supported
Kantrowitz: Gemini 3 crushed the ARC-AGI benchmark and topped Chatbot Arena
“Gemini three smashes the benchmarks crushed on the Arc AGI test. It's currently at the top of the LM arena leaderboards”
Alex Kantrowitz Nov 24, 2025 ▶ 2:21 Google Pushes OpenAI, Bezos Returns, AI’s No. 1 Hit
20VC Prediction Not checkable as stated
Frosst: Enterprise clients will never demand Arc AGI pixel manipulation features
“Stuff like the Arc AGI challenge is a benchmark that people talk about, but that's like a pixel manipulation challenge. It's like, you know, taking in like a grid of pixels and based on rules, predicting the next one. That's not a thing any of our customers ha…”
Nick Frosst Sep 1, 2025 ▶ 16:56 Cohere Founder, Nick Frosst: How To Compete with OpenAI & Anthropic, and Sam Altman’s AI Disservice · 20VC with Harry Stebbings
Lambert: AI benchmarks like ARC-AGI should prioritize testing without harnesses
“Harnesses are cool, but they're gonna, they're, They're a handicap that's changing the learning dynamics substantially. So it's good. It's good demos, but I feel like the core thrust has to be no harnesses.”
Nathan Lambert Jul 31, 2025 ▶ 34:02 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
LATENT SPACE Assertion Supported
Kamradt: Mike Knoop Put Up $1M for ARC Prize Bounty
“Mike actually put he put up a million dollars of his own money and said, Hey, I'm going to put a bounty. So for anybody who can beat this benchmark, they're going to get a million dollars.”
Greg Kamradt Jul 18, 2025 ▶ 1:50 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
BIG TECHNOLOGY Assertion Supported
Kantrowitz: Grok 4 beats all models on ARC-AGI by significant margin
“And then of course in the Arc AGI test, it outperforms Every model by a significant margin.”
Alex Kantrowitz Jul 14, 2025 ▶ 12:50 OpenAI’s Windsurf Crash, Grok’s Wild Week, Replace Tim Cook? — With Aaron Levie
TBPN Disclosure
Douglas: Frontier Labs Avoid RL Training Directly on ARC-AGI
“And I mean, I think if you are old on Arc AGI, then it would, you'd probably get superhuman at it pretty fast. But I think we're all trying not to RL on it so that it functions as like an interesting held out.”
Sholto Douglas Jun 7, 2025 ▶ 1:55:15 Weekly Recap - Elon Vs Trump, Ukraine's Drone Attack, Cluely Update & OpenAI CRO
TBPN Opinion
Knoop: Language models operate by memorization rather than solving novel patterns
“Language models. Generally working like a memorization style regime where they're right. Learning lots of data. They're able to apply it to very similar types of patterns that they've seen before, but not novel patterns. That's what RKGI shows.”
Mike Knoop Apr 6, 2025 ▶ 17:32 Mike Knoop (Arc Prize) on Why Scaling AI Won’t Get Us to AGI
MAD Assertion Not checkable as stated
Chollet: OpenAI o3 cost $10k–$20k per ARC puzzle on maximum compute
“For instance OpenAI O.S. On the highest compute settings that we tried it on for Arc, it was consuming somewhere between, like, 10,000 dollars to 20,000 dollars per task, like, for one little puzzle, which you could normally solve with a base of an API for a f…”
Francois Chollet Apr 3, 2025 ▶ 20:00 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
TBPN Assertion Supported
Coogan: OpenAI's o3 High-Compute Mode Spent $3,000 to Solve a Benchmark Task
“Oh three, which isn't out yet, but is even more advanced in terms of reasoning. They have a high compute model. That spends almost 3000 dollars per task. And it just thinks for hours and hours and hours basically, and it was able to break arc that that AI, AGI…”
John Coogan Jan 28, 2025 ▶ 19:38 DeepSeek Update, Market Crash, Timeline in Turmoil, Is VC Cooked, Zero Cope Policy
TBPN Assertion Supported
Coogan: OpenAI o3 high-compute configuration costs $2,000 per solve
“O-three is a reasoning model. The high version costs 2000 dollars per solve.”
John Coogan Dec 31, 2024 ▶ 1:03:39 Year of Tech in Review (ep25)
NO PRIORS Opinion
Knoop: ARC-AGI is the only true AGI evaluation that exists
“Arc AGI to best of my knowledge is the only true. AGI eval that actually exists in the world and measures a actually good definition, correct definition of what AGI is, which we can talk about.”
Mike Knoop Jun 11, 2024 ▶ 2:28 No Priors Ep. 68 | With Zapier Co-Founder and Head of AI Mike Knoop
NO PRIORS Assertion Supported
Knoop: ARC-AGI benchmark performance only moved from 20% to 34% in four years
“There's an AI lab called lab 42 out of Switzerland that's been running a small annual contest over the last four years to try and beat this eval and state of the art today is. 34% state of the art four years ago when it was first introduced was 20%. So we've m…”
Mike Knoop Jun 11, 2024 ▶ 2:38 No Priors Ep. 68 | With Zapier Co-Founder and Head of AI Mike Knoop
NO PRIORS Prediction Open · timeframe Jun 2029
Knoop: ARC solution will likely need under 10k code lines, not massive LLMs
“It's quite likely actually that the solution it can be like written in like 10,000 lines of code or less. And it's not gonna require these like, you know, gigantic You know, two hundred billion large parameter models in order to solve it.”
Mike Knoop Jun 11, 2024 ▶ 15:15 No Priors Ep. 68 | With Zapier Co-Founder and Head of AI Mike Knoop

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.