ARC AGI 3

product on 4 shows · 15 statements across 4 episodes · said 40 times in 6 episodes since 2025

Latent Space 19 the Y Combinator Startup Podcast 12 TBPN 9 the MAD Podcast

Mentions by year, every show

tap a year for its mentions
0010220420252026episodesmentions
02420252026episodes it came up in
005210420252026episodesmentions per episode

Latent Space 19the Y Combinator Startup Podcast 12TBPN 9

2026 20 mentions in 4 episodes 5 per episode
2025 20 mentions in 2 episodes 10 per episode

every mention on every show, scene by scene, with the transcript →

15 statements about ARC AGI 3, every show

TBPN Assertion Supported
Coogan: GPT-6 Astra achieves 99.9% on ARC-AGI-3 benchmark
“The RKGI three number is crazy. The score is 99.9%, so it feels like they just beat that. They just beat RKGI B three.”
John Coogan Sep 4, 2026 ▶ 26:40 Model Mayhem, GPT-6 Astra, Why Nvidia Bought Hugging Face | Diet TBPN
Y COMBINATOR Disclosure
Chollet: ARC Prize Built a Video Game Studio Generating 250+ Games
“We set up an entire video game studio, right, to create them. So we got over 250 games.”
François Chollet Mar 27, 2026 ▶ 29:39 François Chollet: Why Scaling Alone Isn’t Enough for AGI · Y Combinator
Chollet: RL Benchmarks Like Dota and Atari Test Memorization, Not Intelligence
“If you look at Atari games, for instance, or even Dota, you're training on, on the same environment as what you use for testing. So effectively, you're just trying to memorize the best strategies. You're trying to at training time, explore the full space of po…”
François Chollet Mar 27, 2026 ▶ 33:08 François Chollet: Why Scaling Alone Isn’t Enough for AGI · Y Combinator
LATENT SPACE Disclosure
Kamradt Announces ARC-AGI-3 Will Feature 100 Novel Game Environments
“We're coming out with RKGI three. And what this is gonna be is it's gonna be a series of a hundred different novel environments, or you could simply call them a hundred different novel games that we're making ourselves.”
Greg Kamradt Jul 18, 2025 ▶ 7:05 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Disclosure
Kamradt: ARC-AGI-3 Preview Launches Five Games
“So as a part of the preview, we're launching five games. Now, three of them are going to be public on day one, and two of them are going to be private.”
Greg Kamradt Jul 18, 2025 ▶ 8:34 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Disclosure
Kamradt: ARC-AGI-3 Provides AI Agents With a 64x64 JSON Grid
“So we'll show the same thing to AI, except that AI is gonna get a JSON grid list of lists. So those get a bunch of numbers, 64 by 64, and they can choose to turn that into an image if they want to, or agnostic, do whatever you want with it if they want to do m…”
Greg Kamradt Jul 18, 2025 ▶ 9:11 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Disclosure
Kamradt: ARC-AGI-3 Agents Interact via 64x64 Frames and Integer Actions
“What agents will get is agents will get a series of frames and those frames will be 64 by 64. Now generally it's just going to be one frame, but you might be able to get like maybe two in a row or three in a row, and that would show an animation. And so beginn…”
Greg Kamradt Jul 18, 2025 ▶ 12:42 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Assertion Supported
Kamradt: No AI Has Beaten Any ARC-AGI-3 Game Level Yet
“It's still true. We have yet to have an AI successfully beat any level on any of these games. So it hasn't happened yet.”
Greg Kamradt Jul 18, 2025 ▶ 14:33 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Assertion Supported
Kamradt: Random Brute Force Agent Fails ARC-AGI-3 Locksmith Game
“One of the quality checks that we do is we run a random agent at a million steps to see if it beats it or not. And no, it doesn't beat lockstep at all or locksmith.”
Greg Kamradt Jul 18, 2025 ▶ 18:28 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Kamradt: Action Step Count is Core Metric for AI Learning Efficiency
“When we report learning efficiency for this, especially with AI versus humans, it's all going to be around how many actions do you take in order to complete the goal of the environment, which not only does that encompass learning what the environment entails, …”
Greg Kamradt Jul 18, 2025 ▶ 19:14 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Disclosure
Kamradt: ARC-AGI-3 Features Large Percentage of Non-Agent Puzzle Games
“We have a requirement that games must be novel from each other. We have A large percentage of games that are non-agent based. So think of it as like solitaire or connect four or like Simon or memory or something like that. Those are non-agent based games.”
Greg Kamradt Jul 18, 2025 ▶ 20:07 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Disclosure
Kamradt: ARC-AGI-3 Targets 120 Benchmark Games by Q1 2026
“Our goal is to come out with a 120 by Q one of next year.”
Greg Kamradt Jul 18, 2025 ▶ 22:45 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
Kamradt: Human-Built Benchmarks Prevent AI From Reverse-Engineering Generation Code
“And the problem with that is that we don't want to incentivize AI to derive the program that made the game. Right? And so if we continue to have humans make the game, then the AI is incentivized to try to reverse engineer the G inside of humans, and that's kin…”
Greg Kamradt Jul 18, 2025 ▶ 23:07 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
LATENT SPACE Prediction Open · timeframe Dec 2029
Kamradt Predicts ARC-AGI-3 Benchmark Will Remain Unbeaten For 3 Years
“And then V three, our durability estimate for that is three years. And that's what we're aiming for is 36 months for V three.”
Greg Kamradt Jul 18, 2025 ▶ 28:35 ⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
MAD Disclosure
Chollet: Work has begun on ARC-AGI-3 featuring a brand-new format
“We're already starting to work on version three which you have a brand new format.”
Francois Chollet Apr 3, 2025 ▶ 50:31 Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.