Jul 18, 2025 · 39m · latent-space
⚡️ARC-AGI-3: The Interactive Reasoning Benchmark
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Greg Kamradt from the ARC Prize Foundation joins the Latent Space Lightning Pod to announce ARC-AGI-3, an interactive reasoning benchmark comprising 100 novel 2D game environments designed to measure sample-efficient skill acquisition. The discussion covers the philosophical definition of general intelligence, technical agent specifications, the foundation's human-designed pipeline, and recent frontier model evaluations including xAI's Grok 4.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 6.7% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Greg forcefully pushes back on swyx's suggestion that AGI is a continuous spectrum, stating that any definition involving profit has ulterior motives and insisting AGI is a binary threshold based on human efficiency.
Hardest push from the hosts ▶ 30:45 Challenging ARC's moving goalpostsswyx refuses the premise that declaring a single AGI moment is a worthwhile goal, arguing that ARC is constantly moving the goalposts across versions while labs like OpenAI treat AGI as continuous levels.
Biggest teaching moment ▶ 6:07 Developer bias in RL simulation environmentsGreg educates the hosts on why scaling RL in custom environments does not equal generalization, explaining that developers inadvertently inject their own intelligence into the simulated reward structure.
The host holds their own ▶ 16:43 Citing Noam Brown on agent scaffoldingswyx demonstrates deep domain expertise by challenging Greg's belief in complex harnesses, quoting researcher Noam Brown to argue that agent scaffolding will inevitably be made obsolete by frontier models.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Origins, Mission, and Growth of ARC Prize Foundation | 5 | 5 | 1 | 1 | The hosts set up the background on ARC Prize gaining mainstream recognition over the past year. Greg explains the origins from François Chollet's 2019 paper to Mike Knoop funding the $1M bounty and turning it into a foundation. | |
| Defining Intelligence: Skill Acquisition Efficiency and Denominators | 5 | 7 | 2 | 1 | swyx asks how Chollet defines AGI and intelligence acquisition. Greg breaks down the core thesis: intelligence is skill acquisition efficiency, defined by the denominators of energy and training data compared to the human brain. | |
| The Transition to Interactive Reasoning in ARC-AGI-3 | 6 | 6 | 2 | 4 | Alessio steelmans the counterargument that compute costs are dropping and RL can simply be parallelized across domains. Greg explains that simulated RL environments typically have developer intelligence injected into them rather than general model intelligence. | |
| Interactive Demonstration of the Locksmith Game | 4 | 5 | 1 | 1 | Greg shares his screen and plays through the Locksmith game demo. The hosts observe and comment as Greg illustrates how rules, energy constraints, and rotation mechanics force exploration and long-term planning. | |
| Technical Specifications and Agent Action Interface | 7 | 5 | 2 | 4 | Alessio and swyx probe the API interface, text vs. multimodal representations, and scaffolding. swyx pushes back with prior findings that vision does not aid ARC and cites Noam Brown's view that agent scaffolds will eventually become obsolete. | |
| Evaluating Efficiency: Cost and Action Step Metrics | 6 | 6 | 1 | 2 | Alessio asks how ARC accounts for model price fluctuations when measuring cost efficiency. Greg explains that while cost acts as a proxy for closed models, ARC-AGI-3 introduces action step efficiency as the primary metric. | |
| Game Diversity, Non-Agent Formats, and Cooperative Mechanics | 6 | 6 | 2 | 4 | swyx critiques the benchmark for appearing strictly embodied and single-agent. Greg clarifies that many ARC-3 tasks are non-agent board games or require cooperative alignment, explaining the human-in-the-loop game construction pipeline. | |
| Foundation Team Operations, Funding, and Hiring | 5 | 6 | 1 | 2 | Alessio inquires about foundation staffing and hiring needs, while swyx asks about the future roadmap beyond ARC-3. Greg emphasizes keeping human solvability as their core operational anchor instead of designing esoteric PhD-level tests. | |
| Benchmark Durability and Frontier Model Projections | 6 | 6 | 5 | 5 | Alessio asks about benchmark durability and Grok 4's 16% score. swyx challenges the premise of declaring a single AGI moment, pointing out shifting goalposts. Greg firmly counters, dismissing revenue-based AGI definitions and insisting on a binary learning-efficiency milestone. | |
| Inside the xAI Grok 4 Launch Event and Meeting Elon Musk | 4 | 6 | 1 | 2 | swyx asks about the atmosphere at the Grok 4 launch event. Greg shares behind-the-scenes details of validating xAI's evaluation scores and pitching ARC-AGI-3's video game paradigm directly to Elon Musk. | |
| Assessing Grok 4 Capabilities and Frontier Model Trends | 6 | 5 | 1 | 1 | swyx and Greg evaluate Grok 4's frontier benchmark performance, discussing RL scaling, rumors surrounding dedicated coding models, and competing lab dynamics before Greg closes with a call for agent developers. |