Coogan: GPT-6 Astra achieves 99.9% on ARC-AGI-3 benchmark
“The RKGI three number is crazy. The score is 99.9%, so it feels like they just beat that. They just beat RKGI B three.”
Chollet: ARC Prize Built a Video Game Studio Generating 250+ Games
“We set up an entire video game studio, right, to create them. So we got over 250 games.”
Chollet: RL Benchmarks Like Dota and Atari Test Memorization, Not Intelligence
“If you look at Atari games, for instance, or even Dota, you're training on, on the same environment as what you use for testing. So effectively, you're just trying to memorize the best strategies. You're trying to at training time, explore the full space of po…”
Kamradt Announces ARC-AGI-3 Will Feature 100 Novel Game Environments
“We're coming out with RKGI three. And what this is gonna be is it's gonna be a series of a hundred different novel environments, or you could simply call them a hundred different novel games that we're making ourselves.”
Kamradt: ARC-AGI-3 Preview Launches Five Games
“So as a part of the preview, we're launching five games. Now, three of them are going to be public on day one, and two of them are going to be private.”
Kamradt: ARC-AGI-3 Provides AI Agents With a 64x64 JSON Grid
“So we'll show the same thing to AI, except that AI is gonna get a JSON grid list of lists. So those get a bunch of numbers, 64 by 64, and they can choose to turn that into an image if they want to, or agnostic, do whatever you want with it if they want to do m…”
Kamradt: ARC-AGI-3 Agents Interact via 64x64 Frames and Integer Actions
“What agents will get is agents will get a series of frames and those frames will be 64 by 64. Now generally it's just going to be one frame, but you might be able to get like maybe two in a row or three in a row, and that would show an animation. And so beginn…”
Kamradt: No AI Has Beaten Any ARC-AGI-3 Game Level Yet
“It's still true. We have yet to have an AI successfully beat any level on any of these games. So it hasn't happened yet.”
Kamradt: Random Brute Force Agent Fails ARC-AGI-3 Locksmith Game
“One of the quality checks that we do is we run a random agent at a million steps to see if it beats it or not. And no, it doesn't beat lockstep at all or locksmith.”
Kamradt: Action Step Count is Core Metric for AI Learning Efficiency
“When we report learning efficiency for this, especially with AI versus humans, it's all going to be around how many actions do you take in order to complete the goal of the environment, which not only does that encompass learning what the environment entails, …”
Kamradt: ARC-AGI-3 Features Large Percentage of Non-Agent Puzzle Games
“We have a requirement that games must be novel from each other. We have
A large percentage of games that are non-agent based. So think of it as like solitaire or connect four or like Simon or memory or something like that. Those are non-agent based games.”
Kamradt: ARC-AGI-3 Targets 120 Benchmark Games by Q1 2026
“Our goal is to come out with a 120 by Q one of next year.”
Kamradt: Human-Built Benchmarks Prevent AI From Reverse-Engineering Generation Code
“And the problem with that is that we don't want to incentivize AI to derive the program that made the game. Right? And so if we continue to have humans make the game, then the AI is incentivized to try to reverse engineer the G inside of humans, and that's kin…”
Kamradt Predicts ARC-AGI-3 Benchmark Will Remain Unbeaten For 3 Years
“And then V three, our durability estimate for that is three years. And that's what we're aiming for is 36 months for V three.”
Chollet: Work has begun on ARC-AGI-3 featuring a brand-new format
“We're already starting to work on version three which you have a brand new format.”