Jul 18, 2025 · 39m · latent-space

⚡️ARC-AGI-3: The Interactive Reasoning Benchmark

Greg Kamradt · 27m spoken Shawn Wang · 5m spoken Alessio Fanelli · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Greg Kamradt from the ARC Prize Foundation joins the Latent Space Lightning Pod to announce ARC-AGI-3, an interactive reasoning benchmark comprising 100 novel 2D game environments designed to measure sample-efficient skill acquisition. The discussion covers the philosophical definition of general intelligence, technical agent specifications, the foundation's human-designed pipeline, and recent frontier model evaluations including xAI's Grok 4.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 6.7% of the talking time here. How this is scored →

The hosts as informed peer 5.5 Guest teaching 5.7 Guest disagreement 1.7 The hosts pushing back 2.5
05100:0010:0020:0030:001:04–3:18 · The hosts as informed peer 5/10 Origins, Mission, and Growth of ARC Prize Foundation The hosts set up the background on ARC Prize gaining mainstream recognition over the past year. Greg explains the origins from François Chollet's 2019 paper to Mike Knoop funding the $1M bounty and turning it into a foundation.3:19–5:35 · The hosts as informed peer 5/10 Defining Intelligence: Skill Acquisition Efficiency and Denominators swyx asks how Chollet defines AGI and intelligence acquisition. Greg breaks down the core thesis: intelligence is skill acquisition efficiency, defined by the denominators of energy and training data compared to the human brain.5:37–8:02 · The hosts as informed peer 6/10 The Transition to Interactive Reasoning in ARC-AGI-3 Alessio steelmans the counterargument that compute costs are dropping and RL can simply be parallelized across domains. Greg explains that simulated RL environments typically have developer intelligence injected into them rather than general model intelligence.8:03–12:22 · The hosts as informed peer 4/10 Interactive Demonstration of the Locksmith Game Greg shares his screen and plays through the Locksmith game demo. The hosts observe and comment as Greg illustrates how rules, energy constraints, and rotation mechanics force exploration and long-term planning.12:23–17:02 · The hosts as informed peer 7/10 Technical Specifications and Agent Action Interface Alessio and swyx probe the API interface, text vs. multimodal representations, and scaffolding. swyx pushes back with prior findings that vision does not aid ARC and cites Noam Brown's view that agent scaffolds will eventually become obsolete.17:03–19:31 · The hosts as informed peer 6/10 Evaluating Efficiency: Cost and Action Step Metrics Alessio asks how ARC accounts for model price fluctuations when measuring cost efficiency. Greg explains that while cost acts as a proxy for closed models, ARC-AGI-3 introduces action step efficiency as the primary metric.19:31–23:28 · The hosts as informed peer 6/10 Game Diversity, Non-Agent Formats, and Cooperative Mechanics swyx critiques the benchmark for appearing strictly embodied and single-agent. Greg clarifies that many ARC-3 tasks are non-agent board games or require cooperative alignment, explaining the human-in-the-loop game construction pipeline.23:29–27:07 · The hosts as informed peer 5/10 Foundation Team Operations, Funding, and Hiring Alessio inquires about foundation staffing and hiring needs, while swyx asks about the future roadmap beyond ARC-3. Greg emphasizes keeping human solvability as their core operational anchor instead of designing esoteric PhD-level tests.27:08–32:04 · The hosts as informed peer 6/10 Benchmark Durability and Frontier Model Projections Alessio asks about benchmark durability and Grok 4's 16% score. swyx challenges the premise of declaring a single AGI moment, pointing out shifting goalposts. Greg firmly counters, dismissing revenue-based AGI definitions and insisting on a binary learning-efficiency milestone.32:04–36:04 · The hosts as informed peer 4/10 Inside the xAI Grok 4 Launch Event and Meeting Elon Musk swyx asks about the atmosphere at the Grok 4 launch event. Greg shares behind-the-scenes details of validating xAI's evaluation scores and pitching ARC-AGI-3's video game paradigm directly to Elon Musk.36:06–38:42 · The hosts as informed peer 6/10 Assessing Grok 4 Capabilities and Frontier Model Trends swyx and Greg evaluate Grok 4's frontier benchmark performance, discussing RL scaling, rumors surrounding dedicated coding models, and competing lab dynamics before Greg closes with a call for agent developers.1:04–3:18 · Guest teaching 5/10 Origins, Mission, and Growth of ARC Prize Foundation The hosts set up the background on ARC Prize gaining mainstream recognition over the past year. Greg explains the origins from François Chollet's 2019 paper to Mike Knoop funding the $1M bounty and turning it into a foundation.3:19–5:35 · Guest teaching 7/10 Defining Intelligence: Skill Acquisition Efficiency and Denominators swyx asks how Chollet defines AGI and intelligence acquisition. Greg breaks down the core thesis: intelligence is skill acquisition efficiency, defined by the denominators of energy and training data compared to the human brain.5:37–8:02 · Guest teaching 6/10 The Transition to Interactive Reasoning in ARC-AGI-3 Alessio steelmans the counterargument that compute costs are dropping and RL can simply be parallelized across domains. Greg explains that simulated RL environments typically have developer intelligence injected into them rather than general model intelligence.8:03–12:22 · Guest teaching 5/10 Interactive Demonstration of the Locksmith Game Greg shares his screen and plays through the Locksmith game demo. The hosts observe and comment as Greg illustrates how rules, energy constraints, and rotation mechanics force exploration and long-term planning.12:23–17:02 · Guest teaching 5/10 Technical Specifications and Agent Action Interface Alessio and swyx probe the API interface, text vs. multimodal representations, and scaffolding. swyx pushes back with prior findings that vision does not aid ARC and cites Noam Brown's view that agent scaffolds will eventually become obsolete.17:03–19:31 · Guest teaching 6/10 Evaluating Efficiency: Cost and Action Step Metrics Alessio asks how ARC accounts for model price fluctuations when measuring cost efficiency. Greg explains that while cost acts as a proxy for closed models, ARC-AGI-3 introduces action step efficiency as the primary metric.19:31–23:28 · Guest teaching 6/10 Game Diversity, Non-Agent Formats, and Cooperative Mechanics swyx critiques the benchmark for appearing strictly embodied and single-agent. Greg clarifies that many ARC-3 tasks are non-agent board games or require cooperative alignment, explaining the human-in-the-loop game construction pipeline.23:29–27:07 · Guest teaching 6/10 Foundation Team Operations, Funding, and Hiring Alessio inquires about foundation staffing and hiring needs, while swyx asks about the future roadmap beyond ARC-3. Greg emphasizes keeping human solvability as their core operational anchor instead of designing esoteric PhD-level tests.27:08–32:04 · Guest teaching 6/10 Benchmark Durability and Frontier Model Projections Alessio asks about benchmark durability and Grok 4's 16% score. swyx challenges the premise of declaring a single AGI moment, pointing out shifting goalposts. Greg firmly counters, dismissing revenue-based AGI definitions and insisting on a binary learning-efficiency milestone.32:04–36:04 · Guest teaching 6/10 Inside the xAI Grok 4 Launch Event and Meeting Elon Musk swyx asks about the atmosphere at the Grok 4 launch event. Greg shares behind-the-scenes details of validating xAI's evaluation scores and pitching ARC-AGI-3's video game paradigm directly to Elon Musk.36:06–38:42 · Guest teaching 5/10 Assessing Grok 4 Capabilities and Frontier Model Trends swyx and Greg evaluate Grok 4's frontier benchmark performance, discussing RL scaling, rumors surrounding dedicated coding models, and competing lab dynamics before Greg closes with a call for agent developers.1:04–3:18 · Guest disagreement 1/10 Origins, Mission, and Growth of ARC Prize Foundation The hosts set up the background on ARC Prize gaining mainstream recognition over the past year. Greg explains the origins from François Chollet's 2019 paper to Mike Knoop funding the $1M bounty and turning it into a foundation.3:19–5:35 · Guest disagreement 2/10 Defining Intelligence: Skill Acquisition Efficiency and Denominators swyx asks how Chollet defines AGI and intelligence acquisition. Greg breaks down the core thesis: intelligence is skill acquisition efficiency, defined by the denominators of energy and training data compared to the human brain.5:37–8:02 · Guest disagreement 2/10 The Transition to Interactive Reasoning in ARC-AGI-3 Alessio steelmans the counterargument that compute costs are dropping and RL can simply be parallelized across domains. Greg explains that simulated RL environments typically have developer intelligence injected into them rather than general model intelligence.8:03–12:22 · Guest disagreement 1/10 Interactive Demonstration of the Locksmith Game Greg shares his screen and plays through the Locksmith game demo. The hosts observe and comment as Greg illustrates how rules, energy constraints, and rotation mechanics force exploration and long-term planning.12:23–17:02 · Guest disagreement 2/10 Technical Specifications and Agent Action Interface Alessio and swyx probe the API interface, text vs. multimodal representations, and scaffolding. swyx pushes back with prior findings that vision does not aid ARC and cites Noam Brown's view that agent scaffolds will eventually become obsolete.17:03–19:31 · Guest disagreement 1/10 Evaluating Efficiency: Cost and Action Step Metrics Alessio asks how ARC accounts for model price fluctuations when measuring cost efficiency. Greg explains that while cost acts as a proxy for closed models, ARC-AGI-3 introduces action step efficiency as the primary metric.19:31–23:28 · Guest disagreement 2/10 Game Diversity, Non-Agent Formats, and Cooperative Mechanics swyx critiques the benchmark for appearing strictly embodied and single-agent. Greg clarifies that many ARC-3 tasks are non-agent board games or require cooperative alignment, explaining the human-in-the-loop game construction pipeline.23:29–27:07 · Guest disagreement 1/10 Foundation Team Operations, Funding, and Hiring Alessio inquires about foundation staffing and hiring needs, while swyx asks about the future roadmap beyond ARC-3. Greg emphasizes keeping human solvability as their core operational anchor instead of designing esoteric PhD-level tests.27:08–32:04 · Guest disagreement 5/10 Benchmark Durability and Frontier Model Projections Alessio asks about benchmark durability and Grok 4's 16% score. swyx challenges the premise of declaring a single AGI moment, pointing out shifting goalposts. Greg firmly counters, dismissing revenue-based AGI definitions and insisting on a binary learning-efficiency milestone.32:04–36:04 · Guest disagreement 1/10 Inside the xAI Grok 4 Launch Event and Meeting Elon Musk swyx asks about the atmosphere at the Grok 4 launch event. Greg shares behind-the-scenes details of validating xAI's evaluation scores and pitching ARC-AGI-3's video game paradigm directly to Elon Musk.36:06–38:42 · Guest disagreement 1/10 Assessing Grok 4 Capabilities and Frontier Model Trends swyx and Greg evaluate Grok 4's frontier benchmark performance, discussing RL scaling, rumors surrounding dedicated coding models, and competing lab dynamics before Greg closes with a call for agent developers.1:04–3:18 · The hosts pushing back 1/10 Origins, Mission, and Growth of ARC Prize Foundation The hosts set up the background on ARC Prize gaining mainstream recognition over the past year. Greg explains the origins from François Chollet's 2019 paper to Mike Knoop funding the $1M bounty and turning it into a foundation.3:19–5:35 · The hosts pushing back 1/10 Defining Intelligence: Skill Acquisition Efficiency and Denominators swyx asks how Chollet defines AGI and intelligence acquisition. Greg breaks down the core thesis: intelligence is skill acquisition efficiency, defined by the denominators of energy and training data compared to the human brain.5:37–8:02 · The hosts pushing back 4/10 The Transition to Interactive Reasoning in ARC-AGI-3 Alessio steelmans the counterargument that compute costs are dropping and RL can simply be parallelized across domains. Greg explains that simulated RL environments typically have developer intelligence injected into them rather than general model intelligence.8:03–12:22 · The hosts pushing back 1/10 Interactive Demonstration of the Locksmith Game Greg shares his screen and plays through the Locksmith game demo. The hosts observe and comment as Greg illustrates how rules, energy constraints, and rotation mechanics force exploration and long-term planning.12:23–17:02 · The hosts pushing back 4/10 Technical Specifications and Agent Action Interface Alessio and swyx probe the API interface, text vs. multimodal representations, and scaffolding. swyx pushes back with prior findings that vision does not aid ARC and cites Noam Brown's view that agent scaffolds will eventually become obsolete.17:03–19:31 · The hosts pushing back 2/10 Evaluating Efficiency: Cost and Action Step Metrics Alessio asks how ARC accounts for model price fluctuations when measuring cost efficiency. Greg explains that while cost acts as a proxy for closed models, ARC-AGI-3 introduces action step efficiency as the primary metric.19:31–23:28 · The hosts pushing back 4/10 Game Diversity, Non-Agent Formats, and Cooperative Mechanics swyx critiques the benchmark for appearing strictly embodied and single-agent. Greg clarifies that many ARC-3 tasks are non-agent board games or require cooperative alignment, explaining the human-in-the-loop game construction pipeline.23:29–27:07 · The hosts pushing back 2/10 Foundation Team Operations, Funding, and Hiring Alessio inquires about foundation staffing and hiring needs, while swyx asks about the future roadmap beyond ARC-3. Greg emphasizes keeping human solvability as their core operational anchor instead of designing esoteric PhD-level tests.27:08–32:04 · The hosts pushing back 5/10 Benchmark Durability and Frontier Model Projections Alessio asks about benchmark durability and Grok 4's 16% score. swyx challenges the premise of declaring a single AGI moment, pointing out shifting goalposts. Greg firmly counters, dismissing revenue-based AGI definitions and insisting on a binary learning-efficiency milestone.32:04–36:04 · The hosts pushing back 2/10 Inside the xAI Grok 4 Launch Event and Meeting Elon Musk swyx asks about the atmosphere at the Grok 4 launch event. Greg shares behind-the-scenes details of validating xAI's evaluation scores and pitching ARC-AGI-3's video game paradigm directly to Elon Musk.36:06–38:42 · The hosts pushing back 1/10 Assessing Grok 4 Capabilities and Frontier Model Trends swyx and Greg evaluate Grok 4's frontier benchmark performance, discussing RL scaling, rumors surrounding dedicated coding models, and competing lab dynamics before Greg closes with a call for agent developers.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 12.9% · guest 87.1%0:00 · the hosts 12.9% · guest 87.1%3:00 · the hosts 12.7% · guest 87.3%3:00 · the hosts 12.7% · guest 87.3%6:00 · the hosts 10.5% · guest 89.5%6:00 · the hosts 10.5% · guest 89.5%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 6.2% · guest 93.8%12:00 · the hosts 6.2% · guest 93.8%15:00 · the hosts 13% · guest 87%15:00 · the hosts 13% · guest 87%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 7.2% · guest 92.8%21:00 · the hosts 7.2% · guest 92.8%24:00 · the hosts 3.7% · guest 96.3%24:00 · the hosts 3.7% · guest 96.3%27:00 · the hosts 15.2% · guest 84.8%27:00 · the hosts 15.2% · guest 84.8%30:00 · the hosts 4.6% · guest 95.4%30:00 · the hosts 4.6% · guest 95.4%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 1.8% · guest 98.2%36:00 · the hosts 1.8% · guest 98.2%39:00 · the hosts 1.6% · guest 98.4%39:00 · the hosts 1.6% · guest 98.4%
Sharpest disagreement ▶ 31:20 Rejecting commercial definitions of AGI

Greg forcefully pushes back on swyx's suggestion that AGI is a continuous spectrum, stating that any definition involving profit has ulterior motives and insisting AGI is a binary threshold based on human efficiency.

Hardest push from the hosts ▶ 30:45 Challenging ARC's moving goalposts

swyx refuses the premise that declaring a single AGI moment is a worthwhile goal, arguing that ARC is constantly moving the goalposts across versions while labs like OpenAI treat AGI as continuous levels.

Biggest teaching moment ▶ 6:07 Developer bias in RL simulation environments

Greg educates the hosts on why scaling RL in custom environments does not equal generalization, explaining that developers inadvertently inject their own intelligence into the simulated reward structure.

The host holds their own ▶ 16:43 Citing Noam Brown on agent scaffolding

swyx demonstrates deep domain expertise by challenging Greg's belief in complex harnesses, quoting researcher Noam Brown to argue that agent scaffolding will inevitably be made obsolete by frontier models.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins, Mission, and Growth of ARC Prize Foundation 5511 The hosts set up the background on ARC Prize gaining mainstream recognition over the past year. Greg explains the origins from François Chollet's 2019 paper to Mike Knoop funding the $1M bounty and turning it into a foundation.
Defining Intelligence: Skill Acquisition Efficiency and Denominators 5721 swyx asks how Chollet defines AGI and intelligence acquisition. Greg breaks down the core thesis: intelligence is skill acquisition efficiency, defined by the denominators of energy and training data compared to the human brain.
The Transition to Interactive Reasoning in ARC-AGI-3 6624 Alessio steelmans the counterargument that compute costs are dropping and RL can simply be parallelized across domains. Greg explains that simulated RL environments typically have developer intelligence injected into them rather than general model intelligence.
Interactive Demonstration of the Locksmith Game 4511 Greg shares his screen and plays through the Locksmith game demo. The hosts observe and comment as Greg illustrates how rules, energy constraints, and rotation mechanics force exploration and long-term planning.
Technical Specifications and Agent Action Interface 7524 Alessio and swyx probe the API interface, text vs. multimodal representations, and scaffolding. swyx pushes back with prior findings that vision does not aid ARC and cites Noam Brown's view that agent scaffolds will eventually become obsolete.
Evaluating Efficiency: Cost and Action Step Metrics 6612 Alessio asks how ARC accounts for model price fluctuations when measuring cost efficiency. Greg explains that while cost acts as a proxy for closed models, ARC-AGI-3 introduces action step efficiency as the primary metric.
Game Diversity, Non-Agent Formats, and Cooperative Mechanics 6624 swyx critiques the benchmark for appearing strictly embodied and single-agent. Greg clarifies that many ARC-3 tasks are non-agent board games or require cooperative alignment, explaining the human-in-the-loop game construction pipeline.
Foundation Team Operations, Funding, and Hiring 5612 Alessio inquires about foundation staffing and hiring needs, while swyx asks about the future roadmap beyond ARC-3. Greg emphasizes keeping human solvability as their core operational anchor instead of designing esoteric PhD-level tests.
Benchmark Durability and Frontier Model Projections 6655 Alessio asks about benchmark durability and Grok 4's 16% score. swyx challenges the premise of declaring a single AGI moment, pointing out shifting goalposts. Greg firmly counters, dismissing revenue-based AGI definitions and insisting on a binary learning-efficiency milestone.
Inside the xAI Grok 4 Launch Event and Meeting Elon Musk 4612 swyx asks about the atmosphere at the Grok 4 launch event. Greg shares behind-the-scenes details of validating xAI's evaluation scores and pitching ARC-AGI-3's video game paradigm directly to Elon Musk.
Assessing Grok 4 Capabilities and Frontier Model Trends 6511 swyx and Greg evaluate Grok 4's frontier benchmark performance, discussing RL scaling, rumors surrounding dedicated coding models, and competing lab dynamics before Greg closes with a call for agent developers.

Statements from this episode (24)

Assertion Supported
Kamradt: Mike Knoop Put Up $1M for ARC Prize Bounty
“Mike actually put he put up a million dollars of his own money and said, Hey, I'm going to put a bounty. So for anybody who can beat this benchmark, they're going to get a million dollars.”
Greg Kamradt Jul 18, 2025 ▶ 1:50
Disclosure
Kamradt: ARC Prize Measures Intelligence via Energy and Data Efficiency
“So humans, they do not have an internet's worth of training data in their training data, but yet they can still do generally intelligent things. Whereas we're not seeing that with AI right now. So energy and training data are the two denominators that we use f…”
Greg Kamradt Jul 18, 2025 ▶ 5:23
Insight
Kamradt: Synthetic RL Transfers Developer Intelligence Rather Than Creating True Intelligence
“Often what happens is the human or developer intelligence is often injected into that environment itself, and so the model isn't actually Intelligent. You're just almost like taking the intelligence from the developer, injecting it into the environment, and th…”
Greg Kamradt Jul 18, 2025 ▶ 6:24
Disclosure
Kamradt Announces ARC-AGI-3 Will Feature 100 Novel Game Environments
“We're coming out with RKGI three. And what this is gonna be is it's gonna be a series of a hundred different novel environments, or you could simply call them a hundred different novel games that we're making ourselves.”
Greg Kamradt Jul 18, 2025 ▶ 7:05
Prediction Not checkable as stated
Kamradt Predicts AGI Will Be Declared Via An Interactive Benchmark
“My hypothesis is that when AGI is declared, it will happen via an interactive benchmark. We're not going to know that AGI is here just via a static benchmark.”
Greg Kamradt Jul 18, 2025 ▶ 7:38
Disclosure
Kamradt: ARC-AGI-3 Preview Launches Five Games
“So as a part of the preview, we're launching five games. Now, three of them are going to be public on day one, and two of them are going to be private.”
Greg Kamradt Jul 18, 2025 ▶ 8:34
Disclosure
Kamradt: ARC-AGI-3 Provides AI Agents With a 64x64 JSON Grid
“So we'll show the same thing to AI, except that AI is gonna get a JSON grid list of lists. So those get a bunch of numbers, 64 by 64, and they can choose to turn that into an image if they want to, or agnostic, do whatever you want with it if they want to do m…”
Greg Kamradt Jul 18, 2025 ▶ 9:11
Disclosure
Kamradt: ARC-AGI-3 Agents Interact via 64x64 Frames and Integer Actions
“What agents will get is agents will get a series of frames and those frames will be 64 by 64. Now generally it's just going to be one frame, but you might be able to get like maybe two in a row or three in a row, and that would show an animation. And so beginn…”
Greg Kamradt Jul 18, 2025 ▶ 12:42
Assertion Supported
Kamradt: No AI Has Beaten Any ARC-AGI-3 Game Level Yet
“It's still true. We have yet to have an AI successfully beat any level on any of these games. So it hasn't happened yet.”
Greg Kamradt Jul 18, 2025 ▶ 14:33
Prediction Not checkable as stated
Kamradt Predicts AGI Will Require Heavy Scaffolding and Cooperating Components
“Now I know that sounds kind of like a weird question, but my current hypothesis that AGI will be heavily scaffolded. And why do you have that? Well, my hypothesis is that you're going to need different components that are working together in order to get the e…”
Greg Kamradt Jul 18, 2025 ▶ 15:54
Assertion Supported
Kamradt: Random Brute Force Agent Fails ARC-AGI-3 Locksmith Game
“One of the quality checks that we do is we run a random agent at a million steps to see if it beats it or not. And no, it doesn't beat lockstep at all or locksmith.”
Greg Kamradt Jul 18, 2025 ▶ 18:28
Insight
Kamradt: Action Step Count is Core Metric for AI Learning Efficiency
“When we report learning efficiency for this, especially with AI versus humans, it's all going to be around how many actions do you take in order to complete the goal of the environment, which not only does that encompass learning what the environment entails, …”
Greg Kamradt Jul 18, 2025 ▶ 19:14
Disclosure
Kamradt: ARC-AGI-3 Features Large Percentage of Non-Agent Puzzle Games
“We have a requirement that games must be novel from each other. We have A large percentage of games that are non-agent based. So think of it as like solitaire or connect four or like Simon or memory or something like that. Those are non-agent based games.”
Greg Kamradt Jul 18, 2025 ▶ 20:07
Disclosure
Kamradt: ARC-AGI-3 Targets 120 Benchmark Games by Q1 2026
“Our goal is to come out with a 120 by Q one of next year.”
Greg Kamradt Jul 18, 2025 ▶ 22:45
Insight
Kamradt: Human-Built Benchmarks Prevent AI From Reverse-Engineering Generation Code
“And the problem with that is that we don't want to incentivize AI to derive the program that made the game. Right? And so if we continue to have humans make the game, then the AI is incentivized to try to reverse engineer the G inside of humans, and that's kin…”
Greg Kamradt Jul 18, 2025 ▶ 23:07
Disclosure
Kamradt: ARC-AGI-4 Will Expand Beyond 2D 64x64 Grids
“So without knowing what the exact answer is, I do know that RKGI four or five or whatever it may be, will need to allow us to have more axes of freedom that are above a two D 64 by 64 type of grid that comes from there.”
Greg Kamradt Jul 18, 2025 ▶ 25:56
Insight
Kamradt Defines AGI as When AI Can Solve All Human-Created Tasks
“Because our hypothesis and our definition of AGI is as long as we can come up with problems that humans can do and AI cannot, then we do not have AGI. And then the flip side of that is also true, which is When us as ArcPrize, we're like, we consider ourselves,…”
Greg Kamradt Jul 18, 2025 ▶ 26:28
Prediction Open · timeframe Dec 2026
Kamradt Predicts OpenAI Will Delay o5 Release Until 2026
“O five is not coming out this year is my guess, you know, it's going to be coming out next year.”
Greg Kamradt Jul 18, 2025 ▶ 28:26
Prediction Didn’t hold up
Kamradt Predicts ARC-AGI-2 Will Not Be Beaten For 12 Months
“My guess is it's not going to be beat for the next 12 months.”
Greg Kamradt Jul 18, 2025 ▶ 28:29
Prediction Open · timeframe Dec 2029
Kamradt Predicts ARC-AGI-3 Benchmark Will Remain Unbeaten For 3 Years
“And then V three, our durability estimate for that is three years. And that's what we're aiming for is 36 months for V three.”
Greg Kamradt Jul 18, 2025 ▶ 28:35
Opinion
Kamradt: Any AGI Definition Involving Profit Has Ulterior Motives
“Any AGI definition that involves money has ulterior motives. I mean, simple as that. Money has nothing to do with intelligence, right?”
Greg Kamradt Jul 18, 2025 ▶ 31:25
Assertion Supported
Kamradt: ARC Prize Validated Grok 4 Benchmark Scores Privately
“We ran it on semi-private. It worked out, looked great, validated.”
Greg Kamradt Jul 18, 2025 ▶ 32:52
Assertion Supported
Kamradt: xAI Increased RL Compute on Grok 4 Tenfold
“They tend X the RL that they put on top of Grok for that.”
Greg Kamradt Jul 18, 2025 ▶ 36:42
Assertion Not checkable as stated
Kamradt: xAI Delaying Coder Model Release To Beat Specific Rival
“I heard rumors that, that Grok doesn't want to release the coding model until it's better than one specific other lab out there. So they're going to wait and see when it's actually better for the, to, they can have that marketing point.”
Greg Kamradt Jul 18, 2025 ▶ 37:18
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.