Mar 4, 2025 · 37m · latent-space

How Claude Plays Pokémon was made

David Hershey · 24m spoken Alessio Fanelli · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Anthropic engineer David Hershey joins Latent Space to detail how he built 'Claude Plays Pokémon,' examining agent architecture, spatial reasoning hurdles, context window management, and the evolution of autonomous capabilities in Claude 3.7 Sonnet.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 14.8% of the talking time here. How this is scored →

The hosts as informed peer 4.0 Guest teaching 4.9 Guest disagreement 1.3 The hosts pushing back 1.1
05100:0010:0020:0030:000:56–6:27 · The hosts as informed peer 4/10 Origins and Evolution of Claude Plays Pokémon Alessio sets up the premise comparing Claude Plays Pokémon to Twitch Plays Pokémon and asks about game mechanics like isometric views. David details the project's progression across Claude model versions from 3.5 Sonnet to 3.7 Sonnet.6:28–12:01 · The hosts as informed peer 4/10 System Architecture and Emulator Integration David walks through the harness architecture diagram, explaining tool definitions, RAM reverse-engineering, and hallucination workarounds. Vibu and Alessio ask technical clarifying questions about Game Boy coordinates and emulator state extraction.12:04–15:27 · The hosts as informed peer 3/10 Spatial Vision Challenges and the Navigator Tool Alessio asks how much domain knowledge the model has and how Claude handles spatial self-awareness. David educates on how Claude's visual reasoning struggles with 2D Game Boy sprites, requiring the custom Navigator tool patch.15:28–19:20 · The hosts as informed peer 4/10 Token Economics and Context Window Management Vibu probes into context window mechanics, token economics, and truncation boundaries. David details the exact breakdown of token budgets, 30-step rollout cycles, and summarization trigger points.19:20–24:12 · The hosts as informed peer 4/10 Model Reasoning and Prompt Simplification Alessio checks Claude's pathfinding plan for Mount Moon, and Vibu questions whether pure guide memorization is ideal. David explains the counterintuitive discovery that removing prompt band-aids yielded better reasoning performance on Claude 3.7.24:13–29:05 · The hosts as informed peer 5/10 Emergent Behaviors and Cross-Game Skill Transfer Alessio draws an analogy to cross-game conceptual learning in Magic: The Gathering and trading card games. David agrees and explains how the model writes its own meta-commentary inside the prompt dictionary to learn from tactical mistakes.29:06–34:24 · The hosts as informed peer 4/10 Overcoming Game Bottlenecks and Milestone Highlights Vibu inquires about community optimization opportunities and formal evaluation metrics. David pushes back on the idea that prompt tweaks can fix fundamental vision gaps, recounting Claude repeatedly entering and exiting Oak's lab in an infinite loop.0:56–6:27 · Guest teaching 3/10 Origins and Evolution of Claude Plays Pokémon Alessio sets up the premise comparing Claude Plays Pokémon to Twitch Plays Pokémon and asks about game mechanics like isometric views. David details the project's progression across Claude model versions from 3.5 Sonnet to 3.7 Sonnet.6:28–12:01 · Guest teaching 5/10 System Architecture and Emulator Integration David walks through the harness architecture diagram, explaining tool definitions, RAM reverse-engineering, and hallucination workarounds. Vibu and Alessio ask technical clarifying questions about Game Boy coordinates and emulator state extraction.12:04–15:27 · Guest teaching 6/10 Spatial Vision Challenges and the Navigator Tool Alessio asks how much domain knowledge the model has and how Claude handles spatial self-awareness. David educates on how Claude's visual reasoning struggles with 2D Game Boy sprites, requiring the custom Navigator tool patch.15:28–19:20 · Guest teaching 5/10 Token Economics and Context Window Management Vibu probes into context window mechanics, token economics, and truncation boundaries. David details the exact breakdown of token budgets, 30-step rollout cycles, and summarization trigger points.19:20–24:12 · Guest teaching 6/10 Model Reasoning and Prompt Simplification Alessio checks Claude's pathfinding plan for Mount Moon, and Vibu questions whether pure guide memorization is ideal. David explains the counterintuitive discovery that removing prompt band-aids yielded better reasoning performance on Claude 3.7.24:13–29:05 · Guest teaching 4/10 Emergent Behaviors and Cross-Game Skill Transfer Alessio draws an analogy to cross-game conceptual learning in Magic: The Gathering and trading card games. David agrees and explains how the model writes its own meta-commentary inside the prompt dictionary to learn from tactical mistakes.29:06–34:24 · Guest teaching 5/10 Overcoming Game Bottlenecks and Milestone Highlights Vibu inquires about community optimization opportunities and formal evaluation metrics. David pushes back on the idea that prompt tweaks can fix fundamental vision gaps, recounting Claude repeatedly entering and exiting Oak's lab in an infinite loop.0:56–6:27 · Guest disagreement 1/10 Origins and Evolution of Claude Plays Pokémon Alessio sets up the premise comparing Claude Plays Pokémon to Twitch Plays Pokémon and asks about game mechanics like isometric views. David details the project's progression across Claude model versions from 3.5 Sonnet to 3.7 Sonnet.6:28–12:01 · Guest disagreement 1/10 System Architecture and Emulator Integration David walks through the harness architecture diagram, explaining tool definitions, RAM reverse-engineering, and hallucination workarounds. Vibu and Alessio ask technical clarifying questions about Game Boy coordinates and emulator state extraction.12:04–15:27 · Guest disagreement 1/10 Spatial Vision Challenges and the Navigator Tool Alessio asks how much domain knowledge the model has and how Claude handles spatial self-awareness. David educates on how Claude's visual reasoning struggles with 2D Game Boy sprites, requiring the custom Navigator tool patch.15:28–19:20 · Guest disagreement 1/10 Token Economics and Context Window Management Vibu probes into context window mechanics, token economics, and truncation boundaries. David details the exact breakdown of token budgets, 30-step rollout cycles, and summarization trigger points.19:20–24:12 · Guest disagreement 2/10 Model Reasoning and Prompt Simplification Alessio checks Claude's pathfinding plan for Mount Moon, and Vibu questions whether pure guide memorization is ideal. David explains the counterintuitive discovery that removing prompt band-aids yielded better reasoning performance on Claude 3.7.24:13–29:05 · Guest disagreement 1/10 Emergent Behaviors and Cross-Game Skill Transfer Alessio draws an analogy to cross-game conceptual learning in Magic: The Gathering and trading card games. David agrees and explains how the model writes its own meta-commentary inside the prompt dictionary to learn from tactical mistakes.29:06–34:24 · Guest disagreement 2/10 Overcoming Game Bottlenecks and Milestone Highlights Vibu inquires about community optimization opportunities and formal evaluation metrics. David pushes back on the idea that prompt tweaks can fix fundamental vision gaps, recounting Claude repeatedly entering and exiting Oak's lab in an infinite loop.0:56–6:27 · The hosts pushing back 1/10 Origins and Evolution of Claude Plays Pokémon Alessio sets up the premise comparing Claude Plays Pokémon to Twitch Plays Pokémon and asks about game mechanics like isometric views. David details the project's progression across Claude model versions from 3.5 Sonnet to 3.7 Sonnet.6:28–12:01 · The hosts pushing back 1/10 System Architecture and Emulator Integration David walks through the harness architecture diagram, explaining tool definitions, RAM reverse-engineering, and hallucination workarounds. Vibu and Alessio ask technical clarifying questions about Game Boy coordinates and emulator state extraction.12:04–15:27 · The hosts pushing back 1/10 Spatial Vision Challenges and the Navigator Tool Alessio asks how much domain knowledge the model has and how Claude handles spatial self-awareness. David educates on how Claude's visual reasoning struggles with 2D Game Boy sprites, requiring the custom Navigator tool patch.15:28–19:20 · The hosts pushing back 1/10 Token Economics and Context Window Management Vibu probes into context window mechanics, token economics, and truncation boundaries. David details the exact breakdown of token budgets, 30-step rollout cycles, and summarization trigger points.19:20–24:12 · The hosts pushing back 2/10 Model Reasoning and Prompt Simplification Alessio checks Claude's pathfinding plan for Mount Moon, and Vibu questions whether pure guide memorization is ideal. David explains the counterintuitive discovery that removing prompt band-aids yielded better reasoning performance on Claude 3.7.24:13–29:05 · The hosts pushing back 1/10 Emergent Behaviors and Cross-Game Skill Transfer Alessio draws an analogy to cross-game conceptual learning in Magic: The Gathering and trading card games. David agrees and explains how the model writes its own meta-commentary inside the prompt dictionary to learn from tactical mistakes.29:06–34:24 · The hosts pushing back 1/10 Overcoming Game Bottlenecks and Milestone Highlights Vibu inquires about community optimization opportunities and formal evaluation metrics. David pushes back on the idea that prompt tweaks can fix fundamental vision gaps, recounting Claude repeatedly entering and exiting Oak's lab in an infinite loop.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 50.6% · guest 49.4%0:00 · the hosts 50.6% · guest 49.4%3:00 · the hosts 15.3% · guest 84.7%3:00 · the hosts 15.3% · guest 84.7%6:00 · the hosts 8% · guest 92%6:00 · the hosts 8% · guest 92%9:00 · the hosts 1.8% · guest 98.2%9:00 · the hosts 1.8% · guest 98.2%12:00 · the hosts 13.6% · guest 86.4%12:00 · the hosts 13.6% · guest 86.4%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 15.3% · guest 84.7%18:00 · the hosts 15.3% · guest 84.7%21:00 · the hosts 0.1% · guest 99.9%21:00 · the hosts 0.1% · guest 99.9%24:00 · the hosts 32.1% · guest 67.9%24:00 · the hosts 32.1% · guest 67.9%27:00 · the hosts 11.8% · guest 88.2%27:00 · the hosts 11.8% · guest 88.2%30:00 · the hosts 18.7% · guest 81.3%30:00 · the hosts 18.7% · guest 81.3%33:00 · the hosts 14.9% · guest 85.1%33:00 · the hosts 14.9% · guest 85.1%36:00 · the hosts 5.9% · guest 94.1%36:00 · the hosts 5.9% · guest 94.1%
Sharpest disagreement ▶ 30:00 David dismisses prompt engineering as a cure for navigation

David politely but firmly rejects the Twitch community's assertions that clever prompt crafting can fix Claude's spatial disorientation.

Hardest push from the hosts ▶ 20:42 Alessio queries Claude live on the Mount Moon walkthrough

Alessio directly tests Claude's pathing logic live on Claude.ai to challenge whether navigation failure stems from missing game knowledge.

Biggest teaching moment ▶ 22:30 David explains the counter-intuitive power of prompt deletion

David educates the hosts on how human intuition and prescriptive prompting actually degraded performance relative to giving raw reasoning models free reign.

The host holds their own ▶ 27:29 Alessio connects game learning to TCG tempo transfer

Alessio demonstrates domain depth by comparing agent cross-game knowledge transfer to universal tempo concepts across Magic and Flesh and Blood.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins and Evolution of Claude Plays Pokémon 4311 Alessio sets up the premise comparing Claude Plays Pokémon to Twitch Plays Pokémon and asks about game mechanics like isometric views. David details the project's progression across Claude model versions from 3.5 Sonnet to 3.7 Sonnet.
System Architecture and Emulator Integration 4511 David walks through the harness architecture diagram, explaining tool definitions, RAM reverse-engineering, and hallucination workarounds. Vibu and Alessio ask technical clarifying questions about Game Boy coordinates and emulator state extraction.
Spatial Vision Challenges and the Navigator Tool 3611 Alessio asks how much domain knowledge the model has and how Claude handles spatial self-awareness. David educates on how Claude's visual reasoning struggles with 2D Game Boy sprites, requiring the custom Navigator tool patch.
Token Economics and Context Window Management 4511 Vibu probes into context window mechanics, token economics, and truncation boundaries. David details the exact breakdown of token budgets, 30-step rollout cycles, and summarization trigger points.
Model Reasoning and Prompt Simplification 4622 Alessio checks Claude's pathfinding plan for Mount Moon, and Vibu questions whether pure guide memorization is ideal. David explains the counterintuitive discovery that removing prompt band-aids yielded better reasoning performance on Claude 3.7.
Emergent Behaviors and Cross-Game Skill Transfer 5411 Alessio draws an analogy to cross-game conceptual learning in Magic: The Gathering and trading card games. David agrees and explains how the model writes its own meta-commentary inside the prompt dictionary to learn from tactical mistakes.
Overcoming Game Bottlenecks and Milestone Highlights 4521 Vibu inquires about community optimization opportunities and formal evaluation metrics. David pushes back on the idea that prompt tweaks can fix fundamental vision gaps, recounting Claude repeatedly entering and exiting Oak's lab in an infinite loop.

Statements from this episode (16)

Insight
Pokémon is ideal for testing AI agents because delays bring no penalty
“Pokemon's actually really nice because, like, if you don't do anything for five seconds, like, there's typically not a consequence by the nature of, like, doing inference on a model every, like, snapshot of time. It's actually a pretty good game to be able to …”
David Hershey Mar 4, 2025 ▶ 6:10
Assertion Not checkable as stated
Claude aggressively hallucinates game zone transitions without explicit negative feedback
“Claude will, like, pretty aggressively hallucinate that it succeeded in transitioning between zones if you don't, like, tell it did not.”
David Hershey Mar 4, 2025 ▶ 11:09
Assertion Not checkable as stated
Claude spent 12 hours overnight mistaking a Pokémon doormat for a textbox
“I once saw it see like a red box on the screen that was like the doormat and think it was a text box and spend 12 hours pressing A overnight to try to clear the text box, which you see that happen once and you add in some helpful reminders to not do that.”
David Hershey Mar 4, 2025 ▶ 11:44
Assertion Not checkable as stated
Claude still struggles with spatial awareness and visual positioning on screen
“Quad doesn't particularly understand, like, the middle of a Game Boy screen and a whole bunch of concepts like that, which means, like, you can prompt all around everywhere, but, like, this kind of, like, spatial awareness and where something is with respect t…”
David Hershey Mar 4, 2025 ▶ 14:05
Disclosure
Claude Plays Pokémon API requests max out around 100,000 tokens
“So in practice, this rollout ends up, like, at max, ending up around a 100,000 tokens, I think, is where it is, like, the longest message you ever send to the API on one of these turns, and it will fluctuate in, like, summarization, depending on the state of k…”
David Hershey Mar 4, 2025 ▶ 17:27
Assertion Not checkable as stated
Running Claude Plays Pokémon experiments costs thousands of dollars in API tokens
“There's like at least thousands of dollars of tokens being consumed. So it's not a, it is not a cheap rollout.”
David Hershey Mar 4, 2025 ▶ 18:35
Insight
AI agents have an optimal effective context length where intelligence peaks
“I think like one thing you see a lot when you talk to people building agents is there's like some effective context length that actually like has the model be the smartest. And that seems to vary slightly model by model, but for this model, for whatever purpos…”
David Hershey Mar 4, 2025 ▶ 18:57
Insight
Prompting Claude cannot improve its spatial navigation without explicit instructions
“You can try to prompt Quad a lot of different ways to understand how to navigate better, and anything short of telling it exactly what to do does not improve its, like, actual navigation.”
David Hershey Mar 4, 2025 ▶ 19:55
Disclosure
Upgrading Claude models in Pokémon primarily involves deleting prompt scaffolding
“Literally every model that has come out with Pokemon, like, the main change that I have made to this agent is deleting prompt stuff.”
David Hershey Mar 4, 2025 ▶ 22:51
Assertion Not checkable as stated
Prompting Claude to nickname its Pokémon caused it to exhibit protective behaviors
“And one thing we found when we started doing that is it got more protective of the Pokemon it nicknamed. Like, it's pretty obvious, like, when it catches a Pokemon, Now that it has a nickname, it will, like, go heal it right away if it's hurt, and that did not…”
David Hershey Mar 4, 2025 ▶ 25:13
Assertion Not checkable as stated
Claude 3.7 Sonnet uniquely generates meta-commentary reflecting on its perceptual errors
“And actually like one of the things that's most unique about. 3.7 sonnet that I've seen is like, it will have like meta commentary on what it's good at and bad at and it's knowledge base. Like I misperceived this thing. And so like, I need to be careful doing …”
David Hershey Mar 4, 2025 ▶ 26:23
Disclosure
The Claude Plays Pokémon knowledge base is simply a prompted Python dictionary
“I think my knowledge base is frankly, like, kind of kludgy of an implementation right now. Like, it's, like, more or less a Python dictionary that's appended to the prompt.”
David Hershey Mar 4, 2025 ▶ 26:50
Prediction Held up
Hershey predicts the Claude stream won't reach Victory Road within 16 days
“I think we have a little ways before we can beat the game in 16 days. I do not have a lot of faith that the current stream is gonna, gonna be standing in Victory Road in 13 days.”
David Hershey Mar 4, 2025 ▶ 32:42
Assertion Supported
Anthropic's Pokémon research graph reflects a single run passing Lt. Surge
“The run that you saw that's, like, on the graph we put out alongside, like, in our research blog is, like, a single run that I have watched, like, get through At least surges Jim. And then it got a little past that. And the reason that that's where we stopped …”
David Hershey Mar 4, 2025 ▶ 33:55
Prediction Not checkable as stated
Claude 3.7's error-correction capabilities will enable better real-world AI agents
“I really do think like this is just demonstrating like a thing that is going to make agents better with this model, you know, like This is a very fun way to see it, but, like, I think the thing is that it, like, has some ability to, like, course correct, updat…”
David Hershey Mar 4, 2025 ▶ 35:39
Insight
Measuring AI game progress is an integration test, not a unit test
“And so I think like how quickly it's able to make progress is actually a pretty reason or a reasonable like eval if a slightly expensive one to calculate. It's an integration test, not a unit test.”
David Hershey Mar 4, 2025 ▶ 37:04
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.