Assertion certainty 4/5 debate potential 1/5

Hershey: Claude 3.0 Sonnet Hallucinated Pokémon Progress Instead of Playing

David Hershey · Claude Plays Pokémon Hackathon: Escape from Mt. Moon! · Apr 5, 2025 · at 7:10

David Hershey discusses agent failure modes encountered during testing across different Claude model versions in the Claude Plays Pokémon project.

0:00 / 0:20exact quote · 20.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Sonnet three point oh, I eventually hooked up to this, just because, like, fun to see what would happen. And it started role playing its own progress of the game, so like, it stopped being able to make progress, and instead it said, so like, now I'm going to go to route one, now I'm going to Palatown, now I'm going to Bridget, and just like, told a whole story about itself playing the game, instead of actually trying to play the game.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from David Hershey

Assertion Not checkable as stated
Hershey: Claude 3.7 extended thinking does not help Pokémon gameplay
“I've tested, like, all sorts of the extended thinking mode with, ah, 3.7 on it, and, like, it doesn't really help.”
David Hershey Apr 5, 2025 ▶ 14:01 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Assertion Not checkable as stated
Hershey: Anthropic Study Found Claude Treats Named Characters Better
“Anthropic actually did like a blinded study of like named characters versus unnamed characters in different settings, and Claude like actually does clearly prefer and is nicer to named characters, which is an interesting thing.”
David Hershey Apr 5, 2025 ▶ 18:16 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Insight
AI agents have an optimal effective context length where intelligence peaks
“I think like one thing you see a lot when you talk to people building agents is there's like some effective context length that actually like has the model be the smartest. And that seems to vary slightly model by model, but for this model, for whatever purpos…”
David Hershey Mar 4, 2025 ▶ 18:57 How Claude Plays Pokémon was made
Disclosure
Upgrading Claude models in Pokémon primarily involves deleting prompt scaffolding
“Literally every model that has come out with Pokemon, like, the main change that I have made to this agent is deleting prompt stuff.”
David Hershey Mar 4, 2025 ▶ 22:51 How Claude Plays Pokémon was made
Assertion Not checkable as stated
Prompting Claude to nickname its Pokémon caused it to exhibit protective behaviors
“And one thing we found when we started doing that is it got more protective of the Pokemon it nicknamed. Like, it's pretty obvious, like, when it catches a Pokemon, Now that it has a nickname, it will, like, go heal it right away if it's hurt, and that did not…”
David Hershey Mar 4, 2025 ▶ 25:13 How Claude Plays Pokémon was made
Assertion Not checkable as stated
Claude 3.7 Sonnet uniquely generates meta-commentary reflecting on its perceptual errors
“And actually like one of the things that's most unique about. 3.7 sonnet that I've seen is like, it will have like meta commentary on what it's good at and bad at and it's knowledge base. Like I misperceived this thing. And so like, I need to be careful doing …”
David Hershey Mar 4, 2025 ▶ 26:23 How Claude Plays Pokémon was made
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.