The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 8 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Insight
AI agents have an optimal effective context length where intelligence peaks
“I think like one thing you see a lot when you talk to people building agents is there's like some effective context length that actually like has the model be the smartest. And that seems to vary slightly model by model, but for this model, for whatever purpos…”
David Hershey Mar 4, 2025 ▶ 18:57 How Claude Plays Pokémon was made
Insight
Hershey: Pokémon serves as an effective multi-day evaluation benchmark for AI models
“Building evals that actually test, like, time performance over, like, days of token sampling are quite hard. Like, it's actually really, really hard to build things that you can measure In a reasonable way. And one, this is one of them. Like, it's an eval that…”
David Hershey Apr 5, 2025 ▶ 4:58 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Insight
Hershey: Newer Claude Models Tenaciously Keep Trying Rather Than Quitting
“This is, like, maybe the thing that is the best about the new models is they have a tendency to, like, tenaciously still try things.”
David Hershey Apr 5, 2025 ▶ 6:56 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Insight
Hershey: Prompt engineering cannot fix Claude's visual comprehension limitations
“Vision is, like, pretty beyond fixing with a prompt. Again, go for it. Have fun. I've spent a lot of hours, like, overlaying grids, overlaying images, stretching, compressing, contrast colors, all sorts of stuff.”
David Hershey Apr 5, 2025 ▶ 12:54 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Insight
Hershey: Stacking Negative Scaffolding Instructions Degrades Agent Performance
“If you keep adding that type of instructions, you eventually make the model stupider. Like, if you add 50 of that type of instruction, you just end up, like, with the whole agent doing worse”
David Hershey Apr 5, 2025 ▶ 25:37 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Insight
Prompting Claude cannot improve its spatial navigation without explicit instructions
“You can try to prompt Quad a lot of different ways to understand how to navigate better, and anything short of telling it exactly what to do does not improve its, like, actual navigation.”
David Hershey Mar 4, 2025 ▶ 19:55 How Claude Plays Pokémon was made
Insight
Pokémon is ideal for testing AI agents because delays bring no penalty
“Pokemon's actually really nice because, like, if you don't do anything for five seconds, like, there's typically not a consequence by the nature of, like, doing inference on a model every, like, snapshot of time. It's actually a pretty good game to be able to …”
David Hershey Mar 4, 2025 ▶ 6:10 How Claude Plays Pokémon was made
Insight
Measuring AI game progress is an integration test, not a unit test
“And so I think like how quickly it's able to make progress is actually a pretty reason or a reasonable like eval if a slightly expensive one to calculate. It's an integration test, not a unit test.”
David Hershey Mar 4, 2025 ▶ 37:04 How Claude Plays Pokémon was made
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.