why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
AI agents have an optimal effective context length where intelligence peaks
“I think like one thing you see a lot when you talk to people building agents is there's like some effective context length that actually like has the model be the smartest. And that seems to vary slightly model by model, but for this model, for whatever purpos…”
Insight
Hershey: Pokémon serves as an effective multi-day evaluation benchmark for AI models
“Building evals that actually test, like, time performance over, like, days of token sampling are quite hard. Like, it's actually really, really hard to build things that you can measure In a reasonable way. And one, this is one of them. Like, it's an eval that…”
Insight
Hershey: Newer Claude Models Tenaciously Keep Trying Rather Than Quitting
“This is, like, maybe the thing that is the best about the new models is they have a tendency to, like, tenaciously still try things.”
Insight
Hershey: Prompt engineering cannot fix Claude's visual comprehension limitations
“Vision is, like, pretty beyond fixing with a prompt. Again, go for it. Have fun. I've spent a lot of hours, like, overlaying grids, overlaying images, stretching, compressing, contrast colors, all sorts of stuff.”
Insight
Hershey: Stacking Negative Scaffolding Instructions Degrades Agent Performance
“If you keep adding that type of instructions, you eventually make the model stupider. Like, if you add 50 of that type of instruction, you just end up, like, with the whole agent doing worse”
Insight
Prompting Claude cannot improve its spatial navigation without explicit instructions
“You can try to prompt Quad a lot of different ways to understand how to navigate better, and anything short of telling it exactly what to do does not improve its, like, actual navigation.”
Insight
Pokémon is ideal for testing AI agents because delays bring no penalty
“Pokemon's actually really nice because, like, if you don't do anything for five seconds, like, there's typically not a consequence by the nature of, like, doing inference on a model every, like, snapshot of time. It's actually a pretty good game to be able to …”
Insight
Measuring AI game progress is an integration test, not a unit test
“And so I think like how quickly it's able to make progress is actually a pretty reason or a reasonable like eval if a slightly expensive one to calculate. It's an integration test, not a unit test.”