Anthropic engineer David Hershey reflects on tuning conversation history length for agent harnesses on Latent Space.
Assertion Not checkable as stated
Hershey: Claude 3.7 extended thinking does not help Pokémon gameplay
“I've tested, like, all sorts of the extended thinking mode with, ah, 3.7 on it, and, like, it doesn't really help.”
Assertion Not checkable as stated
Hershey: Anthropic Study Found Claude Treats Named Characters Better
“Anthropic actually did like a blinded study of like named characters versus unnamed characters in different settings, and Claude like actually does clearly prefer and is nicer to named characters, which is an interesting thing.”
Disclosure
Upgrading Claude models in Pokémon primarily involves deleting prompt scaffolding
“Literally every model that has come out with Pokemon, like, the main change that I have made to this agent is deleting prompt stuff.”
Assertion Not checkable as stated
Prompting Claude to nickname its Pokémon caused it to exhibit protective behaviors
“And one thing we found when we started doing that is it got more protective of the Pokemon it nicknamed. Like, it's pretty obvious, like, when it catches a Pokemon, Now that it has a nickname, it will, like, go heal it right away if it's hurt, and that did not…”
Assertion Not checkable as stated
Claude 3.7 Sonnet uniquely generates meta-commentary reflecting on its perceptual errors
“And actually like one of the things that's most unique about. 3.7 sonnet that I've seen is like, it will have like meta commentary on what it's good at and bad at and it's knowledge base. Like I misperceived this thing. And so like, I need to be careful doing …”
Opinion
Hershey: Claude is exceptionally bad at point-to-point spatial navigation
“The thing that Claude's the worst at is understanding how to get from point A to point B. It's, like, really, really god-awful at hitting the buttons to go from point A to point B on a screen.”