Ameisen: LLMs use internal circuits to backwards-plan rhyming poetry lines
Emmanuel Ameisen · The Utility of Interpretability — Emmanuel Amiesen · Jun 6, 2025 · at 1:18:59
Anthropic research scientist Emmanuel Ameisen explains how mechanistic interpretability circuit tracing revealed backward planning behavior in Claude during poetry generation.
“And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called like backwards planning, where it's like, well, because I need to finish with green, I'm not going to say illuminating the peaceful night, because then I'd be like illuminating the peaceful green. That doesn't make sense. I need to say a completely different sentence that lets me finish with green. And so there's a circuit in the model that decides on the rhyme and then works backwards from the rhyme. To set up your sentence.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →