“I would describe it more as not unlearning, but maybe suppression. I think there's, like, really, like, I guess, guarantees that you've fully removed information from a model is, is, I don't think it's been convincingly showed anywhere yet”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Jack Merullo
Insight
Merullo: LLM memorization spans a gradient from reasoning to rote recall
“You can actually see, like the way that we, like, disentangle memorization, you can kind of see this like, gradient of memorization in between both mechanistically and behaviorally with, like, logical reasoning tasks being quite distinct from rote memorization…”
Jack MerulloDec 31, 2025▶ 6:02[State of MechInterp] SAEs in Production, Circuit Tracing, AI4Science, "Pragmatic" Interp — Goodfire
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.