Emmanuel Ameisen

14 statements across 1 episodes · 6 bullish · 3 bearish · 1 people on the record · first statement Jun 6, 2025 by Emmanuel Ameisen · said 3 times in 3 episodes since 2025 · across every show →

On the record as a speaker too: Emmanuel Ameisen's record, appearances and statements → this page counts the times other people say the name.

Mentions by year

brought up most by Shawn Wang (2)

tap a year for its mentions
00112220252026episodesmentions
01220252026episodes it came up in
000.511220252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Emmanuel Ameisen, oldest first

Jun 6, 2025 neutral
Assertion Supported
Ameisen: Model internal representations show measurable bias toward English logits
“And it does seem like Does sort of like inner representations have a higher connection to like the output logits for English logits. And so there's like some bias towards English at least in the model we studied here.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:11:13 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Insight
Ameisen: Circuit tracing diagnoses model failures by exposing incorrect internal representations
“Like you, you're not limited to studying what the model can do, right? Like if the model's failing at something like, you know, counting the number of letters in strawberry or whatever you could just try that and try to figure out the circuit for like, well, i…”
Emmanuel Ameisen Jun 6, 2025 ▶ 23:02 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Assertion Not checkable as stated
Ameisen: Every interpretability team member joined partly due to Anthropic's interactive papers
“When we had a team meeting, like it was a couple months ago, somebody on the team asked how many of the people on this team are here, at least in part because they like read one of these papers and thought like, wow, this is so compelling. Like this like makes…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:46:07 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 neutral
Insight
Ameisen: Superposition is more severe in language models than in vision models
“That means that like language models pack a lot more in less space than Vision models. So maybe like a kind of like really hand wavy analogy, right? It's like, well, if you want curve detectors, like you don't need that many curve detectors. You know, if each …”
Emmanuel Ameisen Jun 6, 2025 ▶ 35:01 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Insight
Ameisen: Validated circuit models enable predictable steering via feature swapping
“If you understood the circuit well, and if you identified where it's thinking about Huskies or where it's thinking about like kind of like breeding two different breeds, then you should be able to like swap these in and out and get it to kind of like say whate…”
Emmanuel Ameisen Jun 6, 2025 ▶ 20:45 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 neutral
Assertion Supported
Ameisen: Language model neurons are far less directly interpretable than vision neurons
“If you look at just the neurons of a lot of vision models, you can See neurons that are curve detectors or that are edge detectors or that are high, low frequency detectors. And so you can sort of like make sense of the neurons mostly. But if you look at neuro…”
Emmanuel Ameisen Jun 6, 2025 ▶ 34:31 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 bearish
Insight
Ameisen: Naive model pruning fails because superposition distributes critical representations
“Well, right, and it's like, on, on each example, maybe this neuron is like at the bottom of, like, what matters, but actually it's participating, like, five percent to, like, understanding English, like, doing integrals and, you know, like, whatever, like, cra…”
Emmanuel Ameisen Jun 6, 2025 ▶ 49:59 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Assertion Supported
Ameisen: Swapping Internal Features Proves Single-Pass LLM Multi-Step Reasoning
“We claim that this is like the Texas representation. Let's get another one and replace it. And we just change like that feature in the middle of the model and we change it to like California. And if you change it to California, sure enough, it says Sacramento.…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:00:02 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 negative
Opinion
Ameisen: Current LLM Chain of Thought Is Unfaithful and Untrustworthy
“So I think there's like a sense in which right now the chain of thought is, is unfaithful, or at least you can't read the chain of thought and trust that that's how the model did it.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:37:12 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 bullish
Assertion Not checkable as stated
Ameisen: Tracing prompt computation in models takes only minutes with built infrastructure
“One of the reasons that we're really excited about this method is once you've built your like infrastructure, like to go from a prompt to like what happened is, you know, O of minutes.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:47:52 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:33:39 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025
Disclosure
Ameisen: Golden Gate Claude was chosen organically after an internal demo
“Golden Gate Claude was like a pure, as far as I remember, at least, like, a pure, just like, Weird random thing where, like, somebody found it, initially went an internal demo of it, everybody thought it was hilarious, and then that's sort of how it came out. …”
Emmanuel Ameisen Jun 6, 2025 ▶ 44:15 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 positive
Assertion Supported
Ameisen: Larger language models share more concept representations across languages
“If you look inside the model, if you look at the middle of the model, which is the middle of this plot here, models share more features. They share more of these representations in the middle of the model, and bigger models share even more. And so the, like, t…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:07:57 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 negative
Opinion
Ameisen: Stochastic Parrots Label Ignores Complex Multi-Step LLM Reasoning
“It's, like, activating many different distributed representations, like, combining them, and sort of, like, doing something pretty complicated. And so, yeah, I think it's funny, because in my opinion, that's like, yeah, like, oh god, stochastic parrots is not …”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:02:24 The Utility of Interpretability — Emmanuel Amiesen
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.