attention

3 statements across 3 episodes · 0 bullish · 1 bearish · 3 people on the record · first statement Aug 3, 2023 by Tri Dao · across every show →

Everything said about attention, oldest first

Aug 3, 2023 neutral
Assertion Supported
Memory read/write dominates standard attention computation time
“We ended up focusing a lot more on Memory reading and writing, because that turned out to be the majority of time when you're doing attention is reading and writing memory.”
Tri Dao Aug 3, 2023 ▶ 5:54 FlashAttention-2: Making Transformers 800% faster AND exact
Jun 6, 2025 negative
Assertion Not checkable as stated
Ameisen: Interpretability researchers lack good methods for analyzing attention layers
“So like, I think that right now we have some pretty good solutions for like understanding what's in the residual stream, understanding what's, is it in MLPs? We don't have good solutions for like attention.”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:28:25 The Utility of Interpretability — Emmanuel Amiesen
Nov 3, 2025
Insight
Training frontends matter little if attention and MLP kernels are highly optimized
“Most of that is an attention and MOPs, right? So if you have good kernels for attention, MOPs and norms and so on, then it doesn't much matter what the front end to, you know, send tensors to and from those kernels is”
Quentin Anthony Nov 3, 2025 ▶ 6:11 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.