FlashAttention-2, every mention
6 scenes · ← back to FlashAttention-2
tap a year for its mentions
every year anyone Quentin Anthony 2Alessio Fanelli 1
Verbatim, from the transcripts: the passages where FlashAttention-2 comes up
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 3:08 Quentin Anthony So we released a blog at Zyphra on porting flash attention to over to my 300 X, just to sort of play out, can we move our own stack over 2 times in the scene
⚡️ Beyond Transformers with Power Retention
- ▶ 8:27 unnamed speaker Um, and then I also know you rewrote the flash attention to kernels, uh, and actually faster than the ones that had three that originally implemented.
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 20:12 Alessio Fanelli First of all, shout out to our friend Tridao, uh, who released Flash Attention Three, Flash Attention Two.
The Four Wars of the AI Stack - Dec 2023 Recap
- ▶ 34:18 unnamed speaker And I mean, together it's doing so much for like three dial and like fresh attention to and whatnot.