FlashAttention, every mention
1 scene · ← back to FlashAttention
tap a year for its mentions
every year anyone Stuart (Stu) 1
Verbatim, from the transcripts: the passages where FlashAttention comes up
Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
- ▶ 8:03 Stuart (Stu) Um, for instance, we have IO aware algorithms like flash attention, linear attention models, mega kernels, and many, many programming libraries and frameworks for, um, writing efficient kernels across many architects.