Triton, every mention
12 scenes · ← back to Triton
tap a year for its mentions
every year anyone Quentin Anthony 14Dylan Patel 4Jeremy Howard 3Diego Bachman 2Yining Zhang 1George Hotz 1Chris Lattner 1Batuhan Taskaya 1
Verbatim, from the transcripts: the passages where Triton comes up
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 6:13 Quentin Anthony So if you have good kernels for attention, MOPs and norms and so on, then it doesn't much matter what the front end to, you know, send tensors to and from those kernels is now on things like Triton, for example, that you, okay, you are… 4 times in the scene
- ▶ 10:52 Quentin Anthony Um, this is only if I need absolute control over the hardware and I've tried Triton. 3 times in the scene
- ▶ 13:12 Quentin Anthony So if I'm trying to write some like fusion of two kernels, or if I'm trying to write something in Triton that's very high level, models are great. 3 times in the scene
- ▶ 16:23 Quentin Anthony Um, if Triton doesn't do what I want, I need to go a level lower. 4 times in the scene
⚡️ Beyond Transformers with Power Retention
- ▶ 9:26 Diego Bachman And so we actually couldn't use something like Triton or Palace. 2 times in the scene
A Technical History of Generative Media
- ▶ 13:32 Batuhan Taskaya Trace the execution of, of, of, uh, your neural net and generate Triton kernels that are fused, like, that are more efficient.
The Shape of Compute (Chris Lattner of Modular)
- ▶ 20:06 Chris Lattner And the way that has always worked is you have, for example, CUDA or things like Triton Lang or things like this on the inside.
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 11:54 Yining Zhang So you, you should use something like Tweeten,
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 13:40 Dylan Patel Um, likewise, there's OpenAI's Triton, like, what they're trying to do there, and like, uh, you know, everyone's really coalescing around Triton. 3 times in the scene
- ▶ 59:01 Dylan Patel O and Triton one that I did.
The End of Finetuning — with Jeremy Howard of Fast.ai
- ▶ 1:15:57 Jeremy Howard But, I mean, the, the honest truth is, particularly before Triton, like, everybody knew that tiling is the right way to solve anything, and everybody knew that attention, fused attention, wasn't tiled, and that was stupid. 3 times in the scene
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 14:31 George Hotz Um, but I know they're generating some Triton stuff, which is going to generate the kernels on the fly.