Triton

4 statements across 4 episodes · 2 bullish · 2 bearish · 4 people on the record · first statement Oct 20, 2023 by Jeremy Howard · said 27 times in 8 episodes since 2023 · across every show →

Mentions by year

brought up most by Quentin Anthony (14), Dylan Patel (4), Jeremy Howard (3), Diego Bachman (2), Yining Zhang (1), George Hotz (1), Chris Lattner (1), Batuhan Taskaya (1)

tap a year for its mentions
00103205202320242025episodesmentions
035202320242025episodes it came up in
0022.545202320242025episodesmentions per episode
2025 19 mentions in 5 episodes 4 per episode
2023 8 mentions in 3 episodes 3 per episode

every mention, scene by scene, with the transcript →

Everything said about Triton, oldest first

Oct 20, 2023 bullish
Prediction Not checkable as stated
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
Jeremy Howard Oct 20, 2023 ▶ 1:16:28 The End of Finetuning — with Jeremy Howard of Fast.ai
Dec 5, 2023 bullish
Opinion
Patel: Hardware vendors and developers are coalescing around OpenAI's Triton
“Likewise, there's OpenAI's Triton, like, what they're trying to do there, and like you know, everyone's really coalescing around Triton. You know, people, you know, third-party hardware vendors”
Dylan Patel Dec 5, 2023 ▶ 13:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Sep 23, 2025 negative
Disclosure
Bachman: Manifest Switched Power Retention from Triton to Custom CUDA
“Actually, our initial version of power retention was written in Triton, but we realized quickly that it just didn't offer the flexibility to really squeeze the performance that we wanted out of the GPU. So we took a step back and dove into CUDA.”
Diego Bachman Sep 23, 2025 ▶ 9:33 ⚡️ Beyond Transformers with Power Retention
Nov 3, 2025 negative
Insight
AI coding assistants degrade quickly on low-level CUDA and PTX code
“Getting Below Triton or level or anything else down to like CUDA or below there's orders of magnitude less public good kernels at that level. And I think that really shows and models capabilities. So when I try and get a model to do something in CUDA or PTX or…”
Quentin Anthony Nov 3, 2025 ▶ 13:35 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.