Triton

product on 8 shows · 4 statements across 4 episodes · said 43 times in 17 episodes since 2023

Latent Space 27 the Y Combinator Startup Podcast 5 TBPN 2 In Depth 1 BG2 Pod 1 the MAD Podcast 1 the a16z Podcast 1 20VC 1

Mentions by year, every show

tap a year for its mentions
0013525102023202420252026episodesmentions
05102023202420252026episodes it came up in
002.555102023202420252026episodesmentions per episode

Latent Space 27the Y Combinator Startup Podcast 5TBPN 220VC 1BG2 Pod 1the a16z Podcast 1In Depth 1the MAD Podcast 1

2026 5 mentions in 1 episode
2025 24 mentions in 10 episodes 2 per episode
2024 1 mention in 1 episode
2023 9 mentions in 4 episodes 2 per episode

every mention on every show, scene by scene, with the transcript →

4 statements about Triton, every show

AI coding assistants degrade quickly on low-level CUDA and PTX code
“Getting Below Triton or level or anything else down to like CUDA or below there's orders of magnitude less public good kernels at that level. And I think that really shows and models capabilities. So when I try and get a model to do something in CUDA or PTX or…”
Quentin Anthony Nov 3, 2025 ▶ 13:35 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
LATENT SPACE Disclosure
Bachman: Manifest Switched Power Retention from Triton to Custom CUDA
“Actually, our initial version of power retention was written in Triton, but we realized quickly that it just didn't offer the flexibility to really squeeze the performance that we wanted out of the GPU. So we took a step back and dove into CUDA.”
Diego Bachman Sep 23, 2025 ▶ 9:33 ⚡️ Beyond Transformers with Power Retention
Patel: Hardware vendors and developers are coalescing around OpenAI's Triton
“Likewise, there's OpenAI's Triton, like, what they're trying to do there, and like you know, everyone's really coalescing around Triton. You know, people, you know, third-party hardware vendors”
Dylan Patel Dec 5, 2023 ▶ 13:40 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
LATENT SPACE Prediction Not checkable as stated
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
Jeremy Howard Oct 20, 2023 ▶ 1:16:28 The End of Finetuning — with Jeremy Howard of Fast.ai

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.