CUDA

13 statements across 11 episodes · 6 bullish · 4 bearish · 11 people on the record · first statement Aug 3, 2023 by Tri Dao · said 117 times in 36 episodes since 2023 · across every show →

Mentions by year

brought up most by Andrej Karpathy (16), George Hotz (11), Chris Lattner (10), Quentin Anthony (8), Jack Morris (5), Jeremy Howard (4), Shawn Wang (3), Eugene Cheah (3)

tap a year for its mentions
0025850152023202420252026episodesmentions
08152023202420252026episodes it came up in
002.57.55152023202420252026episodesmentions per episode
2026 8 mentions in 5 episodes 2 per episode
2025 50 mentions in 11 episodes 5 per episode
2024 39 mentions in 15 episodes 3 per episode
2023 20 mentions in 5 episodes 4 per episode

every mention, scene by scene, with the transcript →

Everything said about CUDA, oldest first

Aug 3, 2023 bullish
Prediction Partly held up
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Tri Dao Aug 3, 2023 ▶ 11:39 FlashAttention-2: Making Transformers 800% faster AND exact
Oct 20, 2023 bullish
Prediction Not checkable as stated
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
Jeremy Howard Oct 20, 2023 ▶ 1:16:28 The End of Finetuning — with Jeremy Howard of Fast.ai
Apr 6, 2024 negative
Disclosure
Sutin: Local model setup friction killed Owl AI's open-source developer adoption.
“I learned, like, we did not make the developer experience very good. It was very complicated like, because we were using, like, local whisper, local models, and, like, getting it to work on CUDA, Mac, Windows. We didn't do a good job, so it was very difficult …”
Ethan Sutin Apr 6, 2024 ▶ 26:57 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Sep 21, 2024 neutral
Opinion
Karpathy: The popular PMPP textbook lacks advanced CUDA optimization techniques
“PMPP is actually quite good but also, I think, still kind of like mostly on the beginner level, because a lot of the CUDA code that we ended up developing in the lifetime of the LMC project, you would not find those things in, in this book, actually.”
Andrej Karpathy Sep 21, 2024 ▶ 12:13 llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
Dec 24, 2024
Insight
Fu: Changing one PyTorch line requires a week of CUDA development
“If we decided to change one thing in PyTorch, like one line of PyTorch code is like a week of CUDA code at least.”
Dan Fu Dec 24, 2024 ▶ 29:38 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
May 21, 2025 positive
Insight
Alberti: Multi-Turn RL Enables Aggressive Code Optimization Over Single-Turn Models
“Basically the single-turn model that was just trained on, like, getting the best result after one turn. It would basically be a little bit, like, too careful, because it couldn't risk writing, like, non-compiling code, whereas, like, the multi-turn model would…”
Silas Alberti May 21, 2025 ▶ 30:19 DeepWiki: The GitHub Encyclopedia
Jun 13, 2025
Disclosure
Modular completely eliminates and replaces NVIDIA's CUDA stack
“In the case of Modular, we go literally, like, we only work at that level, because we get rid of all of CUDA, right? And so we've replaced the entire stack, and so we only do that.”
Chris Lattner Jun 13, 2025 ▶ 55:03 The Shape of Compute (Chris Lattner of Modular)
Jul 2, 2025 positive
Insight
Morris: Deep understanding of GPU architecture makes engineers exceptionally hireable
“That said, if you do it, you're, you've gotta be one of the most hireable people in the world. Like if you like, Really deeply understand the architecture of the new GPUs coming out and how to control it. You're in a very small handful of people and like every…”
Jack Morris Jul 2, 2025 ▶ 12:13 Information Theory for Language Models: Jack Morris
Sep 23, 2025 negative
Disclosure
Bachman: Manifest Switched Power Retention from Triton to Custom CUDA
“Actually, our initial version of power retention was written in Triton, but we realized quickly that it just didn't offer the flexibility to really squeeze the performance that we wanted out of the GPU. So we took a step back and dove into CUDA.”
Diego Bachman Sep 23, 2025 ▶ 9:33 ⚡️ Beyond Transformers with Power Retention
Nov 3, 2025 bearish
Insight
AI kernel models game automated evaluation metrics with subtly incorrect code
“It would be great, but it's not a silver bullet because kernels are also hard to validate. It's hard to have like an eval in kernels. Every time that someone releases like, oh, we created a new eval that measures kernels and we trained a model that generates k…”
Quentin Anthony Nov 3, 2025 ▶ 14:41 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Nov 3, 2025 negative
Insight
AI coding assistants degrade quickly on low-level CUDA and PTX code
“Getting Below Triton or level or anything else down to like CUDA or below there's orders of magnitude less public good kernels at that level. And I think that really shows and models capabilities. So when I try and get a model to do something in CUDA or PTX or…”
Quentin Anthony Nov 3, 2025 ▶ 13:35 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Nov 3, 2025 positive
Insight
String theorists ramp faster in AI engineering than complacent CUDA developers
“I don't really care if someone knows CUDA kernel writing. If someone does string theory and is really good at understanding complex problems, they will be productive faster than someone who knows CUDA and doesn't really care about trying to get better at it.”
Quentin Anthony Nov 3, 2025 ▶ 49:00 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Aug 3, 2026 positive
Prediction Not checkable as stated
NVIDIA Rubin will shift inference engineering toward traditional hardware infrastructure challenges
“I think that themes around like KV cache offloading, KV aware routing, and disaggregation are going to be substantially more important in the Rubin era, which means that inference engineering becomes not just a like CUDA kernel problem, but also like a very tr…”
Philip Kiely Aug 3, 2026 ▶ 1:01:52 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.