CUDA, every mention
64 scenes · ← back to CUDA
tap a year for its mentions
every year anyone Andrej Karpathy 16George Hotz 11Chris Lattner 10Quentin Anthony 8Jack Morris 5Jeremy Howard 4Shawn Wang 3Eugene Cheah 3Dan Fu 3Ben Firshman 3
Verbatim, from the transcripts: the passages where CUDA comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 1:01:52 Philip Kiely Um, so I think that themes around like KV cache offloading, KV aware routing, and, and disaggregation are going to be substantially more important in the Rubin era, which means that inference engineering becomes not just a like CUDA kernel…
- ▶ 1:05:20 Ali Taha I can write CUDA to control it and change its operations.
Satya Nadella on AI: @NoPriorsPodcast x Latent Space Crossover Special at Microsoft Build 2026
- ▶ 14:31 Satya Nadella We built DX and he built, you know, CUDA on top of it.
The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition
- ▶ 21:23 Peter Ludwig and you have a start with an assumption that, oh, I'm gonna, I'm gonna use CUDA and I'm gonna run this, uh, on an NVIDIA chip, then you don't really have to think about the hardware in that sense. 2 times in the scene
AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
- ▶ 1:02:50 Mikhail Parakhin And even Liquid, we had to work a lot with NVIDIA and to, because almost everything is not designed in CUDA for, or in, in the current stack for, for low latency.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 11:03 Kyle Kranen They don't know what CUDA is. 2 times in the scene
World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
- ▶ 15:54 Pim de Witte so I was very, very familiar with CUDA and like the GPU side, and all the video infrastructure that we were using for this stuff, but the modeling side itself was, was still quite foreign.
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 8:42 unnamed speaker Because I think the other question is, like, well, I'm gonna do all this work versus, like, I just write CUDA code that then, when the BG-G-Hundred comes online for my cluster, I'll just switch it over right away.
- ▶ 10:23 unnamed speaker And I think obviously if you understand CUDA is like, you know, 4 times in the scene
- ▶ 13:36 Quentin Anthony Below Triton or level or anything else down to like CUDA or below, um, there's orders of magnitude less public good kernels at that level. 4 times in the scene
- ▶ 22:42 Quentin Anthony So like it's very, Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like 2 times in the scene
- ▶ 49:00 Quentin Anthony I don't really care if someone knows CUDA kernel writing. 2 times in the scene
⚡️ Beyond Transformers with Power Retention
- ▶ 8:14 unnamed speaker And so you're also releasing, um, just to go from the bottom up, uh, Vidrill, which is a framework for CUDA kernel. 4 times in the scene
- ▶ 29:53 unnamed speaker If you are excited by those ideas, especially of a strong background in deep learning, uh, CUDA programming, or anything mathematically adjacent to those ideas, please reach out.
A Technical History of Generative Media
- ▶ 10:45 unnamed speaker Kuda kernel specialist at the time, right? 2 times in the scene
⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
- ▶ 22:05 Thomas Sohmers You know, the way that we say that we're in the CUDA ecosystem, it's by nature of the fact that we are ingesting the, the output of everything that exists in, in the CUDA infrastructure. 2 times in the scene
Greg Brockman on OpenAI's Road to AGI
- ▶ 52:14 Greg Brockman Cuda kernels are a good example of a very self-contained problem that actually our models should get very good at very soon, but it's just difficult because it requires a lot of domain expertise, a lot of like real abstract thinking. 2 times in the scene
Information Theory for Language Models: Jack Morris
- ▶ 11:20 Jack Morris Yeah, I'll comment on that quickly because if someone has been listening to this and also following me online for a while, I think I've made a couple of comments like saying something like you shouldn't learn about CUDA or things to that… 5 times in the scene
The Shape of Compute (Chris Lattner of Modular)
- ▶ 3:10 Chris Lattner And I said, oh, by the way, we can't use CUDA.
- ▶ 8:24 Chris Lattner Also, let's not just build, uh, you know, air quotes, a CUDA replacement. 3 times in the scene
- ▶ 11:47 Chris Lattner It doesn't use CUDA. 3 times in the scene
- ▶ 30:10 Chris Lattner Massively hand coded CUDA kernels and all this stuff for any new thing.
- ▶ 41:30 Chris Lattner But I will admit that their time to market and revenue growth and stuff like that has been much faster because they didn't have to like build an entire replacement for CUDA to get there. 2 times in the scene
- ▶ 52:54 Alessio Fanelli And I think specifically in your case, you know, they worked at the BTX layer of the GPU, which is like even lower and more proprietary than CUDA. 2 times in the scene
- ▶ 1:10:37 Shawn Wang That they are training or benchmarking their models for writing CUDA kernels, right? 2 times in the scene
DeepWiki: The GitHub Encyclopedia
- ▶ 29:43 Silas Alberti And CUDA channels are just so cool because you can just, like, compare the Python implementation with the CUDA implementation. 2 times in the scene
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 1:08:14 William Beauchamp He kind of explained to me, he was like, look, if you know hardware really well, you can write the CUDA kernels really well.
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 11:48 Yining Zhang And currently, even you use something like, uh, kuda or kublas, you cannot support that.
- ▶ 19:32 unnamed speaker The kernels that they come with, I'm yet to see folks do better than what NVIDIA can do when it comes to CUDA kernels.
- ▶ 28:58 unnamed speaker The MLA technique, for example, that, that Inang mentioned, or the, the CUDA kernels that, that are being used.
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
- ▶ 5:57 Andrej Karpathy ...in PyTorch of LayerNorm, because it's kind of like a block and eventually pulls into some CUDA kernels.
- ▶ 10:39 Andrej Karpathy So this is where we go to the dev CUDA part of the repo, and we start to develop all the kernels. 9 times in the scene
- ▶ 13:02 Andrej Karpathy Okay, so next what happened is, uh, I was basically struggling with the CUDA code, uh, a little bit, and I was reading through the book, and I was implementing all these CUDA kernels, and they're like, okay CUDA kernels, but they're not,… 3 times in the scene
- ▶ 14:01 Andrej Karpathy Ok, so we've converted all the layers to CUDA.
- ▶ 20:56 Andrej Karpathy I think also the C++ Cuda fork is quite nice, and so a lot of folks also encourage you to also fork LN.C.
- ▶ 22:25 Andrej Karpathy But actually, don't you want to write all code in custom CUDA kernels and so on?
Personal benchmarks vs HumanEval - with Nicholas Carlini of DeepMind
- ▶ 18:06 Nicholas Carlini And in order to understand whether or not the solution was right, you just have to trust me on it because you know, like I got my machine in a state that like CUDA was not talking to whatever, some other thing, the versions were…
Answer.ai & AI Magic with Jeremy Howard
- ▶ 24:10 Jeremy Howard Because again, he quite correctly identified, like, you know, this is a way that not everybody has to use CUDA.
- ▶ 39:25 Jeremy Howard And, you know, because it's like, the interesting bits are all written in CUDA, it's hard to like, to step through it and see what's happening.
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 12:22 unnamed speaker They talked about the batching, GPU utilization, memory, like utilization, all that stuff, like CUDA optimizations, their whole training recipe, um, what they released performance stuff.
- ▶ 1:07:53 unnamed speaker So like maybe there's a server server in West coast, East coast, that's different GPUs, different CUDA kernels, different drivers. 2 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
- ▶ 1:38:18 Shawn Wang And from that, uh, I've been using it to translate, um, Andrei Karpathy's, um, LLM-II.py to LLM-II.c, and it needs to write about a raw C code and test it, um, debug, you know, memory issues and CUDA issues and all that.
Breaking down the OG GPT Paper by Alec Radford
- ▶ 21:23 unnamed speaker Yeah, AlexNet did definitely use GPUs, but back then I think he were writing, you know, Q the code.
Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
- ▶ 27:11 Ethan Sutin It was very complicated, um, like, because we were using, like, local whisper, local models, and, like, getting it to work on CUDA, Mac, Windows.
Making Transformers Sing - with Mikey Shulman of Suno
- ▶ 28:39 Alessio Fanelli But, it's funny that it knows about CUDA chorus.
Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
- ▶ 9:09 Soumith Chintala Like, generating the, the final CUDA code or, like, CPU code.
- ▶ 28:39 unnamed speaker Thing here, like, uh, I think, like, Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster, like, you know, hardware kernels in general.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 46:11 Ben Firshman So COG is really just a Docker container that attaches to a CUDA device if it needs a GPU that has a open API specification as a label on the Docker image. 2 times in the scene
- ▶ 53:51 Ben Firshman Like COG is really designed around servers and attaching to CUDA devices and, and NVIDIA GPUs and this kind of thing.
Building an open AI company - with Ce and Vipul of Together AI
- ▶ 1:09:21 Vipul Ved Prakash Uh, CUDA and Kernel Hacker, we have lots of exciting projects. 2 times in the scene
The Four Wars of the AI Stack - Dec 2023 Recap
- ▶ 32:24 unnamed speaker Like, uh, Fireworks, um, recently announced, uh, Fire Attention, where they wrote a custom cruder kernel for mixed drawl, uh, on H 100.
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 13:04 Dylan Patel You know, the forever shouldn't be that the person doing, you know, you know, like PyTorch-level code, right, that high, should also be writing custom CUDA kernels, right? 2 times in the scene
The End of Finetuning — with Jeremy Howard of Fast.ai
- ▶ 1:16:13 Jeremy Howard But, not everybody's got his ability to, like, be like, oh, well, I, I, I'm confident enough in Kuta and or Triton. 2 times in the scene
RWKV: Reinventing RNNs for the Transformer Era
- ▶ 4:58 Eugene Cheah You're going to train it directly with CUDA and so on because it just performs much better. 2 times in the scene
- ▶ 1:43:37 Eugene Cheah Oh, uh, especially, uh, especially Qt, Qt GPUs tends to work really well when they align to the batch size of multiples of .
FlashAttention-2: Making Transformers 800% faster AND exact
- ▶ 11:39 Tri Dao So, um, maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.
- ▶ 33:23 unnamed speaker And since it's a NVIDIA library, can you only run this on like CUDA runtime?
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 8:42 George Hotz TPUs are a lot closer to what I'm talking about than, than like, like CUDA. 2 times in the scene
- ▶ 17:05 George Hotz When I do A times B in PyTorch, it's going to launch a CUDA kernel to do A times B. 4 times in the scene
- ▶ 24:03 George Hotz I have, I have a GitHub repo called CUDA IO control sniffer. 4 times in the scene
- ▶ 41:39 George Hotz It's a compute cluster, which is totally legal under the CUDA license agreement.