CUDA, every mention
317 scenes across 22 shows · ← back to CUDA
Latent Space 117
Acquired 100
the MAD Podcast 40
BG2 Pod 34
TBPN 30
No Priors 29
20VC 25
Founders 2214 more shows
every year every show
Latent Space 117
Acquired 100
the MAD Podcast 40
BG2 Pod 34
TBPN 30
No Priors 29
20VC 25
Founders 22
How I Built This 21
All-In 21
Big Technology 17
the Neon Show 15
the Y Combinator Startup Podcast 12
the a16z Podcast 9
Cheeky Pint 6
Invest Like the Best 4
WTF is with Nikhil Kamath 3
In Depth 2
A Product Market Fit Show 2
Sourcery 2
My First Million 1
the Startup Ideas Podcast 1
Verbatim, from the transcripts: passages where CUDA comes up on Latent Space, Acquired, the MAD Podcast, BG2 Pod, TBPN
When AI Improves Itself | Richard Socher (Recursive)
- ▶ 1:08:54 Richard Socher Um, we also showed that they can build new CUDA kernels, which is very useful for faster inference.
Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
- ▶ 1:09:55 Neil Movva I don't look for, you know, CUDA experience at all. 2 times in the scene
What Happens When the AI Boom Runs Out of Money · Invest Like The Best
- ▶ 1:19:26 Ben Thompson CUDA's moat is dramatically diminished because the models don't care what they run on, and that's what actually matters, what's built on top of the models, but it still matters.
He spent $150K on brand before he had a product—closed $1.5M ARR in 1 month. | Endra
- ▶ 15:19 Niklas Lindgren And what he did was that he built his, he spent a year building his own LLM from CUDA level up to the chat interface.
How The AI Bet Pays Off + AI Lab Strategy Game — With David Cahn
- ▶ 36:42 unnamed speaker You look at CUDA, you look at sort of the ecosystem he's built around the chip.
- ▶ 55:28 David Cahn Kuda's is amazing. 2 times in the scene
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 1:01:52 Philip Kiely Um, so I think that themes around like KV cache offloading, KV aware routing, and, and disaggregation are going to be substantially more important in the Rubin era, which means that inference engineering becomes not just a like CUDA kernel…
- ▶ 1:05:20 Ali Taha I can write CUDA to control it and change its operations.
Multi-GPU Kernels, Intelligence per Watt, Heterogeneous Inference, and More | YC Paper Club · Y Combinator
- ▶ 1:06 Francois Chaubard You can't do that with inference, and there's so many things like that that we'll talk about, and there's so much juice left to squeeze on the CUDA side, on the kernel side,
- ▶ 7:30 Stuart (Stu) Um, one quick note before I begin, um, the final deliverable of this paper is a CUDA framework that we refer to as parallel kittens, but the goal of today's talk is not to promote my open source library, but rather to convey the set of…
- ▶ 19:14 Stuart (Stu) So Parallel Kittens builds on all the trade-offs and principles that we discussed, and it's a highly opinionated set of CUDA programming primitives that extends Thunder Kittens, which is one of our previous works for single GPU kernels.
- ▶ 33:03 unnamed speaker Ultimately, these dispatch down to individual CUDA kernels that may or may not be very fast on modern hardware, which is why we get programming languages like Triton, which are a tile-based. 2 times in the scene
- ▶ 34:44 unnamed speaker And at least, like, for this archetype of people that just want state-of-the-art perf, as far as I can tell, like, they really prefer, like, CUDA, and it was only when we were doing lots of gem-related problems that people loved using…
- ▶ 35:32 unnamed speaker I've never written an open and CUDA book before.
- ▶ 37:23 unnamed speaker I'm gonna make you, if you've never written a CUDA kernel before, you're gonna do one with me right now.
- ▶ 1:14:36 unnamed speaker Like, I see a lot of programming language systems are kind of trending towards, okay, it's still the Cuda programming model, but it's Python syntax on top, which is great, like, simplifies things, easier to look at, but it doesn't actually…
Jensen's Open-Weights Letter | Google Cloud Grows 82% But The Market Tanks
- ▶ 3:16 Jason Lemkin Open doesn't need CUDA.
RSI Is Closer Than People Think, Per Tae Kim
- ▶ 25:07 unnamed speaker Talk about the NVIDIA CUDA mode. 5 times in the scene
Will Open-Source Threaten Anthropic's Business & Do Margins Matter in a World of AI | Matt Murphy
- ▶ 1:03:31 Matt Murphy To, uh, kind of obfuscate the underlying chips and technology stacks like CUDA, etc.
Cerebras CEO: Why GPUs Can't Do Fast Inference
- ▶ 1:02 Matt Turck We started from what is a wafer and built up step by step, why GPUs struggle with fast inference, the three shortages nobody talks about, the decade in the desert when nobody wanted this chip, and why Andrew believes that CUDA is no longer…
- ▶ 1:01:11 Matt Turck Cuda as well. 5 times in the scene
Former Intel CEO on What Went Wrong, What's Next + Lovable CEO on the Real Promise of Vibe Coding
- ▶ 8:43 Pat Gelsinger You know, it was always the big CPU and those little GPUs, but when they started to build a real software stack with it, right, you know, sort of, okay, this CUDA thing and SIMT as a technology, you know, so, uh, you know, uh,… 2 times in the scene
Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
- ▶ 18:07 Bryan Catanzaro Follows through over long time periods, you know, uh, and I've seen that with CUDA.
- ▶ 30:33 Bryan Catanzaro You know, we followed through over 10 plus years with CUDA, and we're doing that with Nemo Tron now.
The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
- ▶ 52:22 Rob Wachen Not to support arbitrary PyTorch, not to support arbitrary CUDA, not to support arbitrary Onyx graphs, but instead we envisioned a world where there was going to be under a hundred models that actually mattered, and they were all going to…
The GPU Myth: State of AI Compute 2026 | Stephen Balaban
- ▶ 27:11 Stephen Balaban It's not just CUDA. 2 times in the scene
Re-engineering the Semiconductor Supply Chain with Intel CEO Lip Bu Tan
- ▶ 31:26 Lip-Bu Tan You know, he focused on Kodak.
We Need An Ecosystem in AI, And Every Company Can Win A Place In It
- ▶ 15:14 Satya Nadella Adobe built Autodesk built, uh, or even like take what Jensen said, we built DX and he built, you know, CUDA on top of it.
Satya Nadella on AI: @NoPriorsPodcast x Latent Space Crossover Special at Microsoft Build 2026
- ▶ 14:31 Satya Nadella We built DX and he built, you know, CUDA on top of it.
NVIDIA: Jensen Huang. From near collapse to becoming the world’s biggest company
- ▶ 3:09 Guy Raz Jensen poured billions of dollars into developing a platform and software layer called CUDA, which is basically the instruction manual that lets you use all those line cooks in entirely new ways. 4 times in the scene
- ▶ 30:15 Jensen Huang And it was during that time, um, other types of techniques for general purpose computing was coming along, uh, which ultimately led to what we now call CUDA.
- ▶ 30:24 Guy Raz All right, so let's jump into this, because you launched this project called CUDA, or C-U-D-A, in roughly, and to put this in, in very basic terms, CUDA is this platform that makes your graphics chips a lot more versatile, right? 12 times in the scene
- ▶ 37:35 Jensen Huang Well, we, at that time, uh, several different groups reached out to us to ask us for help on using CUDA to accelerate deep learning, and the reason for that was because there was a contest coming up for computer vision called ImageNet, and… 4 times in the scene
Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud
- ▶ 27:09 Tuhin Srivastava Like, how good they are at that CUDA, how good CUDA is, the developer ecosystem around it, um, 2 times in the scene
The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition
- ▶ 21:23 Peter Ludwig and you have a start with an assumption that, oh, I'm gonna, I'm gonna use CUDA and I'm gonna run this, uh, on an NVIDIA chip, then you don't really have to think about the hardware in that sense. 2 times in the scene
AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
- ▶ 1:02:50 Mikhail Parakhin And even Liquid, we had to work a lot with NVIDIA and to, because almost everything is not designed in CUDA for, or in, in the current stack for, for low latency.
Jensen On The Ropes, Sam Altman’s Conflicts, Allbirds’ GPU Pivot
- ▶ 6:52 Ranjan Roy I mean, he said CUDA over and over again.
Jensen on Dwarkesh, Cursor x XAI, Netflix Stock Sinks | Diet TBPN
- ▶ 1:21 unnamed speaker Somebody clearly made this video just because they're enthusiasts of Fast and the Furious, and then, uh, the semi-analysis team was able to quickly recontextualize that to be about, uh, Jensen's answer on, uh, NVIDIA's moat, and the CUDA…
- ▶ 2:45 unnamed speaker Uh, but once the AI boom kicked off, the CUDA ecosystem significantly sped up development of AI systems and training of AI models. 5 times in the scene
- ▶ 21:37 unnamed speaker Uh, clearly, uh, with, with all of this backdrop and just the idea of more chips and maybe the CUDA ecosystem being something that you can work around, uh, can an American fab 2 times in the scene
Arm’s $15B Chip Bet, Sanders & AOC vs Datacenters, Meta & YouTube Lose Trial | Diet TBPN
- ▶ 9:48 Turner Novak The X-AVI moat is not as strong as the CUDA moat, but there's still this like dynamic of, of NVIDIA and ARM are going up against Intel and AMD in the same way that, uh, different GPU makers are going up against the CUDA moat. 2 times in the scene
Skill Issue: Andrej Karpathy on Code Agents, AutoResearch, and the Loopy Era of AI
- ▶ 23:49 Andrej Karpathy So for example, like writing kernels for more efficient CUDA, you know, code for various parts of a model, etc., are the perfect fit.
Jensen Huang: Nvidia's Future, Physical AI, Rise of the Agent, Inference Explosion, AI PR Crisis
- ▶ 43:28 Jensen Huang And so, there's a whole, whole part of our market, about 40% of our, of our business, most people don't realize this, 40% of our business, unless you have the CUDA stack, unless you can build an entire AI factory, you have, the customers…
Nvidia Restarts China Sales, Vibe Coding Backlash, Peptide Craze | Diet TBPN
- ▶ 3:54 unnamed speaker Uh, the better argument might just be dependent on a functional TSMC fab, but, uh, there is, there are benefits to the CUDA ecosystem and to the, the idea that whatever models get built there will be applicable here.
AI vs. Dog Cancer, Timothée Chalamet Under Fire, ‘Agents Over Bubbles' | Diet TBPN
- ▶ 13:59 Jensen Huang This is the 20th anniversary of CUDA. 2 times in the scene
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 11:03 Kyle Kranen They don't know what CUDA is. 2 times in the scene
Where SMALL models will Win | Sudarshan kamath, Smallest ai
- ▶ 23:24 Sudarshan Kamath Um, and in our, like WhatsApp chats, uh, my co-founder was asking about, oh, has anyone trained this model and like this issue we are facing some QDA errors, et cetera.
The SaaS Apocalypse: Who Lives & Who Dies | Insight Partners Co-Founder, Jerry Murdock · 20VC with Harry Stebbings
- ▶ 9:02 Jerry Murdock It's about making sure that CUDA can also support ASICS chips, and they're going to have to go out and make sure that they can get, you know, they know what's coming. 3 times in the scene
Reiner Pope of MatX on accelerating AI with transformer-optimized chips
- ▶ 34:31 John Collison So there's an art about there that NVIDIA, a huge part of the defensibility comes not from the chips, which are good, but from the, uh, software layer and the ability for engineers to write these really parallel workloads, uh, and the fact… 3 times in the scene
Is AI changing EVEN the Top Roles in startups?
- ▶ 48:08 Umesh Padval So he started with CUDA, which is a brilliant move because now you get
Dylan Patel: NVIDIA's New Moat & Why China is "Semiconductor Pilled”
- ▶ 10:07 Matt Turck And do you think, uh, CUDA is going to remain that mode? 15 times in the scene
- ▶ 39:10 Dylan Patel And now does that like weaken the CUDA moat?
FULL INTERVIEW: Dylan Patel Says We’re Still Underestimating AI
- ▶ 27:14 unnamed speaker The, the Ben Thompson line was something like, ah, he's, he's okay selling chips because he wants dependency on the NVIDIA ecosystem, CUDA, but he would ban, ah, lithography tools from going to China, and I'm always, I've, I've been…
AI’s Steve Jobs?, Big Tech AI Chaos, 2026 Crystal Ball
- ▶ 6:53 M.G. Siegler They have, of course, software layers with CUDA and things like that.
World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI
- ▶ 15:54 Pim de Witte so I was very, very familiar with CUDA and like the GPU side, and all the video infrastructure that we were using for this stuff, but the modeling side itself was, was still quite foreign.
ChatGPT Turns Three, OpenAI x Thrive, David Sacks vs. The New York Times | Diet TBPN
- ▶ 17:39 unnamed speaker They are asking, is this potentially the end of the CUDA moat? 2 times in the scene
Anthropic Raises $30B from Microsoft & NVIDIA & NVIDIA’s Core Business Faces TPU Threat · 20VC with Harry Stebbings
- ▶ 8:49 unnamed speaker TPU that can also be rolled out to everyone with all the cutest support that Nvidia has, but even if it's just used internally, first of all, and I can save that twenty billion dollars of profit, hell, I got to look at that if I'm, if I'm… 2 times in the scene
What 20 Years of Investing taught Somesh Dash of IVP after backing Perplexity, Figma, Dropbox
- ▶ 26:17 Somesh Dash The growth engines of the top companies of S&P 500 carry the markets around the world, and if you look at NVIDIA as an example of this, you know, there's one argument that's saying if you look at the PE ratio, NVIDIA is very expensive,…
- ▶ 43:00 Somesh Dash And so they're all going to be still using NVIDIA chips and CUDA, like all the usual stuff's going to go in there.
Amjad Masad & Adam D’Angelo: How Far Are We From AGI?
- ▶ 47:11 Amjad Masad Um, and the main idea there was that if we put a verify on the loop, I remember reading DeepSeq, uh, a paper from Nvidia about how they, um, used DeepSeq to write CUDA kernels and they were able to run DeepSeq for like 20 minutes.
How Larry Ellison Thinks
- ▶ 40:26 David Senra This is again, Jensen did this at NVIDIA and with, with CUDA and GPUs.
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 8:42 unnamed speaker Because I think the other question is, like, well, I'm gonna do all this work versus, like, I just write CUDA code that then, when the BG-G-Hundred comes online for my cluster, I'll just switch it over right away.
- ▶ 10:23 unnamed speaker And I think obviously if you understand CUDA is like, you know, 4 times in the scene
- ▶ 13:36 Quentin Anthony Below Triton or level or anything else down to like CUDA or below, um, there's orders of magnitude less public good kernels at that level. 4 times in the scene
- ▶ 22:42 Quentin Anthony So like it's very, Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like 2 times in the scene
- ▶ 49:00 Quentin Anthony I don't really care if someone knows CUDA kernel writing. 2 times in the scene
"Is there an AI bubble?” Gavin Baker and David George
- ▶ 24:08 Gavin Baker You know, not, it was a semiconductor company, then a software company with CUDA, now a systems company with these rack-level solutions.
How Jensen Works
- ▶ 47:44 David Senra This is when they create the new programming model called CUDA, which stands for Compute Unified Device Architecture. 10 times in the scene
Cerebras CEO, Andrew Feldman on Why Raise $1BN and Delay the IPO & Why NVIDIA’s Worried About Growth · 20VC with Harry Stebbings
- ▶ 24:59 Andrew Feldman In inference, the truth is, nobody cares about CUDA.
Tobi Lütke is still captivated by internet commerce, 20 years later
- ▶ 1:06:01 Tobi Lütke And, I mean, this is in no area more, more true than, ah, in the NVIDIA case, where, like, I mean, these CUDA cores, the Tensor cores, were around on the, they, they, they, they were on these cards, which, which everyone, they were part of… 2 times in the scene
- ▶ 1:07:36 John Collison And kind of like you're saying with CUDA, you know, it's hard to justify at that moment in time, but it seems like the right technical decision.
Google Part III: The AI Company. Google is amazingly well-positioned... will they win in AI? (Audio) · Acquired
- ▶ 50:33 David Rosenthal The Toronto team rewrites their neural network algorithms in CUDA, NVIDIA's programming language.
- ▶ 3:28:41 Ben Gilbert So if you're willing to not use CUDA and build on Google stack, they have an abundant amount of TPUs for you.
- ▶ 3:31:45 Ben Gilbert They're trying to build an ecosystem around their chips the way that CUDA does, and you're only gonna credibly be able to do that if your chips are accessible in anywhere that someone's running their existing workloads.
Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI · Y Combinator
- ▶ 14:47 Ankit Gupta But not necessarily at the level of abstraction of, you know, writing custom CUDA kernels, or like was that also in the space where you guys were thinking about things?
- ▶ 51:16 Nick Joseph I was, like, working at the torch.matmol, but, like, I didn't know CUDA.
Groq Founder, Jonathan Ross: OpenAI & Anthropic Will Build Their Own Chips & Will NVIDIA Hit $10TRN · 20VC with Harry Stebbings
- ▶ 1:22:06 Harry Stebbings Cuda locking is bullshit. 2 times in the scene
NVIDIA: OpenAI, Future of Compute, and the American Dream | BG2 w/ Bill Gurley and Brad Gerstner · Bg2 Pod
- ▶ 36:09 Jensen Huang We invented CUDA, invented GPUs, and we invented the idea of co-design at a very large scale.
- ▶ 43:37 Jensen Huang If not for the fact that CUDA is easy to operate on and iterate on, how do they try all of their vast number of experiments to decide which one of the Transformer versions, what kind of 2 times in the scene
- ▶ 1:02:51 Jensen Huang Sriram is the only person in Washington DC that I think knows CUDA, um, and, and, which is strange anyways, but, but I, I just love the fact
⚡️ Beyond Transformers with Power Retention
- ▶ 8:14 unnamed speaker And so you're also releasing, um, just to go from the bottom up, uh, Vidrill, which is a framework for CUDA kernel. 4 times in the scene
- ▶ 29:53 unnamed speaker If you are excited by those ideas, especially of a strong background in deep learning, uh, CUDA programming, or anything mathematically adjacent to those ideas, please reach out.
A Technical History of Generative Media
- ▶ 10:45 unnamed speaker Kuda kernel specialist at the time, right? 2 times in the scene
From Chrome extension to $5B platform | Postman’s journey | Abhinav Asthana (Co-founder & CEO)
- ▶ 18:23 Abhinav Asthana If you can imagine, we were working with CUDA back then. 2 times in the scene
⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
- ▶ 22:05 Thomas Sohmers You know, the way that we say that we're in the CUDA ecosystem, it's by nature of the fact that we are ingesting the, the output of everything that exists in, in the CUDA infrastructure. 2 times in the scene
Weekly Recap: Sam vs. Elon, Cluely's Crazy Ad, UFC x Paramount, Apple's AI Robot
- ▶ 1:31:43 Delian Asparouhov For sure, some of it propped up by like CUDA and they're like, you know, sort of software side of the house, but like NVIDIA is the one that is performing the best of all of those.
Greg Brockman on OpenAI's Road to AGI
- ▶ 52:14 Greg Brockman Cuda kernels are a good example of a very self-contained problem that actually our models should get very good at very soon, but it's just difficult because it requires a lot of domain expertise, a lot of like real abstract thinking. 2 times in the scene
Weekly Recap: Casey Neistat, OpenAI Cracks Math, The Future of ChatGPT, Apple x F1, Intel Layoffs
- ▶ 1:20:21 John Coogan Is there something that's coming down the pipe that could disrupt the NVIDIA GPU monopoly, the CUDA ecosystem, or even the
Winning the AI Race Part 3: Jensen Huang, Lisa Su, James Litinsky, Chase Lochmiller
- ▶ 50:43 Jensen Huang The reason for that is because CUDA is so programmable, and we're constantly, the whole world, not just us, the whole world is doing open source development, improving its effectiveness.
Information Theory for Language Models: Jack Morris
- ▶ 11:20 Jack Morris Yeah, I'll comment on that quickly because if someone has been listening to this and also following me online for a while, I think I've made a couple of comments like saying something like you shouldn't learn about CUDA or things to that… 5 times in the scene
IPOs and SPACs are Back, Mag 7 Showdown, Zuck on Tilt, Apple's Fumble, GENIUS Act passes Senate
- ▶ 22:25 Chamath Palihapitiya Which is to say, can you take a CUDA workload, and then can you redirect it away from NVIDIA to different hardware?
The Shape of Compute (Chris Lattner of Modular)
- ▶ 3:10 Chris Lattner And I said, oh, by the way, we can't use CUDA.
- ▶ 8:24 Chris Lattner Also, let's not just build, uh, you know, air quotes, a CUDA replacement. 3 times in the scene
- ▶ 11:47 Chris Lattner It doesn't use CUDA. 3 times in the scene
- ▶ 30:10 Chris Lattner Massively hand coded CUDA kernels and all this stuff for any new thing.
- ▶ 41:30 Chris Lattner But I will admit that their time to market and revenue growth and stuff like that has been much faster because they didn't have to like build an entire replacement for CUDA to get there. 2 times in the scene
- ▶ 52:54 Alessio Fanelli And I think specifically in your case, you know, they worked at the BTX layer of the GPU, which is like even lower and more proprietary than CUDA. 2 times in the scene
- ▶ 1:10:37 Shawn Wang That they are training or benchmarking their models for writing CUDA kernels, right? 2 times in the scene
China, AI Immigration, Rare Earths & Chips, Tariffs, Markets | BG2 w/ Bill Gurley & Brad Gerstner · Bg2 Pod
- ▶ 39:15 Brad Gerstner It keeps them in that CUDA ecosystem, allows NVIDIA to compete, and I think slows down their ability to run the table around the rest of the world.
Windsurf CEO & Co-Founder, Varun Mohan: AI's Biggest Acquisition to Date! · 20VC with Harry Stebbings
- ▶ 21:44 Varun Mohan And I think maybe an example of this that I like to bring up is even a company like NVIDIA, I think everyone outside looking in is like, CUDA is the real moat. 2 times in the scene