Movva: Screening performance engineers for CUDA experience is a red herring
“I don't look for lots of AI experience. I don't look for, you know, CUDA experience at all. That's actually a huge red herring. I mean, CUDA as a concept or GPS as a concept have evolved so much in the last five years. There's no point asking for 10 years of e…”
Thompson: NVIDIA's CUDA moat is dramatically diminished for AI models
“CUDA's moat is dramatically diminished because the models don't care what they run on, and that's what actually matters, what's built on top of the models, but it still matters. It, it's still something of a mode.”
Lindgren: Endra CTO spent a year building an LLM from CUDA up
“And what he did was that he built his, he spent a year building his own LLM from CUDA level up to the chat interface.”
NVIDIA Rubin will shift inference engineering toward traditional hardware infrastructure challenges
“I think that themes around like KV cache offloading, KV aware routing, and disaggregation are going to be substantially more important in the Rubin era, which means that inference engineering becomes not just a like CUDA kernel problem, but also like a very tr…”
Lemkin: Open-source AI models will bypass Nvidia's CUDA platform
“He's got to go frontier and open and open is dangerous. Open doesn't need CUDA. Open is cheaper and it's lower margins. Open will bypass him, but he's got it.”
Feldman: Nvidia CUDA lost 70% of frontier AI model training market share
“I think two years ago every state of the art model was trained in a Cuda flow. And right now, Gemini is trained without Cuda. Anthropical is trained without Cuda. Open AI as strange as could. So in a one or two year period, they lost 70% share. Of training mod…”
Balaban: NVIDIA's real software moat is cuDNN, not just CUDA
“One of the big moats they've got is just The QDNN stack. It's not just CUDA. It's, you know, CUDA is sure. That's like the water we all swim, but like CUDNN has got so many, you know, matrix multiplication, routine optimizations baked into it.”
Distributing CUDA for free on GPUs severely depressed NVIDIA's gross margins
“The answer was, let's use GeForce, which is the GPU that is now everywhere in the world used for playing video games. And let's have GeForce carry on its back CUDA to every single computer in the world. Now, of course, by doing so, our gross margins would go f…”
Karpathy: AutoResearch is limited strictly to domains with easily evaluatable objective metrics
“This is extremely well suited to anything that has objective metrics that are easy to evaluate. So for example, like writing kernels for more efficient CUDA, you know, code for various parts of a model, etc., are the perfect fit. Because you have inefficient c…”
Huang: 40% of Nvidia's business requires CUDA and full AI stack
“About 40% of our business, most people don't realize this, 40% of our business, unless you have the CUDA stack, unless you can build an entire AI factory, you have, the customers don't know what to do with you.”
Huang: Nvidia has radiation-hardened CUDA hardware deployed in satellites
“We're already radiation-hardened. We have CUDA in satellites around the world.”
Rory O'Driscoll: Nvidia customers will not abandon CUDA due to switching costs
“For most users, because they have such dominance, such validation of the CUDA software layer, for most customers, it's going to be too much brain debt to switch from the cheap, you know, the GPU you know and love to something new, right? Because there's probab…”
AI coding assistants degrade quickly on low-level CUDA and PTX code
“Getting Below Triton or level or anything else down to like CUDA or below there's orders of magnitude less public good kernels at that level. And I think that really shows and models capabilities. So when I try and get a model to do something in CUDA or PTX or…”
AI kernel models game automated evaluation metrics with subtly incorrect code
“It would be great, but it's not a silver bullet because kernels are also hard to validate. It's hard to have like an eval in kernels. Every time that someone releases like, oh, we created a new eval that measures kernels and we trained a model that generates k…”
String theorists ramp faster in AI engineering than complacent CUDA developers
“I don't really care if someone knows CUDA kernel writing. If someone does string theory and is really good at understanding complex problems, they will be productive faster than someone who knows CUDA and doesn't really care about trying to get better at it.”
Nvidia's stock fell 80% during its massive initial investment in CUDA
“The company invested so much in converting its GPUs for CUDA compatibility that its gross margin fell from 45 to 35%. At the same time it's increasing its spending on CUDA, the global financial crisis destroyed consumer demand And Nvidia's stock fell by more t…”
In AI Inference, Nobody Cares About Nvidia's CUDA or PyTorch
“In inference, the truth is, nobody cares about CUDA. Nobody even cares about PyTorch. All right. What they want is an API.”
AlexNet Was Trained on Two Off-the-Shelf Nvidia GTX 580 GPUs
“They buy two NVIDIA GeForce GTX five eighties, which were NVIDIA's top of the line gaming cards at the time. The Toronto team rewrites their neural network algorithms in CUDA, NVIDIA's programming language. They train it on these two off the shelf GTX five eig…”
Ross: NVIDIA's software moat applies to training, not inference
“That NVIDIA's software is a moat. Yeah it's true for training, but it's not true for inference.”
Bachman: Manifest Switched Power Retention from Triton to Custom CUDA
“Actually, our initial version of power retention was written in Triton, but we realized quickly that it just didn't offer the flexibility to really squeeze the performance that we wanted out of the GPU. So we took a step back and dove into CUDA.”
Morris: Deep understanding of GPU architecture makes engineers exceptionally hireable
“That said, if you do it, you're, you've gotta be one of the most hireable people in the world. Like if you like, Really deeply understand the architecture of the new GPUs coming out and how to control it. You're in a very small handful of people and like every…”
Modular completely eliminates and replaces NVIDIA's CUDA stack
“In the case of Modular, we go literally, like, we only work at that level, because we get rid of all of CUDA, right? And so we've replaced the entire stack, and so we only do that.”
Mohan: CUDA is not Nvidia's real competitive moat
“Even a company like NVIDIA, I think everyone outside looking in is like, CUDA is the real moat. Or something like that. I think that's just inaccurate, right?”
Gerstner: The US should sell AI chips to China to protect the CUDA ecosystem
“Are we better off selling to those countries, those companies, keeping companies like ByteDance and Tencent, et cetera, in the CUDA ecosystem, rather than allowing all of that data, all of those profits to flow right into the Huawei ecosystem and benefit the C…”
Alberti: Multi-Turn RL Enables Aggressive Code Optimization Over Single-Turn Models
“Basically the single-turn model that was just trained on, like, getting the best result after one turn. It would basically be a little bit, like, too careful, because it couldn't risk writing, like, non-compiling code, whereas, like, the multi-turn model would…”
Gurley: DeepSeek bypassed Nvidia's CUDA framework for low-level optimization
“One, it's validated now that they went around CUDA, and I just think that's interesting.”
Palihapitiya: DeepSeek bypassed Nvidia CUDA software lock-in using PTX
“Everybody is used to building models and compiling through CUDA, which is Nvidia's proprietary language, which I've said for a couple of times is their biggest moat, but it's also the biggest threat factor for lock-in. And these guys worked totally around CUDA…”
Manickam: Nvidia's dominance stems from optimizing hardware for its CUDA software
“The reason why NVIDIA is successful is because they actually had this software, AI software, they call it CUDA, right? So now they had that hardware that runs this CUDA very, very efficiently.”
CUDA Investments Pushed Nvidia Gross Margins Down to 35 Percent
“It took four years and the company invested so much in converting its GPUs for CUDA compatibility that its gross margin fell from 45% to 35%.”
Nvidia Stock Dropped Over 80 Percent Amid 2008 CUDA Expenditures
“As NVIDIA increased spending on CUDA, the global financial crisis destroyed consumer demand, and NVIDIA's stock price fell by more than 80% from October, 2007 to November, 2008.”
Fu: Changing one PyTorch line requires a week of CUDA development
“If we decided to change one thing in PyTorch, like one line of PyTorch code is like a week of CUDA code at least.”
Huang: NVIDIA Improved Hopper Performance on LLaMA 5x in One Year
“CUDA made it possible for us to iterate so quickly, just in the last year, and then we just went back and benchmarked when Lama first came out, we've improved the performance of hopper by a factor of five without the layer on top ever changing.”
Gerstner: NVIDIA's CUDA library has over 300 industry-specific acceleration algorithms
“The CUDA library now has over 300 industry specific acceleration algorithms, right? Where they deeply learn the industry, right? So whether this is synthetic biology or this is image generation, or this is autonomous driving, they learn the needs of that indus…”
Madra: Fewer developers will touch CUDA long-term, weakening NVIDIA's software moat
“I think there's going to be fewer people touching that. And I do think that's a point where they're the moat is not as strong as a longer term, as you say, and think about like, you know, the way the analogy that I would go with is like, think about the number…”
Gurley: NVIDIA's competitive advantage is strongest at massive system scale
“NVIDIA's competitive advantage is strongest where the size of the system is largest, which is another way of saying what Renee said. It's flipping it on its head. It's not to say it's weak on the edge, but it's super powerful when you put a whole bunch of them…”
Madra: NVIDIA's CUDA moat does not exist for inference workloads
“There is no
Tie into CUDA that's required to go faster.
That's required to get the models running, right?
Obviously none of the three companies run CUDA.
And so that moat doesn't exist around inference.”
Karpathy: The popular PMPP textbook lacks advanced CUDA optimization techniques
“PMPP is actually quite good but also, I think, still kind of like mostly on the beginner level, because a lot of the CUDA code that we ended up developing in the lifetime of the LMC project, you would not find those things in, in this book, actually.”
Chamath: His firm is building a transpiler to bypass NVIDIA hardware lock-in
“So as part of eighty-ninety, one of the things that we're doing is we're building a transpiler.”
Noone: Zoo built the first GPU-optimized cloud CAD geometry engine using CUDA
“So our view was we would start with this geometry engine. That's the world's first GPU optimized cloud-based API accessible CAD engine. We built that from scratch. So that's CUDA, which is kind of GPU level machine language CUDA level implementation of those c…”
Mayya: AMD cannot compete with Nvidia's CUDA architecture for LLMs
“Everything in AI is going to require NVIDIA, and AMD is not, AMD is good, but it's not like, it can't compete with CUDA. It's good for consumer devices, right? But somebody building the next LLM, you're locked into Nvidia for the time being, right?”
Sutin: Local model setup friction killed Owl AI's open-source developer adoption.
“I learned, like, we did not make the developer experience very good. It was very complicated like, because we were using, like, local whisper, local models, and, like, getting it to work on CUDA, Mac, Windows. We didn't do a good job, so it was very difficult …”
Srivastava: The ease of running CUDA workloads on AMD is significantly overstated
“I've, I personally think that it's pretty overstated how easy it is to run something that looks like CUDA or CUDA in some form on an AMD chip seems, seems like a challenge to me.”
Doshi: AI companies would leave NVIDIA if cheaper compute alternatives existed
“It's not CUDA that's keeping, I think, keeping a lot of us. It's actually that there is nothing really dramatically better than NVIDIA's GPUs. And so if there's nothing dramatically better than, I mean, the reality is the cost for training and inference are so…”
Lamini has achieved software parity on AMD GPUs with CUDA
“We have reached software parity with essentially CUDA.”
Howard: Mojo-like languages will unlock thousands of FlashAttention-scale breakthroughs
“There is a thousand flash attentions out there for us to build. You just got to make it easy for us to build them. So like Triton definitely helps. But it's still, Not easy. You know, it still requires kind of really understanding the VPU architecture, writing…”
NVIDIA's programmable shaders were the first massively parallel processors
“And if you just looked at the pipeline of a programmable shader, it is a processor and is highly parallel, and it is massively threaded, and it is the only processor in the world that does that.”
Huang: CUDA is used across almost all fields of scientific research
“CUDA is used for almost all fields of science. Everything from molecular dynamics to imaging, CT reconstruction to seismic processing to, you know, weather simulations, quantum chemistry, the list goes on, right?”
Huang: GeForce NOW was NVIDIA's first data center product, preceding CUDA supercomputing
“GeForce Now was NVIDIA's first data center product. And our second data center product was remote graphics, putting our GPUs in, in the world's enterprise data centers, which then led us to our third product, which combined CUDA plus our GPU, which became a su…”
Huang: NVIDIA is only accelerator maker with universal architectural compatibility
“We are the only accelerator on the planet where every single accelerator is architecturally compatible with the others. None has ever existed.”
Huang: NVIDIA has 250M to 300M active compatible CUDA GPUs globally
“There are literally a couple of hundred million, right? 250,000,300 million installed base of active CUDA GPUs being used in the world today, and they're all architecturally compatible.”
AlexNet was trained on two consumer Nvidia GeForce GTX 580 GPUs
“And what these guys from Toronto did is they went out probably to their local Best Buy or equivalent in Canada. They bought two GeForce GTX-Five-Eighty's, which were the top-of-the-line cards at the time, and they wrote their algorithm, their convolutional neu…”
Nvidia's CUDA platform reached four million registered developers by May 2023
“If you look at the number of CUDA developers over time, it was released in 2006, It took four years to get the first 100,000 people. Then by twenty-sixteen, 13 years in, they got to a million developers. Then just two years later, they got to two million. So 1…”
AMD lacks Nvidia's TSMC advanced packaging capacity and CUDA developer ecosystem
“AMD doesn't have all this capacity reserved from TSMC, at least not for the 2.5 D packaging process for the high end GPUs. AMD doesn't have the developer ecosystem from CUDA.”
Nvidia has an installed base of 500 million CUDA-capable GPUs globally
“Today there are five hundred million CUDA capable GPUs for developers to target.”
Compilers will automate complex kernel fusion within two years
“Maybe in a year or two, we'll, we'll have compilers that are able to do a lot of these optimizations for you, and you don't have to, for example, spend a couple months writing CUDA to get this stuff to work.”
Huang: NVIDIA made every chip CUDA-compatible despite few initial customers
“And so for the first five, 10 years, you know, we had very few customers for CUDA, but we made every chip CUDA compatible.”
Gilbert: CUDA was not a useful platform for six-plus years after 2006
“Because while CUDA development began in 2006, That was not a useful usable platform for six plus years at NVIDIA.”
Gilbert: NVIDIA has 1,100 employees with 'CUDA' in their LinkedIn title
“I searched LinkedIn for people who work at NVIDIA today and have the word CUDA in their title. There are 1100 employees dedicated specifically to the CUDA platform.”
Gilbert: NVIDIA's investment in CUDA was an iPhone-sized bet
“Those were big bets relative to the company's size at the time, but this bet is like an iPhone-sized bet.”
Rosenthal: NVIDIA has never charged a dollar for CUDA
“NVIDIA to this day, now this may be changing, we'll talk about this at the end of the episode, has never charged a dollar for CUDA”