Nvidia, every mention
349 scenes, the whole family · ← back to Nvidia
tap a year for its mentions
every year anyone Shawn Wang 80Dylan Patel 55Kyle Kranen 43Ali Taha 26Ethan He 24Chris Lattner 24George Hotz 23Philip Kiely 20Doug O'Laughlin 19Sarah Chieng 17
Verbatim, from the transcripts: the passages where Nvidia comes up
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
- ▶ 16:30 Sean Lie I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than, um, than the GPU than NVIDIA. 2 times in the scene
- ▶ 21:20 Sean Lie I think it's really awesome to see that, you know, SRAM architectures are becoming, you know, more accessible, you know, even the biggest of the big guys here, NVIDIA, is embracing SRAM design, acknowledging that,
- ▶ 22:06 unnamed speaker Um, I mean, this is separate Rubin, you know, strategy. 3 times in the scene
- ▶ 22:10 Sean Lie Well, so there's, there's, there's Rubin, but, you know, the LPX itself, 2 times in the scene
- ▶ 30:35 Sean Lie What, what I think is, um, the most untapped opportunity right now, uh, frankly for Cerebrus, but frankly for the entire non-NVIDIA, uh, uh, environment, uh, right? 4 times in the scene
- ▶ 31:05 Sean Lie Like, okay, this thing was designed to run on B. 200, GB.
- ▶ 32:00 unnamed speaker I think before getting into etched and super spicy chips, any takes on AMD and NVIDIA, what they're doing? 3 times in the scene
- ▶ 33:21 Sean Lie And so I, AMD, Tranium, in many ways TPU, like all of these in my mind are all trying to build a better Reuben, right?
🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
- ▶ 2:23 Anima Anandkumar Uh, but yeah, just as a brief introduction, uh, you know, I've been working in AI for more than two decades, uh, in a way, you know, before even deep learning, when a lot of the theoretical foundations had to be built for probabilistic…
- ▶ 1:09:03 Anima Anandkumar And so that's where a lot of, like, the exploration was to make this work well in practice and over into Amazon Web Services, then Nvidia.
- ▶ 1:21:25 Anima Anandkumar Because, you know, and, you know, of course our compute that we have is growing so much more than even a few years ago, thanks to NVIDIA, thanks to others.
⏭️ Forward Deployed: Voice AI on what works in 2026
- ▶ 33:41 unnamed speaker The whole point of that is that, as I think today, Pneumotron three-point-five launched, and there's already an ASR, um, benchmark out, did very well, the NVIDIA one.
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 14:49 Philip Kiely You're taking the model from, generally, these models are not released in NVFP four, um, and we want them to be in NVFP four for maximum blackwell compatibility, so we have to perform that quantization, um, and, you know, calibrate the…
- ▶ 23:58 Ali Taha Like oftentimes, um, the image you run, like NVIDIA will release an image, for instance, and if we will upstream the changes from the latest RTLM image into our stack, we'll find that it, it fixes it. 2 times in the scene
- ▶ 30:58 Ali Taha One of our research interns, Joshua, um, I think it's a tweet on, on how we have 20% better quantized JLM five two than NVIDIA.
- ▶ 37:12 Philip Kiely So like on GLM 5.2, um, if you want to get unquantized, uh, perhaps on hoppers even, um, and you're just using an off the shelf inference engine with no particular optimizations, no, no speculator, um, nothing, nothing extra around like KV…
- ▶ 38:45 Ali Taha Like, if you, if you obviously have a thing where you're serving it on just, like, a node of H-one-hundreds, and then you throw, like, you know, you short the model across, like, four nodes of B-to-hundreds.
- ▶ 38:45 Ali Taha Like, if you, if you obviously have a thing where you're serving it on just, like, a node of H-one-hundreds, and then you throw, like, you know, you short the model across, like, four nodes of B-to-hundreds.
- ▶ 41:24 Ali Taha NVIDIA is going to push one out if no one else does. 2 times in the scene
- ▶ 42:14 Shawn Wang I was waiting for a mention on Dynamo. 6 times in the scene
- ▶ 49:49 Ali Taha And it took off and it was implemented on local devices because your, your memory bandwidth is so slow on like a MacBook, for instance, but try putting the same thing on like an Nvidia GPU on a B 200. 2 times in the scene
- ▶ 54:39 Philip Kiely Like, let's say, let's say you're doing a deployment on H-one hundreds for whatever reason, and you're putting a, a trillion parameter model on there. 4 times in the scene
- ▶ 55:08 Ali Taha Like on a B 200 is one 80 gigawatts per GPU, and then a node of eight, you're talking like one 80 times eight.
- ▶ 55:30 Philip Kiely Do you want to tell me about the T- fours? 2 times in the scene
- ▶ 57:03 Philip Kiely If you look at, for example, NVIDIA Nemotron models, they run very, very well on Blackwell.
- ▶ 57:08 Philip Kiely That's, that's unsurprising.
- ▶ 59:00 Ali Taha With the Rubens, I don't know if you guys saw the Rubens Twitter posts yesterday, but they're also, um, Rubens? 5 times in the scene
- ▶ 59:15 Ali Taha One of, one of the tech leads at NVIDIA, like, launched a Twitter post, like, we're pulling the curtain on Ruben, and here's the, here's the specs.
- ▶ 59:38 Philip Kiely Can I speculate about Ruben for a minute, please? 4 times in the scene
- ▶ 59:54 Shawn Wang And you even had the, the name of the one, Feynman.
- ▶ 1:04:26 Ali Taha Just like seeing NVIDIA more and more specialized, like take its GPUs from a general programming paradigm where you're just, it's a general computer that you can use to program threads. 2 times in the scene
- ▶ 1:04:59 Ali Taha They're almost evolving towards an, like, like, like, as in, as in Ruben, I guess, like, compared to Ampere or, you know, T-Four, Ruben is, is, is basically an ASIC.
- ▶ 1:04:59 Ali Taha They're almost evolving towards an, like, like, like, as in, as in Ruben, I guess, like, compared to Ampere or, you know, T-Four, Ruben is, is, is basically an ASIC. 5 times in the scene
- ▶ 1:04:59 Ali Taha They're almost evolving towards an, like, like, like, as in, as in Ruben, I guess, like, compared to Ampere or, you know, T-Four, Ruben is, is, is basically an ASIC.
- ▶ 1:09:51 Philip Kiely Like the whole full-oh, save full-oh movement, like, you don't gotta have a save llama three movement, you just gotta have an eight 100 somewhere.
- ▶ 1:11:37 Philip Kiely You, you, you need, you need GB 300 to fit it on a single node. 3 times in the scene
- ▶ 1:12:35 Alessio Fanelli With the Rubens, you now have, what, NVL-SVII rack of plenty terabytes? 2 times in the scene
- ▶ 1:12:41 Philip Kiely Yeah, now, you, you still have NVL-SVII on, on, uh, Blackwell as well, but, um, you, you can't necessarily assume you're gonna do inference on that. 2 times in the scene
- ▶ 1:30:29 Ali Taha But more and more so we're seeing techniques like NVIDIA released a quantization aware distillation paper,
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
- ▶ 55:31 Eiso Kant Uh, doesn't make sense yet, because we're still training on hoppers, right?
- ▶ 55:35 Eiso Kant We're a 10 K H 200 cluster company right now.
- ▶ 57:35 unnamed speaker I think a GPT-LSS- one-twenty B was the first because it's a large single GPU, which was the H-one hundred, right? 2 times in the scene
- ▶ 1:35:36 unnamed speaker I think the one entity that has more power than the U S government here is NVIDIA. 6 times in the scene
- ▶ 1:36:46 Eiso Kant The, the difference of a model you could train on hoppers versus GB 300 is the difference in a trillion parameter model and a five or six trillion parameter model.
- ▶ 1:36:46 Eiso Kant The, the difference of a model you could train on hoppers versus GB 300 is the difference in a trillion parameter model and a five or six trillion parameter model.
- ▶ 1:41:39 Eiso Kant Uh, it's definitely something that I'm excited to be doing once we move to, uh, to black wall GPUs.
🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Podcast Crossover: AIE, AGI, frontier lab strategy with @matthew_berman and @swyxtv
- ▶ 3:09 Matthew Berman Um, what do you think about Etched, and are they gonna disrupt NVIDIA? 3 times in the scene
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
- ▶ 3:36 unnamed speaker Yeah, just like add a 100.
- ▶ 24:24 unnamed speaker Um, yeah, I mean, you know, I, I think, uh, all very bullish, like, you know, one of my reflections was also, I did not originally, so obviously when I met you guys, you weren't that much in the GPU game, and now you're all about, uh,…
- ▶ 35:49 Akshat Bubna It'll even run like a NVIDIA inside profiler and it'll like tweak configs and it'll arrive the right thing.
🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
- ▶ 1:41:25 Evan Feinberg We published the state-of-the-art numbers in Runs and Poses last year when we published the Pearl technical report with NVIDIA at GTC at the end of last year.
- ▶ 1:43:18 RJ Haneke I don't know if you can ask this, but are you looking at non NVIDIA? 6 times in the scene
⚡️Every product of the future will be a living system — Ronak Malde, Trajectory.ai
- ▶ 11:29 Ronak Malde Harvey and NVIDIA in order to train Nimachan three, super, in order to get to the Pareto frontier.
- ▶ 15:14 Ronak Malde And, and so I think as time goes on, like, Nvidia is investing a ton of money. 2 times in the scene
Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
- ▶ 26:19 Anjney Midha But that doesn't mean you just like throw 500 GB, 305 100,000 GB, 300 that you're like, you know, suboptimal model scaling and you waste a bunch of
- ▶ 31:28 Anjney Midha I actually had him buy for office hours in the class earlier today, and there was an insight he brought up that I hadn't considered before, which is when they decided to pick the standard for their data center, they picked the NVIDIA… 5 times in the scene
- ▶ 38:21 Anjney Midha We carved out like a couple thousand H-one hundreds, but I do think there's extraordinary research being done on university campuses.
🔬 The Limits of AI in Science - Why We Need Self-Driving Labs — Joseph Krause, Radical AI
- ▶ 30:22 unnamed speaker How long do you think it's going to take for that thing to get into an iPhone or, uh, NVIDIA GPU?
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
- ▶ 1:29 Shawn Wang Uh, you were first coming to us or joining the Latent Space World because you were working on Cosmos and NVIDIA, and you did a great paper.
- ▶ 2:51 Ethan He Before XAI, I was working on Cosmos work model at NVIDIA. 2 times in the scene
- ▶ 2:51 Ethan He Before XAI, I was working on Cosmos work model at NVIDIA. 4 times in the scene
- ▶ 5:13 Ethan He One thing I say, like, thanks to my experience at NVIDIA, because first time when we were building Cosmos together, we built it, uh, for about a year.
- ▶ 11:52 Ethan He Draw some examples from Cosmos. 2 times in the scene
- ▶ 27:34 Ethan He If you think about the cost, say, let's say H-one-hundred costs one dollar per hour, and if you use this eight hours a day and 30 days, so, um, every month you're paying this to 40, you're actually not
- ▶ 33:51 unnamed speaker It's a Gemma level model trained on roughly 40 trillion tokens at this many H 200 over this much time, right? 2 times in the scene
- ▶ 37:17 Ethan He So in Cosmos, we did a lot of optimizations to make it not I.O. bound.
- ▶ 37:55 Ethan He And if you, if you look at a number of tokens, uh, we disclose that in cosmos, it's also like tens of trillions of tokens.
- ▶ 40:23 Ethan He In Cosmos, I believe we have, we have like four steps and eight steps. 2 times in the scene
- ▶ 1:15:26 Ethan He In Cosmos, that could be typically the, these models, they, they have two parts. 4 times in the scene
- ▶ 1:41:26 Ethan He And after that is NVIDIA, Cosmos. 3 times in the scene
⚡️ Google's Open AI Strategy — Omar Sanseviero, Google DeepMind
- ▶ 3:42 Omar Sanseviero So for example, we work with Lama CPP, Olama, MLX, Hogan Faces, BLM, NVIDIA, AMD.
AI Agents Need Computers: 74% MoM Growth, 850K/Day Runs, & New Agent Cloud — Ivan Burazin, Daytona
- ▶ 1:09:07 Shawn Wang It's, this is like, it's whatever Nvidia decides to bless that day. 2 times in the scene
The $15B Physical AI Company: Simulation, Autonomy OS, Neural Sim, & 1K Engineers—Applied Intuition
- ▶ 6:49 Qasar Younis Nvidia does or what an AMD, but we just don't do chips.
- ▶ 21:23 Peter Ludwig and you have a start with an assumption that, oh, I'm gonna, I'm gonna use CUDA and I'm gonna run this, uh, on an NVIDIA chip, then you don't really have to think about the hardware in that sense. 2 times in the scene
AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
- ▶ 44:52 Shawn Wang Yeah, I understand that you, uh, you published a lot of technical detail at GTC, so I was just going to bring it up a little bit. 2 times in the scene
- ▶ 46:42 Mikhail Parakhin CentML recently got acquired by NVIDIA. 2 times in the scene
- ▶ 1:02:50 Mikhail Parakhin And even Liquid, we had to work a lot with NVIDIA and to, because almost everything is not designed in CUDA for, or in, in the current stack for, for low latency.
- ▶ 1:11:12 Mikhail Parakhin Um, Microsoft, uh, and the NVIDIA collaboration model.
Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
- ▶ 3:42 Marc Andreessen Although that's really when the, the, the NVIDIA phenomenon really, it was, it was, I would say it was in that period when it was very clear that at the time the vocabulary was more machine learning, but it was very clear at that time that…
- ▶ 20:14 Marc Andreessen It was like a new venture, but like the money that's being deployed now at scale is Microsoft and, you know, and Amazon and Google and Facebook and NVIDIA and, you know, these, these, these, and now, you know, by the way, OpenAI and… 4 times in the scene
- ▶ 24:23 Shawn Wang One of my early hits was, like, modeling the lifespan of the H-one hundred and H-two hundreds, and, and going, like, you know, usually they advise, like, four to seven years, and it was,
- ▶ 24:23 Shawn Wang One of my early hits was, like, modeling the lifespan of the H-one hundred and H-two hundreds, and, and going, like, you know, usually they advise, like, four to seven years, and it was,
- ▶ 32:00 Shawn Wang NVIDIA is doing a lot. 3 times in the scene
Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
- ▶ 2:28 Fan-yun Sun With actually NVIDIA research during my PhD years on essentially generating interactive worlds to train reinforcement learning agents or embodied AI agents. 3 times in the scene
Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
- ▶ 21:39 Guillaume Lample We actually, uh, provide them some, some missed help purchase, basically what we announced at, uh, GTC this week.
Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
- ▶ 1:20:09 Felix Rieseberg And for a while I really put GPU developers on this pedestal in my head.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 0:47 Shawn Wang Uh, and our friends, uh, Netter and, uh, Kyle from NVIDIA. 7 times in the scene
- ▶ 1:13 Shawn Wang And we're, we're sort of recording this ahead of NVIDIA GTC, which is coming to town, uh, again, uh, or taking over town, uh, which, uh, which we'll all be at. 4 times in the scene
- ▶ 4:00 Nader Khalil And, um, whenever we would talk to users, they wanted a GPU, they wanted an A-one hundred. 4 times in the scene
- ▶ 9:33 Shawn Wang I think there's also like, you almost like we're the right team at the right time when NVIDIA is starting to invest a lot more in developer experience or whatever you call it, uh, UX or I don't know what you call it, like software. 7 times in the scene
- ▶ 11:49 Nader Khalil Even when Grace Blackwell or when, um, uh, DGX Spark was first coming out, uh, getting to be involved in that from the beginning of the developer experience.
- ▶ 13:04 Nader Khalil So there's a tool called NVIDIA sync.
- ▶ 20:05 Kyle Kranen I took a different path to NVIDIA than Nader. 18 times in the scene
- ▶ 27:18 Shawn Wang And let's go right into Dynamo. 6 times in the scene
- ▶ 28:05 Kyle Kranen Dynoa sort of came about at NVIDIA because myself and a couple others were sort of talking about these concepts that like, you know, you have inference engines like VLM, SGLang, TensorRTLM, um, 2 times in the scene
- ▶ 30:26 Kyle Kranen Let's say you're on an H 100, uh, the maximum NVLink domain, domain for H 100 for most DGX H 100 2 times in the scene
- ▶ 42:04 Kyle Kranen And like with our future generations, generations of hardware, we actually announced like with Rubin, this new accelerator that is pre-fill specific.
- ▶ 54:48 Kyle Kranen I use the associate of NVIDIA 11 times in the scene
- ▶ 1:00:46 Shawn Wang You have Nemo-chan and, and we have an internal cluster.
page 1 of 4 · 100 scenes per page · newest episode first next →