Nvidia, every mention
79 scenes (2024), the whole family · ← back to Nvidia
tap a year for its mentions
every year 2024 anyone Shawn Wang 80Dylan Patel 55Kyle Kranen 43Ali Taha 26Ethan He 24Chris Lattner 24George Hotz 23Philip Kiely 20Doug O'Laughlin 19Sarah Chieng 17
Verbatim, from the transcripts: the passages where Nvidia comes up
Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
- ▶ 8:45 Loubna Ben Allal Um, and I, uh, this is a recent paper from NVIDIA, Mnemotron CC.
- ▶ 22:05 Loubna Ben Allal Um, there's also this paper by NVIDIA that was released recently.
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
- ▶ 17:06 Dan Fu Um, there's this, uh, NVIDIA and MIT put out this new
- ▶ 24:19 Eugene Cheah Okay, because we are all GPU poor, um, and to be clear, like, most of this research is done, like, only on a handful H 100, which I had one Google researcher told me that was like his experiment budget for a single researcher.
- ▶ 29:57 Dan Fu So for example, on H-one-hundred, um, everything is, really revolves around a warp group, uh, matrix multiply operation.
Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
- ▶ 29:16 unnamed speaker Uh, we also collaborated with NVIDIA and open sourced another model, Nemo-II-B, uh, another great model.
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 2:38 Isaac Robinson Reminds me of those RTX demonstrations for next generation video games such as cyberpunk, but with
[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
- ▶ 3:35 Sarah Chieng With an NVIDIA GPU or like running inference on an NVIDIA GPU.
- ▶ 4:21 unnamed speaker Uh, for, for some DB scale, if you want to do, uh, tensor parallel, throw four H 100 at it, then you can, 3 times in the scene
- ▶ 7:39 Sarah Chieng And so, you know, I'm sure everyone's wondering, everyone's much more familiar with the NVIDIA H-Hundred. 6 times in the scene
- ▶ 7:39 Sarah Chieng And so, you know, I'm sure everyone's wondering, everyone's much more familiar with the NVIDIA H-Hundred. 9 times in the scene
- ▶ 25:27 Sarah Chieng And so, as I mentioned, this was a little bit covered in kind of my introduction in the very beginning, but going back, you know, we have our AI processor, the wafer scale engine, much larger than anything we've seen with the NVIDIA GPU,…
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 41:15 unnamed speaker Yeah, 20 dollars, they did, like, eight A-One hundreds for an hour, and they're able to match the glue store of basic BERT with their recipe.
[Paper Club] Upcycling Large Language Models into Mixture of Experts
- ▶ 0:11 Ethan He I'm Ethan from NVIDIA.
- ▶ 8:07 Ethan He Megatron Online provides simple bearable training loop, and you can, you can easily hack, and Nemo provide a high-level interface, where you can, you can just provide Pythonic configuration to train these models.
- ▶ 20:49 Ethan He Uh, for five B and from NVIDIA, we have NemoTron three four DB.
Singapore: the AI Engineer Nation — with Minister Josephine Teo
- ▶ 42:32 Alessio Fanelli 15% of NVIDIA's revenue in Q three of 2024.
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- ▶ 46:56 Jesse Hu They give you, or they have you run, 36 vCPUs with a NVIDIA A-TEN GPU.
- ▶ 46:56 Jesse Hu They give you, or they have you run, 36 vCPUs with a NVIDIA A-TEN GPU.
- ▶ 49:19 Jesse Hu So the, does O-one preview run on A-ten? 2 times in the scene
- ▶ 59:41 unnamed speaker So, for next week, right, for next week, we will be going, uh, we have, let me see, Ethan from NVIDIA who's going through the Megatron, uh, NemoTron distillation or Amoe, um, I'm not so sure which one, um, and yep, look forward to it, and… 2 times in the scene
Building the Silicon Brain - Drew Houston of Dropbox
- ▶ 30:35 Drew Houston And, you know, certainly at the bottom with NVIDIA and the semiconductor companies, and then, then it's going to be at the top, like the people who, who have the customer relationship, who have the application layer.
- ▶ 52:59 Drew Houston And then there's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that. 2 times in the scene
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 1:04:23 unnamed speaker If you run it on the right hardware, if you run it on a T four or like anything that's bigger for any modern GPU, basically it's going to be almost real time.
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
- ▶ 18:02 Andrej Karpathy Which was state-of-the-art LLM as of 2019 or so, uh, and you can train it on a single node of H-one-hundreds in about 24 hours, and that costs roughly 600 dollars.
[Paper Club] Writing in the Margins: Chunked Prefill KV Caching for Long Context Retrieval
- ▶ 41:15 Umar Jamil Uh, there, it was, uh, in video, uh, explanation of checks for feeling. 2 times in the scene
Answer.ai & AI Magic with Jeremy Howard
- ▶ 24:17 Jeremy Howard Not everybody has to use NVIDIA.
- ▶ 37:24 unnamed speaker Um, and then you were like, well, what if you could fine tune a 70 B model on like a 40 90?
Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
- ▶ 52:54 Joseph Nelson The smallest model is thirty-eight million parameters and can run at 45 FPS on an A-one hundred, right?
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 1:34 Alessio Fanelli We had, uh, a lot of- NVIDIA, Goldman, Tomasic, uh, SYNCTEL.
- ▶ 20:34 unnamed speaker Something I mentioned there, and it's something that always comes up, even in the Sovereign AI Summit that we did, was, what does NVIDIA's competitors have, have any threat to NVIDIA? 6 times in the scene
- ▶ 21:00 unnamed speaker only was for H-One hundreds.
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 10:43 unnamed speaker So, 16,000 each 100 hours, what failures they hit, why they went for simplicity.
- ▶ 14:22 unnamed speaker Um, we have more intro providers in video stuff, but
- ▶ 19:01 unnamed speaker I also saw lots of papers on this from, um, maybe NVIDIA and Matei Zaharia from Berkeley or Stanford.
- ▶ 47:38 Eugene Cheah It's going to take slower than normal because I spoke to some people in the fine-tuning community, um, the biggest hurdle has been, what do you mean you need at least three nodes of H 100 to start the process?
- ▶ 58:44 unnamed speaker It takes eight H 100 per instance. 2 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
- ▶ 47:00 Yi Tay So H-Hundred, there were major delays, right, so we were sitting around, right, bunched with, like, And to be clear, you don't own the compute, you are renting. 3 times in the scene
- ▶ 47:12 Yi Tay For, for a long period of time, we had, 500 A-One hundreds, because we, we, we, we, we made a commitment, like, uh, and they were constantly being delayed, I think, because of H-One hundred, supply demand, whatever, like, like, reasons…
State of the Art: Training 70B LLMs on 10,000 H100 clusters
- ▶ 12:20 Jonathan Frankle It's not like NVIDIA did a bad job or, you know, Mellanox did a bad job or the, like the server builder or the data center operator or the cloud provider, like the million other parties that are involved in building this.
- ▶ 13:27 unnamed speaker Um, it's, this post is about one cluster that has 4092 H 100 GPUs spread across 511 computers. 2 times in the scene
- ▶ 19:36 Josh Albrecht Uh, so we worked pretty closely with video on the, yeah, that's what I'm saying.
- ▶ 24:03 Josh Albrecht I think it made it a lot easier from our perspective to have direct control over this, instead of having to go to the cloud provider that goes to the data center that goes to the supplier, we could just go direct to Nvidia or Dell or the…
- ▶ 26:48 Jonathan Frankle I think the, I I'm just curious to ask, like, you know, suppose you were to set up another, let's say another H 100 cluster
- ▶ 27:54 Jonathan Frankle With Connect X eight that will have its own fun behavior and all that good stuff.
- ▶ 28:56 Josh Albrecht Say thanks very much to Dell and H five and NVIDIA and the other people that have done a lot of the work, like to bring up this cluster, uh, you know, with 4000 GPUs and three tier networking, networking architecture, you have 12,000… 2 times in the scene
- ▶ 32:53 Josh Albrecht So we use things from NVIDIA's, you know, Megatron stuff.
- ▶ 42:14 Jonathan Frankle And it's kind of neat to think about that, you know, as, as one thing that I think NVIDIA announced for
How AI is Eating Finance - with Mike Conover of Brightwave
- ▶ 9:45 Mike Conover So, um, you can ask Brightwave, um, a question like, how is Nvidia's position in the GPU market impacted by rare earth metal shortages?
- ▶ 9:57 Mike Conover The Matic contributors to an investment decision, um, or, or developing your thesis that, um, in response to export controls on a 100 cards, uh, China has put in place licensors on the transfer of germanium and gallium, which are not rare…
How to train a Million Context LLM — with Mark Huang of Gradient.ai
- ▶ 13:47 Mark Huang Uh, cloud providers, and, um, they were offering up, like, we want to do a collaboration to showcase, uh, their technology, and, uh, it, it just made it really easy for us to, like, scale up with their L-Forties, and those are the…
Breaking down the OG GPT Paper by Alec Radford
Supervise the Process of AI Research — with Jungwon Byun and Andreas Stuhlmüller of Elicit
- ▶ 52:39 unnamed speaker I don't know if that's like the best name to, to think about it, but I was using one of these chat which are documents features, and I put the AMD and my 300 specs on the, um, new, you know, Blackwell chips from NVIDIA, and I was asking… 2 times in the scene
Why Google failed to make GPT-3 -- with David Luan of Adept
- ▶ 10:57 Shawn Wang I think it's, um, underrated how much NVIDIA worked with you in the early days as well. 6 times in the scene
- ▶ 11:01 Shawn Wang I think, um, maybe, I think it was Jensen, I'm not sure who circulated, um, um, a recent photo of him delivering the first, uh, DGX to you, to you guys.
- ▶ 12:11 David Luan One interesting set of stuff is just like, you know, like knowing that a 100 generation that like quad sparsity was going to be a thing.
- ▶ 19:15 Alessio Fanelli Uh, I'm actually giving a talk at NVIDIA GTC about this, but basically software as a service, you're wrapping user productivity in software, um, with agents and services as software is, uh, replacing things that, you know, you would ask…
- ▶ 19:15 Alessio Fanelli Uh, I'm actually giving a talk at NVIDIA GTC about this, but basically software as a service, you're wrapping user productivity in software, um, with agents and services as software is, uh, replacing things that, you know, you would ask…
Making Transformers Sing - with Mikey Shulman of Suno
- ▶ 14:28 Mikey Shulman Uh, so that's a, that's a really awesome collaboration, um, with, with our friends at NVIDIA.
Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
- ▶ 17:55 Soumith Chintala Distributed, uh, NVIDIA and AMD GPUs, like, just like having a generalization of the concept of a backend, how they treat compilation, uh, 3 times in the scene
- ▶ 46:07 Soumith Chintala Uh, that is by the end of this year, and 600 K H-One hundred equivalents. 2 times in the scene
- ▶ 57:49 Soumith Chintala Is they, each large company has a sufficient enough set of verticalized workloads, uh, that have a pattern to them that, say, a more generic accelerator like an NVIDIA or an AMD GPU does not exploit.
A Brief History of the Open Source AI Hacker - with Ben Firshman of Replicate
- ▶ 53:51 Ben Firshman Like COG is really designed around servers and attaching to CUDA devices and, and NVIDIA GPUs and this kind of thing.
- ▶ 58:22 unnamed speaker None, none from NVIDIA, which is your newest investor? 2 times in the scene
Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
- ▶ 16:30 Erik Bernhardsson Yeah, it's like, you just, like, say, you know, on the function decorator, you're like, GPU equals, you know, a 100, and then, or like, GPU equals, you know, uh, a 10 or a T four, something like that, and then you get that GPU, and like,…
- ▶ 16:30 Erik Bernhardsson Yeah, it's like, you just, like, say, you know, on the function decorator, you're like, GPU equals, you know, a 100, and then, or like, GPU equals, you know, uh, a 10 or a T four, something like that, and then you get that GPU, and like,…
- ▶ 16:30 Erik Bernhardsson Yeah, it's like, you just, like, say, you know, on the function decorator, you're like, GPU equals, you know, a 100, and then, or like, GPU equals, you know, uh, a 10 or a T four, something like that, and then you get that GPU, and like,…
Building an open AI company - with Ce and Vipul of Together AI
- ▶ 27:40 Vipul Ved Prakash They're mostly A-hundreds and H-hundreds.
- ▶ 27:40 Vipul Ved Prakash They're mostly A-hundreds and H-hundreds. 2 times in the scene
- ▶ 29:37 Vipul Ved Prakash NVIDIA is obviously motivated to, to help us, uh, both, both as an investor and we are their customers. 2 times in the scene
- ▶ 32:20 Alessio Fanelli And this post he said, our model indicates that together it's better off using two a 180 gig system rather than a each 100 based system. 2 times in the scene
- ▶ 32:20 Alessio Fanelli And this post he said, our model indicates that together it's better off using two a 180 gig system rather than a each 100 based system. 2 times in the scene
- ▶ 34:04 Vipul Ved Prakash Tricks and techniques, um, and we think there's a lot of room for optimization here, so, um, you know, whichever hardware provides better performance, whether it's H-Hundred or A-Hundreds or L-Forties, uh, we can sort of measure price…
- ▶ 42:07 Vipul Ved Prakash It's, uh, um, you know, the value we get from doing specific optimization, even, even for, you know, what works well for a particular model on A hundreds with a particular bus.
- ▶ 42:22 Vipul Ved Prakash Uh, versus Edge hundreds.
The Four Wars of the AI Stack - Dec 2023 Recap
- ▶ 32:24 unnamed speaker Like, uh, Fireworks, um, recently announced, uh, Fire Attention, where they wrote a custom cruder kernel for mixed drawl, uh, on H 100.
The AI-First Graphics Editor - with Suhail Doshi of Playground AI
- ▶ 21:33 Suhail Doshi And we did that for academic research because there's a whole bunch of, you know, we come across people all the time in academia and they have like, they have access to like one a 100 or eight at best.
- ▶ 1:02:19 unnamed speaker Uh, you, I, I, you had a tweet about like how many A 100 you have, but I feel like it's out of date probably.