nvidia-h100, every mention
54 scenes · ← back to nvidia-h100
tap a year for its mentions
every year anyone Dylan Patel 17Sarah Chieng 9Chris Lattner 6Alessio Fanelli 5Philip Kiely 4Evan Conrad 4Yining Zhang 3Yi Tay 3Vipul Ved Prakash 3Varun Mohan 3
Verbatim, from the transcripts: the passages where nvidia-h100 comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 38:45 Ali Taha Like, if you, if you obviously have a thing where you're serving it on just, like, a node of H-one-hundreds, and then you throw, like, you know, you short the model across, like, four nodes of B-to-hundreds.
- ▶ 54:39 Philip Kiely Like, let's say, let's say you're doing a deployment on H-one hundreds for whatever reason, and you're putting a, a trillion parameter model on there. 4 times in the scene
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
- ▶ 57:35 unnamed speaker I think a GPT-LSS- one-twenty B was the first because it's a large single GPU, which was the H-one hundred, right? 2 times in the scene
Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
- ▶ 38:21 Anjney Midha We carved out like a couple thousand H-one hundreds, but I do think there's extraordinary research being done on university campuses.
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
- ▶ 24:23 Shawn Wang One of my early hits was, like, modeling the lifespan of the H-one hundred and H-two hundreds, and, and going, like, you know, usually they advise, like, four to seven years, and it was,
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 30:26 Kyle Kranen Let's say you're on an H 100, uh, the maximum NVLink domain, domain for H 100 for most DGX H 100 2 times in the scene
Dylan Patel Explains the AI War While Cooking | In-Context Cooking
- ▶ 42:46 Dylan Patel It was a large GPU, um, and it was like having, it was like the best memory, the best networking, everything, sort of the best as possible, um, and sort of like one size fits all, uh, with the main line of like A-one hundred, H-one…
Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
- ▶ 35:31 Doug O'Laughlin To just run the swarm, I think it's like a 16 node of H-one hundreds.
Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell
- ▶ 23:41 Mark Bissell It takes a full, like, H-one hundred node. 2 times in the scene
[NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
- ▶ 24:19 Kevin Wang The nice thing is that all of our experiments, even the thousand layer networks, can be run on one single, 80 gigabyte, each 100 GPU.
SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
- ▶ 10:07 unnamed speaker So, so I'm reading in the paper, it's, uh, 10 objects on two HCOs, 28 on four HCOs, and 64 on eight HCOs, something like that. 3 times in the scene
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 3:19 Quentin Anthony Um, we found that it's, it's great, uh, for flash attention to specifically, we were able to be H-one hundred. 2 times in the scene
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 5:27 Barak Lenz So we designed J to have a version that fits on a single GPU, a single AY 100 or H one, 80 gigabytes.
A Technical History of Generative Media
- ▶ 23:15 Batuhan Taskaya And in, in this world, like, we had to build our orchestration layer, we had to build our own distributed file system, we had to build our own container runtimes, all, all the stack to make sure that the cold starts are extremely,…
- ▶ 23:57 unnamed speaker You keep mentioning H-one hundreds. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 45:36 Varun Mohan It's like for the H 100 boxes, you shove eight of these H 100 on a machine between two nodes. 3 times in the scene
The Shape of Compute (Chris Lattner of Modular)
- ▶ 5:03 Chris Lattner Let's, oh yeah, let's add H 100 support. 2 times in the scene
- ▶ 14:28 Chris Lattner it turns out that, uh, an H-one hundred and AMD chip are actually quite different.
- ▶ 30:51 Chris Lattner And so you can go look at how we brought up H-one hundred, built flash attention from scratch in a few weeks, built like all the stuff. 2 times in the scene
- ▶ 56:32 Chris Lattner Go look at VLM.
SF Compute: Commoditizing Compute
- ▶ 24:58 Evan Conrad Like you can go on SF compute today and you can get thousands of H 100 for an hour if you want.
- ▶ 27:59 Michael Swix (Swyx) One of our top pieces from last year was talking about the H-one hundred glut from all the, uh, long-term contracts that were not being fully utilized and being put under the market.
- ▶ 31:53 Evan Conrad Um, lots of bio and pharma, um, was using, um, H 100 training sort of the bio models of sorts. 3 times in the scene
Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research
- ▶ 56:30 William Beauchamp What the money let them do was, if they wanted to fine-tune Alarma-seventyb on eight H-one-hundreds overnight, if you give them money, then they can do it.
DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing)
- ▶ 2:46 Yining Zhang You need, I think, uh, 671 gigabytes for the weights, and you, you also need an extra memory for the KV cache, so it's not possible to run that on H-one hundred.
- ▶ 5:06 Yining Zhang So big model that we should use H 200 or use H 100 multi nodes. 2 times in the scene
- ▶ 30:44 unnamed speaker In other words, a single model might want to horizontally scale up to 200 replicas, each of which is, let's say, two H-one hundreds or four H-one hundreds or even a full node.
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 14:34 Shawn Wang Um, it doesn't even fit on like one node of, uh, of H 100.
2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
- ▶ 24:19 Eugene Cheah Okay, because we are all GPU poor, um, and to be clear, like, most of this research is done, like, only on a handful H 100, which I had one Google researcher told me that was like his experiment budget for a single researcher.
- ▶ 29:57 Dan Fu So for example, on H-one-hundred, um, everything is, really revolves around a warp group, uh, matrix multiply operation.
[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
- ▶ 4:21 unnamed speaker Uh, for, for some DB scale, if you want to do, uh, tensor parallel, throw four H 100 at it, then you can, 3 times in the scene
- ▶ 7:39 Sarah Chieng And so, you know, I'm sure everyone's wondering, everyone's much more familiar with the NVIDIA H-Hundred. 9 times in the scene
llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE
- ▶ 18:02 Andrej Karpathy Which was state-of-the-art LLM as of 2019 or so, uh, and you can train it on a single node of H-one-hundreds in about 24 hours, and that costs roughly 600 dollars.
The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)
- ▶ 21:00 unnamed speaker only was for H-One hundreds.
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 10:43 unnamed speaker So, 16,000 each 100 hours, what failures they hit, why they went for simplicity.
- ▶ 47:38 Eugene Cheah It's going to take slower than normal because I spoke to some people in the fine-tuning community, um, the biggest hurdle has been, what do you mean you need at least three nodes of H 100 to start the process?
- ▶ 58:44 unnamed speaker It takes eight H 100 per instance. 2 times in the scene
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
State of the Art: Training 70B LLMs on 10,000 H100 clusters
- ▶ 13:27 unnamed speaker Um, it's, this post is about one cluster that has 4092 H 100 GPUs spread across 511 computers. 2 times in the scene
- ▶ 26:48 Jonathan Frankle I think the, I I'm just curious to ask, like, you know, suppose you were to set up another, let's say another H 100 cluster
Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
- ▶ 46:07 Soumith Chintala Uh, that is by the end of this year, and 600 K H-One hundred equivalents. 2 times in the scene
Building an open AI company - with Ce and Vipul of Together AI
- ▶ 27:40 Vipul Ved Prakash They're mostly A-hundreds and H-hundreds. 2 times in the scene
- ▶ 32:20 Alessio Fanelli And this post he said, our model indicates that together it's better off using two a 180 gig system rather than a each 100 based system. 2 times in the scene
- ▶ 42:22 Vipul Ved Prakash Uh, versus Edge hundreds.
The Four Wars of the AI Stack - Dec 2023 Recap
- ▶ 32:24 unnamed speaker Like, uh, Fireworks, um, recently announced, uh, Fire Attention, where they wrote a custom cruder kernel for mixed drawl, uh, on H 100.
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 7:03 Dylan Patel We have 512 H-Hundreds coming online in, in August, and it's like, oh, cool, like, but then you're like, you know, going through the supply chain, it's like, dude, you realize there's 400 to 500,000 being, 400,000 manufactured last… 3 times in the scene
- ▶ 15:26 Dylan Patel The H-One-Hundred has 3.35 terabytes a second of memory bandwidth, and it has a thousand teraflops of FP-sixteen, B-foot-sixteen. 5 times in the scene
- ▶ 27:34 Dylan Patel That's, they're very, they, while they do have quite a few GPUs, they made a big announcement about having 4000 H 100, that's still relatively poor, right, when we're talking about hundreds of thousands of like the big labs, uh, like…
- ▶ 30:32 Alessio Fanelli It's like, you know, if you buy an H-one hundred, sure, the next series is gonna be better, but, like, at least the hardware is good. 3 times in the scene
- ▶ 41:08 Dylan Patel They're gonna release a better chip than the H 100, ah, within the next quarter or so, right? 6 times in the scene
- ▶ 48:55 Dylan Patel Um, it's worse performance than the H-H-E-N-H-E-D, uh, but the cost effectiveness of it is, is better for Microsoft internally, just because they don't have to pay the Nvidia tax.
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 1:06:44 Michael Royzen Um, and, like, NVIDIA claims that this strategy, um, that they're kind of demoing with the H-one-hundred has no degradation.
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 48:36 George Hotz Uh, it's, uh, it's 400 grand. 2 times in the scene