nvidia-h100, every mention

54 scenes · ← back to nvidia-h100

tap a year for its mentions
0020840152023202420252026episodesmentions
08152023202420252026episodes it came up in
0047.58152023202420252026episodesmentions per episode

every year anyone Dylan Patel 17Sarah Chieng 9Chris Lattner 6Alessio Fanelli 5Philip Kiely 4Evan Conrad 4Yining Zhang 3Yi Tay 3Vipul Ved Prakash 3Varun Mohan 3

Verbatim, from the transcripts: the passages where nvidia-h100 comes up

loading…

Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten Aug 3, 2026 · 5 mentions

  • ▶ 38:45 Ali Taha Like, if you, if you obviously have a thing where you're serving it on just, like, a node of H-one-hundreds, and then you throw, like, you know, you short the model across, like, four nodes of B-to-hundreds.
  • ▶ 54:39 Philip Kiely Like, let's say, let's say you're doing a deployment on H-one hundreds for whatever reason, and you're putting a, a trillion parameter model on there. 4 times in the scene

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI Jul 22, 2026 · 2 mentions

  • ▶ 57:35 unnamed speaker I think a GPT-LSS- one-twenty B was the first because it's a large single GPU, which was the H-one hundred, right? 2 times in the scene

Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP Jun 18, 2026 · 1 mention

  • ▶ 38:21 Anjney Midha We carved out like a couple thousand H-one hundreds, but I do think there's extraordinary research being done on university campuses.

Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He Jun 1, 2026 · 1 mention

  • ▶ 27:34 Ethan He If you think about the cost, say, let's say H-one-hundred costs one dollar per hour, and if you use this eight hours a day and 30 days, so, um, every month you're paying this to 40, you're actually not

Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different" Apr 3, 2026 · 1 mention

  • ▶ 24:23 Shawn Wang One of my early hits was, like, modeling the lifespan of the H-one hundred and H-two hundreds, and, and going, like, you know, usually they advise, like, four to seven years, and it was,

Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup Mar 8, 2026 · 2 mentions

  • ▶ 30:26 Kyle Kranen Let's say you're on an H 100, uh, the maximum NVLink domain, domain for H 100 for most DGX H 100 2 times in the scene

Dylan Patel Explains the AI War While Cooking | In-Context Cooking Feb 26, 2026 · 1 mention

  • ▶ 42:46 Dylan Patel It was a large GPU, um, and it was like having, it was like the best memory, the best networking, everything, sort of the best as possible, um, and sort of like one size fits all, uh, with the main line of like A-one hundred, H-one…

Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis Feb 24, 2026 · 1 mention

Goodfire AI’s Bet: Interpretability as the Next Frontier of Model Design — Myra Deng & Mark Bissell Feb 5, 2026 · 2 mentions

[NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton Dec 31, 2025 · 1 mention

  • ▶ 24:19 Kevin Wang The nice thing is that all of our experiments, even the thousand layer networks, can be run on one single, 80 gigabyte, each 100 GPU.

SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow) Dec 18, 2025 · 3 mentions

  • ▶ 10:07 unnamed speaker So, so I'm reading in the paper, it's, uh, 10 objects on two HCOs, 28 on four HCOs, and 64 on eight HCOs, something like that. 3 times in the scene

How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony Nov 3, 2025 · 2 mentions

  • ▶ 3:19 Quentin Anthony Um, we found that it's, it's great, uh, for flash attention to specifically, we were able to be H-one hundred. 2 times in the scene

Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21 Oct 11, 2025 · 1 mention

  • ▶ 5:27 Barak Lenz So we designed J to have a version that fits on a single GPU, a single AY 100 or H one, 80 gigabytes.

A Technical History of Generative Media Sep 8, 2025 · 3 mentions

  • ▶ 23:15 Batuhan Taskaya And in, in this world, like, we had to build our orchestration layer, we had to build our own distributed file system, we had to build our own container runtimes, all, all the stack to make sure that the cold starts are extremely,…
  • ▶ 23:57 unnamed speaker You keep mentioning H-one hundreds. 2 times in the scene

🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R) Jul 28, 2025 · 3 mentions

  • ▶ 45:36 Varun Mohan It's like for the H 100 boxes, you shove eight of these H 100 on a machine between two nodes. 3 times in the scene

The Shape of Compute (Chris Lattner of Modular) Jun 13, 2025 · 6 mentions

SF Compute: Commoditizing Compute Apr 11, 2025 · 5 mentions

  • ▶ 24:58 Evan Conrad Like you can go on SF compute today and you can get thousands of H 100 for an hour if you want.
  • ▶ 27:59 Michael Swix (Swyx) One of our top pieces from last year was talking about the H-one hundred glut from all the, uh, long-term contracts that were not being fully utilized and being put under the market.
  • ▶ 31:53 Evan Conrad Um, lots of bio and pharma, um, was using, um, H 100 training sort of the bio models of sorts. 3 times in the scene

Outlasting Noam Shazeer, Crowdsourcing Chai AI w/ 1.4m DAU — with William Beauchamp, Chai Research Jan 26, 2025 · 1 mention

  • ▶ 56:30 William Beauchamp What the money let them do was, if they wanted to fine-tune Alarma-seventyb on eight H-one-hundreds overnight, if you give them money, then they can do it.

DeepSeek V3, SGLang, and the state of Open Model Inference in 2025 (Quantization, MoEs, Pricing) Jan 19, 2025 · 4 mentions

  • ▶ 2:46 Yining Zhang You need, I think, uh, 671 gigabytes for the weights, and you, you also need an extra memory for the KV cache, so it's not possible to run that on H-one hundred.
  • ▶ 5:06 Yining Zhang So big model that we should use H 200 or use H 100 multi nodes. 2 times in the scene
  • ▶ 30:44 unnamed speaker In other words, a single model might want to horizontally scale up to 200 replicas, each of which is, let's say, two H-one hundreds or four H-one hundreds or even a full node.

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 1 mention

2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024] Dec 24, 2024 · 2 mentions

  • ▶ 24:19 Eugene Cheah Okay, because we are all GPU poor, um, and to be clear, like, most of this research is done, like, only on a handful H 100, which I had one Google researcher told me that was like his experiment budget for a single researcher.
  • ▶ 29:57 Dan Fu So for example, on H-one-hundred, um, everything is, really revolves around a warp group, uh, matrix multiply operation.

[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras) Dec 7, 2024 · 12 mentions

  • ▶ 4:21 unnamed speaker Uh, for, for some DB scale, if you want to do, uh, tensor parallel, throw four H 100 at it, then you can, 3 times in the scene
  • ▶ 7:39 Sarah Chieng And so, you know, I'm sure everyone's wondering, everyone's much more familiar with the NVIDIA H-Hundred. 9 times in the scene

llm.c's Origin and the Future of LLM Compilers - Andrej Karpathy at CUDA MODE Sep 21, 2024 · 1 mention

  • ▶ 18:02 Andrej Karpathy Which was state-of-the-art LLM as of 2019 or so, uh, and you can train it on a single node of H-one-hundreds in about 24 hours, and that costs roughly 600 dollars.

The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap) Aug 2, 2024 · 1 mention

  • ▶ 21:00 unnamed speaker only was for H-One hundreds.

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 4 mentions

  • ▶ 10:43 unnamed speaker So, 16,000 each 100 hours, what failures they hit, why they went for simplicity.
  • ▶ 47:38 Eugene Cheah It's going to take slower than normal because I spoke to some people in the fine-tuning community, um, the biggest hurdle has been, what do you mean you need at least three nodes of H 100 to start the process?
  • ▶ 58:44 unnamed speaker It takes eight H 100 per instance. 2 times in the scene

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka Jul 5, 2024 · 3 mentions

  • ▶ 47:00 Yi Tay So H-Hundred, there were major delays, right, so we were sitting around, right, bunched with, like, And to be clear, you don't own the compute, you are renting. 3 times in the scene

State of the Art: Training 70B LLMs on 10,000 H100 clusters Jun 25, 2024 · 3 mentions

  • ▶ 13:27 unnamed speaker Um, it's, this post is about one cluster that has 4092 H 100 GPUs spread across 511 computers. 2 times in the scene
  • ▶ 26:48 Jonathan Frankle I think the, I I'm just curious to ask, like, you know, suppose you were to set up another, let's say another H 100 cluster

Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI Mar 6, 2024 · 2 mentions

Building an open AI company - with Ce and Vipul of Together AI Feb 8, 2024 · 5 mentions

The Four Wars of the AI Stack - Dec 2023 Recap Jan 26, 2024 · 1 mention

  • ▶ 32:24 unnamed speaker Like, uh, Fireworks, um, recently announced, uh, Fire Attention, where they wrote a custom cruder kernel for mixed drawl, uh, on H 100.

The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis Dec 5, 2023 · 19 mentions

  • ▶ 7:03 Dylan Patel We have 512 H-Hundreds coming online in, in August, and it's like, oh, cool, like, but then you're like, you know, going through the supply chain, it's like, dude, you realize there's 400 to 500,000 being, 400,000 manufactured last… 3 times in the scene
  • ▶ 15:26 Dylan Patel The H-One-Hundred has 3.35 terabytes a second of memory bandwidth, and it has a thousand teraflops of FP-sixteen, B-foot-sixteen. 5 times in the scene
  • ▶ 27:34 Dylan Patel That's, they're very, they, while they do have quite a few GPUs, they made a big announcement about having 4000 H 100, that's still relatively poor, right, when we're talking about hundreds of thousands of like the big labs, uh, like…
  • ▶ 30:32 Alessio Fanelli It's like, you know, if you buy an H-one hundred, sure, the next series is gonna be better, but, like, at least the hardware is good. 3 times in the scene
  • ▶ 41:08 Dylan Patel They're gonna release a better chip than the H 100, ah, within the next quarter or so, right? 6 times in the scene
  • ▶ 48:55 Dylan Patel Um, it's worse performance than the H-H-E-N-H-E-D, uh, but the cost effectiveness of it is, is better for Microsoft internally, just because they don't have to pay the Nvidia tax.

Beating GPT-4 with Open Source Models - with Michael Royzen of Phind Nov 3, 2023 · 1 mention

  • ▶ 1:06:44 Michael Royzen Um, and, like, NVIDIA claims that this strategy, um, that they're kind of demoing with the H-one-hundred has no degradation.

Ep 18: Petaflops to the People — with George Hotz of tinycorp Jun 20, 2023 · 2 mentions

Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.