Nvidia A100, every mention
28 scenes · ← back to Nvidia A100
tap a year for its mentions
every year anyone Dylan Patel 7Nader Khalil 4Chris Lattner 3Vipul Ved Prakash 2George Hotz 2Batuhan Taskaya 2Alessio Fanelli 2Yi Tay 1Tri Dao 1Thomas Sohmers 1
Verbatim, from the transcripts: the passages where Nvidia A100 comes up
Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
- ▶ 1:09:51 Philip Kiely Like the whole full-oh, save full-oh movement, like, you don't gotta have a save llama three movement, you just gotta have an eight 100 somewhere.
The Future of AI Infra: from Kubernetes to Agent Sandboxes — Akshat Bubna, Modal CTO
- ▶ 3:36 unnamed speaker Yeah, just like add a 100.
Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
- ▶ 4:00 Nader Khalil And, um, whenever we would talk to users, they wanted a GPU, they wanted an A-one hundred. 4 times in the scene
Dylan Patel Explains the AI War While Cooking | In-Context Cooking
- ▶ 42:46 Dylan Patel It was a large GPU, um, and it was like having, it was like the best memory, the best networking, everything, sort of the best as possible, um, and sort of like one size fits all, uh, with the main line of like A-one hundred, H-one…
[State of AI Papers 2025] Fixing Research with Social Signals, OCR & Implementation — Team AlphaXiv
- ▶ 5:18 unnamed speaker If you host it on your own A-One-Hundreds, and just like, you, you batch things properly, probably Deep Seek is best bang for your buck.
How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
- ▶ 4:32 Quentin Anthony Like, I would say MI-DX was not on the same level of eight 100.
Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
- ▶ 5:27 Barak Lenz So we designed J to have a version that fits on a single GPU, a single AY 100 or H one, 80 gigabytes.
A Technical History of Generative Media
- ▶ 22:50 Batuhan Taskaya And we, like, Kubernetes version at Google Cloud was fine in 2022 when we wanted to get eight A-one-hundreds. 2 times in the scene
⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
- ▶ 14:44 Thomas Sohmers And the A 100 comparison here is interesting because in most of these cases, they're actually, they,
The Shape of Compute (Chris Lattner of Modular)
- ▶ 4:08 Chris Lattner It ran just on a 100, just one model, but it had state of the art performance. 2 times in the scene
- ▶ 56:32 Chris Lattner Go look at VLM.
SF Compute: Commoditizing Compute
- ▶ 22:06 Evan Conrad We just, like, assumed we could go to, like, Lambda, um, or something, and, like, buy thousands of, at the time, A-One-Hundreds.
[Paper Club] BERT: Bidirectional Encoder Representations from Transformers
- ▶ 41:15 unnamed speaker Yeah, 20 dollars, they did, like, eight A-One hundreds for an hour, and they're able to match the glue store of basic BERT with their recipe.
Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
- ▶ 52:54 Joseph Nelson The smallest model is thirty-eight million parameters and can run at 45 FPS on an A-one hundred, right?
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
How AI is Eating Finance - with Mike Conover of Brightwave
- ▶ 9:57 Mike Conover The Matic contributors to an investment decision, um, or, or developing your thesis that, um, in response to export controls on a 100 cards, uh, China has put in place licensors on the transfer of germanium and gallium, which are not rare…
Why Google failed to make GPT-3 -- with David Luan of Adept
- ▶ 12:11 David Luan One interesting set of stuff is just like, you know, like knowing that a 100 generation that like quad sparsity was going to be a thing.
Truly Serverless Infra for AI Engineers - with Erik Bernhardsson of Modal
- ▶ 16:30 Erik Bernhardsson Yeah, it's like, you just, like, say, you know, on the function decorator, you're like, GPU equals, you know, a 100, and then, or like, GPU equals, you know, uh, a 10 or a T four, something like that, and then you get that GPU, and like,…
Building an open AI company - with Ce and Vipul of Together AI
- ▶ 27:40 Vipul Ved Prakash They're mostly A-hundreds and H-hundreds.
- ▶ 32:20 Alessio Fanelli And this post he said, our model indicates that together it's better off using two a 180 gig system rather than a each 100 based system. 2 times in the scene
- ▶ 42:07 Vipul Ved Prakash It's, uh, um, you know, the value we get from doing specific optimization, even, even for, you know, what works well for a particular model on A hundreds with a particular bus.
The AI-First Graphics Editor - with Suhail Doshi of Playground AI
- ▶ 21:33 Suhail Doshi And we did that for academic research because there's a whole bunch of, you know, we come across people all the time in academia and they have like, they have access to like one a 100 or eight at best.
- ▶ 1:02:19 unnamed speaker Uh, you, I, I, you had a tweet about like how many A 100 you have, but I feel like it's out of date probably.
The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
- ▶ 4:40 Dylan Patel Like 20,000 A-one hundreds. 2 times in the scene
- ▶ 16:12 Dylan Patel Uh, just like the A 180 gig did versus the A 140 gig. 4 times in the scene
RWKV: Reinventing RNNs for the Transformer Era
- ▶ 55:31 Eugene Cheah And donated the A-One-Hundreds needed to train the basic models that RWKB had.
FlashAttention-2: Making Transformers 800% faster AND exact
Ep 18: Petaflops to the People — with George Hotz of tinycorp
- ▶ 41:53 George Hotz Um, so, the bandwidth is the, is roughly 10 X less than what you can get with NV-linked A-Hundreds. 2 times in the scene