Hopper architecture

also referred to as: hopper · hoppers

4 statements across 4 episodes · 1 bullish · 2 bearish · 4 people on the record · first statement Aug 18, 2025 by Thomas Sohmers · said 4 times in 2 episodes since 2025 · across every show →

Mentions by year

brought up most by Philip Kiely (2), Chris Lattner (2)

tap a year for its mentions
00112120252026episodesmentions
01120252026episodes it came up in
0010.52120252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Hopper architecture, oldest first

Aug 18, 2025 bearish
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Thomas Sohmers Aug 18, 2025 ▶ 14:53 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Nov 25, 2025 bearish
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Justin Johnson Nov 25, 2025 ▶ 13:01 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Jul 22, 2026 bullish
Assertion Not checkable as stated
Kant: Nvidia GB300 allows models to jump from 1T to 6T parameters
“The difference of a model you could train on hoppers versus GB 300 is the difference in a trillion parameter model and a five or six trillion parameter model.”
Eiso Kant Jul 22, 2026 ▶ 1:36:46 The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
Aug 3, 2026
Assertion Not checkable as stated
Unoptimized GLM-5.2 delivers a baseline 30 to 40 tokens per second
“So let's say you have, as a reasonable baseline, 30 or 40 tokens per second. You can achieve 10 X that. So like on GLM 5.2 if you want to get unquantized perhaps on hoppers even and you're just using an off the shelf inference engine with no particular optimiz…”
Philip Kiely Aug 3, 2026 ▶ 37:05 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.