Insight certainty 4/5 debate potential 2/5

Uberti: AI Token Serving Requires Non-Linear Cluster Scaling Economics

Gavin Uberti · The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best · Jun 30, 2026 · at 1:29:19

Gavin Uberti, co-founder and CEO of AI chipmaker Etched, discusses the necessity of non-linear architectural scaling for large AI clusters.

0:00 / 0:19exact quote · 19.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“That the way you want this to scale is not that, oh, if I want to go serve 10 times more tokens, I buy 10 times more servers. It must be some solution where if I want to serve 10 times more tokens, then I get some economies of scale benefit with my say cluster scale memory tech that allows me to then not charge as much As 10 times more for those next tokens.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Gavin Uberti

Prediction Not checkable as stated
Uberti: Future AI token factories will cost $40B to $100B per mega-cluster
“You could have the same kind of thing for some futuristic mega cluster. Forty billion dollars, hundred billion dollars as a giant mega token factory serving one or a handful of models for a massive number of users to get that same economies of scale thing.”
Gavin Uberti Jun 30, 2026 ▶ 49:27 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
Prediction Not checkable as stated
Uberti: AI models will eventually do all kernel generation superhumanly
“And when the models keep getting smarter, they'll eventually do all of it. They will become superhuman.”
Gavin Uberti Jun 30, 2026 ▶ 51:35 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
Prediction Open · timeframe Jun 2031
Uberti: Individual trillion-dollar data centers are inevitable
“Absolutely. It is a matter of time. It's like asking, what do you say, a billion dollar fab, or a ten billion dollar fab, or a hundred billion dollar fab? It is inevitable that the economies of scale don't stop at, oh, forty billion dollars is the magic number…”
Gavin Uberti Jun 30, 2026 ▶ 1:27:04 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
Assertion Supported
Uberti: NVIDIA Blackwell point-to-point latency is about 4,000 nanoseconds
“For example, on Blackwell chips, it can be about 4000 nanoseconds to go point to point.”
Gavin Uberti Jun 30, 2026 ▶ 11:19 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
Insight
Uberti: High-speed AI decode is bottlenecked by data movement, not math
“When you do this sort of kernel's work, what you realize is that the math is relatively easy. But to get high speed decode, the thing that matters is data movement. Almost all the work that you do is optimizing how do you move data around a single chip or acro…”
Gavin Uberti Jun 30, 2026 ▶ 23:17 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
Assertion Not checkable as stated
Uberti: TSMC funded an experiment on Etched's recommendation and updated its line
“TSMC customer service is way, way better than I have seen at any other company in any other industry. It's the kind of thing where if you say, hey, you can approve your yield by making this change, you can go make them a recommendation, and then we'll go run a…”
Gavin Uberti Jun 30, 2026 ▶ 41:52 The Two Harvard Dropouts Who raised $800M to take on NVIDIA · Invest Like The Best
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 60 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.