Uberti: NVIDIA Blackwell point-to-point latency is about 4,000 nanoseconds
“For example, on Blackwell chips, it can be about 4000 nanoseconds to go point to point.”
Uberti: High-speed AI decode is bottlenecked by data movement, not math
“When you do this sort of kernel's work, what you realize is that the math is relatively easy. But to get high speed decode, the thing that matters is data movement. Almost all the work that you do is optimizing how do you move data around a single chip or acro…”
Uberti: Etched sent a dozen engineers to Bangalore to fix delays
“We went out and shipped a dozen of our top engineers across the world to Bangalore for six months. I was there as well. I lived in Bangalore for four and a half months personally”
Uberti: The best kernels are written by human-AI collaboration
“Today, it's all very hybrid, and the best kernels are still written by human AI collaborations.”
Uberti: Etched aligned two clock signals to within 50 picoseconds
“We realized there was one and only one way to solve it. As we had to go line up two clock signals on our chip to within 50 picoseconds. That is literally 50 trillionths of a second. And we had to go get these signals aligned to this super small granularity and…”
Uberti: Employees quit Etched over deemed unsolvable chip design problem
“We had people quit. That people literally were like, this problem is unsolvable, and, ah, best of luck guys.”
Uberti: TSMC VP Emailed Wanting to Partner with Etched After One Dinner
“And the following day, I get an email from TSMC saying, Gavin, I want to work with Etched. Find a way to make it happen.”
Uberti: AI Token Serving Requires Non-Linear Cluster Scaling Economics
“That the way you want this to scale is not that, oh, if I want to go serve 10 times more tokens, I buy 10 times more servers. It must be some solution where if I want to serve 10 times more tokens, then I get some economies of scale benefit with my say cluster…”