Assertion Open AI assessment confidence: 85% certainty 4/5 debate potential 3/5

Movva: NVIDIA is producing 5 million Blackwell chips this year

Neil Movva · Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper · Aug 25, 2026 · at 1:04:57

Neil Movva, co-founder of Sail Research and ex-NVIDIA engineer, discusses global GPU supply and orchestration inefficiencies.

0:00 / 0:02exact quote · 2.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“NVIDIA's pumping out five million Blackwell chips this year.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Neil Movva

Opinion
Losing TSMC Wouldn't Be Disastrous Because Intel Is Only 2x Behind
“And my contrarian take is that it wouldn't be that bad. Supply would take a shock for sure, but The best processes that we have in the West, like Intel, not that far behind, at worst, like maybe two X worst performance per watt. And the gap is just far smaller…”
Neil Movva Aug 25, 2026 ▶ 1:17:13 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Disclosure
Movva: Sail Research will buy data centers with only 95% uptime
“You'd have missing zero buyers for a data center that is 95% uptime. I'm that first buyer. I will buy 95% uptime.”
Neil Movva Aug 25, 2026 ▶ 57:26 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Assertion Not checkable as stated
Nvidia's BF16 Efficiency Barely Improves Across Hopper, Blackwell, and Rubin
“If you look at, you know, Hopper to Blackwell to Rubin, and you compare like for like, what is the performance per watt of a beefload-sixteen multiply? It hasn't improved all that much.”
Neil Movva Aug 25, 2026 ▶ 1:16:40 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Assertion Not checkable as stated
TSMC Performance Per Watt Barely Changes From 5nm to 2nm
“If you look at TSMC five nanometer versus four versus three versus two, the performance per watt on these chips doesn't change like a dramatic amount.”
Neil Movva Aug 25, 2026 ▶ 1:16:56 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Assertion Not checkable as stated
Cursor Forced Baseten, Fireworks, and Together to Prioritize Low Latency
“The challenge is, all those companies, you could take your pick, Base-Ten, Fireworks together, they all focus on low latency inference. And they were pulled in that direction by one very important customer Cursor.”
Neil Movva Aug 25, 2026 ▶ 3:42 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Prediction Not checkable as stated
Future AI Agents Will Prioritize Long-Horizon Task Completion Over Speed
“You wanted more persistence, more long horizon tasks, and now it's, to me, very obvious that the future of agentic inference is long horizon tasks. You're gonna run the machine for hours or days at a time. It doesn't matter if it spits out tokens at a hundred …”
Neil Movva Aug 25, 2026 ▶ 4:02 Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 60 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.