Nvidia

includes NVIDIA H100, NVIDIA A100, NVIDIA Rubin, NVIDIA Blackwell, NVIDIA H200, NVIDIA Dynamo, NVIDIA Cosmos, NVIDIA GTC, NVIDIA Hopper, NVIDIA T4, NVIDIA GB300, NVIDIA A10 and 32 more

80 statements across 35 episodes · 38 bullish · 14 bearish · 34 people on the record · first statement Jun 20, 2023 by George Hotz · said 721 times in 100 episodes since 2023 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (80), Dylan Patel (55), Kyle Kranen (43), Ali Taha (26), Ethan He (24), Chris Lattner (24), George Hotz (23), Philip Kiely (20)

tap a year for its mentions
0020020400402023202420252026episodesmentions
020402023202420252026episodes it came up in
00102020402023202420252026episodesmentions per episode
2026 306 mentions in 30 episodes 10 per episode
2025 191 mentions in 33 episodes 6 per episode
2024 122 mentions in 31 episodes 4 per episode
2023 102 mentions in 6 episodes 17 per episode

every mention, scene by scene, with the transcript →

Everything said about Nvidia, oldest first

Jun 20, 2023 positive
Opinion
Hotz: Nvidia makes the best training chips
“NVIDIA has the best training chips.”
George Hotz Jun 20, 2023 ▶ 2:46 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Jun 20, 2023 neutral
Assertion Supported
Hotz: Tinygrad is about 5x slower than PyTorch on Nvidia GPUs
“The correctness for both forwards and backwards passes is there, but on Nvidia, it's about five X slower than PyTorch right now.”
George Hotz Jun 20, 2023 ▶ 21:15 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Jun 20, 2023 positive
Opinion
Hotz: Google TPUs are the only successful non-Nvidia training chips
“The only company, there's one other company aside from Nvidia who's succeeded at all at making training chips... Mid journey is trained on TPU, right? Like a lot of startups do actually train on TPUs, and they're the only other successful training chip aside f…”
George Hotz Jun 20, 2023 ▶ 4:33 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Nov 3, 2023 positive
Insight
Royzen: NVIDIA Remains Cloud-Agnostic Because It Wins Regardless
“At NVIDIA, They know that they're going to win regardless. So they don't care where you get the GPUs from. They're like, they're truly neutral, unlike various sales reps that you might encounter at various like clouds and, you know, hardware companies, et cete…”
Michael Royzen Nov 3, 2023 ▶ 1:02:39 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Nov 3, 2023 positive
Assertion Not publicly verifiable
Royzen: NVIDIA Built Custom FasterTransformer Feature for Phind
“They actually implemented a custom feature for us in Faster Transformer which is one of their libraries... They implemented streaming generation for T-Five-based models, which we were running at the time up until we switched to GPT in In February, March of thi…”
Michael Royzen Nov 3, 2023 ▶ 1:04:04 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Dec 5, 2023 bullish
Assertion Supported
Patel: Nvidia manufactured 400k H100s last quarter and will sell 530k this quarter
“There's 400 to 500,000 being, 400,000 manufactured last quarter, and like five 30,000 this quarter being sold, right, of H-Hundreds”
Dylan Patel Dec 5, 2023 ▶ 7:03 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 neutral
Assertion Not checkable as stated
Patel: Nvidia and Google control over 80% of advanced AI silicon manufacturing capacity
“To simplify it, NVIDIA has a little bit more than half, and Google has, like, 30%, right, through Broadcom. So it's like, the total capacity for everyone else is much lower, and they're all sharing it”
Dylan Patel Dec 5, 2023 ▶ 42:01 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 bullish
Prediction Partly held up
Patel: Intel will release a chip surpassing Nvidia H100 within a quarter
“Intel bought that company from him, and then shut it down, and bought this other AI company, and now that company is kind of, ah, you know, got new chips. They're gonna release a better chip than the H 100, ah, within the next quarter or so, right?”
Dylan Patel Dec 5, 2023 ▶ 41:01 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 bullish
Prediction Held up
Patel: Nvidia will sell over 3 million GPUs in 2024
“NVIDIA is going to sell well over three million, you know, total GPUs next year. You know, over a million H 100 this year alone, right?”
Dylan Patel Dec 5, 2023 ▶ 7:46 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 bullish
Prediction Partly held up
Patel: Nvidia to ship next-gen chip in Q2/Q3 2024 with 3x LLM performance
“Nvidia's releasing a new chip, you know, in, you know, they're gonna announce it in March, and they're gonna release it, you know, and ship it, you know, Q-two, Q-three next year anyways, right? And that chip will probably be three or four times as good. Right…”
Dylan Patel Dec 5, 2023 ▶ 44:00 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Dec 5, 2023 neutral
Assertion Not checkable as stated
Patel: Microsoft's custom AI chip performs worse than Nvidia H100
“Microsoft's going to announce their chip soon. It's worse performance than the H-H-E-N-H-E-D but the cost effectiveness of it is, is better for Microsoft internally, just because they don't have to pay the Nvidia tax.”
Dylan Patel Dec 5, 2023 ▶ 48:53 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Feb 8, 2024 neutral
Prediction Held up
Prakash predicts up to 5 million AI GPUs will sell in 2024
“There is four to five million GPUs that will be sold this year. NVIDIA and others.”
Vipul Ved Prakash Feb 8, 2024 ▶ 30:43 Building an open AI company - with Ce and Vipul of Together AI
Mar 6, 2024 bullish
Insight
Chintala: Hyperscalers build custom silicon to exploit vertical workload efficiencies
“Each large company has a sufficient enough set of verticalized workloads that have a pattern to them that, say, a more generic accelerator like an NVIDIA or an AMD GPU does not exploit. So there is some level of power efficiency that you're leaving on the tabl…”
Soumith Chintala Mar 6, 2024 ▶ 57:49 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Mar 6, 2024 bullish
Opinion
Chintala: Nvidia's primary competitive moat is NVLink interconnect, not GPU silicon
“The mode that Nvidia has right now, I feel like, is that they're, they have the interconnect that no one else has. Like, AMD GPUs are pretty good. I'm sure there's very silicon that is not bad at all, but, like, the interconnect like, NVLink is uniquely awesom…”
Soumith Chintala Mar 6, 2024 ▶ 18:57 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Oct 18, 2024 positive
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Drew Houston Oct 18, 2024 ▶ 52:59 Building the Silicon Brain - Drew Houston of Dropbox
Oct 19, 2024
Assertion Supported
Fanelli: Singapore accounted for 15% of NVIDIA's Q3 2024 revenue
“Singapore was 15% of NVIDIA's revenue in Q three of 2024.”
Alessio Fanelli Oct 19, 2024 ▶ 42:31 Singapore: the AI Engineer Nation — with Minister Josephine Teo
Oct 29, 2024 positive
Insight
Ethan He: Peak pre-training learning rate works best for MoE upcycling
“We found that the best is to use the original highest peak learning rate from pre-training, which works the best.”
Ethan He Oct 29, 2024 ▶ 33:34 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Oct 29, 2024 positive
Insight
Ethan He: 64 experts is the sweet spot for MoE upcycling
“We found, 64 experts is kind of like the sweet spot. If you increase the number of experts beyond 64, it provides diminishing return.”
Ethan He Oct 29, 2024 ▶ 35:03 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Oct 29, 2024 negative
Assertion Supported
Ethan He: Hugging Face's sequential GEMM loop for Mixtral is inefficient
“Let's also look at the implementation of Mixtro eight by seven on Hagen-Phys transformer. You will soon notice the, in the expert operation there, You would iterate over all of the experts and compute each of the gem operations one by one. We found that this i…”
Ethan He Oct 29, 2024 ▶ 12:37 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Dec 7, 2024 positive
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Dec 24, 2024
Assertion Supported
Ben Allal: NVIDIA generated 1.9 trillion synthetic tokens for Nemotron-CC
“This is a recent paper from NVIDIA, Mnemotron CC. They took things a bit further and they generated not a few billion tokens, but 1.9 trillion tokens, which is huge.”
Loubna Ben Allal Dec 24, 2024 ▶ 8:45 Best of 2024: Synthetic Data / Smol Models, Loubna Ben Allal, HuggingFace [LS Live! @ NeurIPS 2024]
Jan 1, 2025 bearish
Assertion Not checkable as stated
Swyx: VC appetite for GPU-rich early-stage startups is completely gone
“The appetite for GPU rich startups, like the, you know, the funding plan is we will raise sixty million and we'll give 50 of that to Nvidia. That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like …”
Shawn Wang Jan 1, 2025 ▶ 38:52 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Jan 4, 2025 negative
Opinion
AMD GPUs perform horribly on Windows for generative AI tasks
“AMD works horribly on Windows. Like, on Linux, it works fine. It's lower than the price equivalent NVIDIA GPU. But it works, like, you can use it, generate images, everything works. On Linux, on Windows, You might have a hard time”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 34:09 AI Engineering for Art - with comfyanonymous
Apr 11, 2025 neutral
Insight
Conrad: NVIDIA avoids customer concentration to prevent hyperscaler price setting
“It's really bad for NVIDIA if you have customer concentration, and Microsoft and Google and Amazon, like, Oracle to, like, buy up your entire supply and then you have four or five customers or so who pretty much get to set prices.”
Evan Conrad Apr 11, 2025 ▶ 15:46 SF Compute: Commoditizing Compute
Apr 11, 2025 bearish
Prediction Not checkable as stated
Conrad: Hyperscalers will probably lose significant money reselling Nvidia GPUs
“My intuition is that the hyperscalers are probably going to lose a lot of money, and they know they're going to lose a lot of money on reselling NVIDIA GPUs at least.”
Evan Conrad Apr 11, 2025 ▶ 6:28 SF Compute: Commoditizing Compute
Apr 11, 2025 neutral
Insight
Conrad: NVIDIA launching a cloud would impair hyperscaler sales
“If they launched their own core weave, then it would make it much harder for them to sell to the hyperscalers.”
Evan Conrad Apr 11, 2025 ▶ 15:00 SF Compute: Commoditizing Compute
May 6, 2025 positive
Assertion Supported
NVIDIA's 600M Parameter Parakeet Model Tops Speech Transcription Leaderboards
“And then Parakeet is Nvidia's new speech model, speech transcription model. That's number one on the leaderboards. And it's just like very enterprise tuned, like really, really rock solid, reliable at fairly small number of weights, like six hundred million pa…”
Kwindla Hultman Kramer May 6, 2025 ▶ 8:07 Voice AI Masterclass — Kwindla Hultman Kramer and swyx
Jun 13, 2025 negative
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Chris Lattner Jun 13, 2025 ▶ 12:25 The Shape of Compute (Chris Lattner of Modular)
Jun 13, 2025
Disclosure
Modular's Mojo and MAX are free on NVIDIA and CPUs
“The max framework and the mojo language, free to use on NVIDIA and CPUs, any scale, go nuts, do whatever you want.”
Chris Lattner Jun 13, 2025 ▶ 41:47 The Shape of Compute (Chris Lattner of Modular)
Jun 13, 2025
Disclosure
Modular completely eliminates and replaces NVIDIA's CUDA stack
“In the case of Modular, we go literally, like, we only work at that level, because we get rid of all of CUDA, right? And so we've replaced the entire stack, and so we only do that.”
Chris Lattner Jun 13, 2025 ▶ 55:03 The Shape of Compute (Chris Lattner of Modular)
Jul 31, 2025
Assertion Partly supported
Swix: Nvidia RTX 4090 prices doubled in the past year
“40 and 90 prices have doubled in the last year.”
Shawn Wang Jul 31, 2025 ▶ 1:10:08 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Aug 18, 2025 neutral
Assertion Supported
Sohmers: NVIDIA's TF32 is actually a 19-bit precision format
“NVIDIA's TF-thirty-two number format is a nineteen-bit number format. They just call it thirty-two-bit.”
Thomas Sohmers Aug 18, 2025 ▶ 27:29 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Aug 18, 2025 bullish
Prediction Not checkable as stated
Sohmers: NVIDIA will have a very good decade ahead despite startup challengers
“The reality is NVIDIA is going to have a very, very good decade ahead for them. And The market is growing so fast that all of us in the space trying to take them on can be very happy with, you know, very, very small wins in the space.”
Thomas Sohmers Aug 18, 2025 ▶ 34:26 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Aug 18, 2025 bearish
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Thomas Sohmers Aug 18, 2025 ▶ 14:53 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Aug 18, 2025 bullish
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Thomas Sohmers Aug 18, 2025 ▶ 16:28 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Aug 18, 2025 neutral
Disclosure
Agrawal: Positron cannot beat NVIDIA on matrix-matrix performance per dollar today
“Thomas said that, you know, we are accelerating matrix vector. It doesn't mean that we can't do matrix, matrix. We can, and we can do it fairly well. It's just that the point becomes is like, are you really beating NVIDIA on it from a perf per dollar? And if y…”
Mitesh Agrawal Aug 18, 2025 ▶ 35:54 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Aug 18, 2025 bullish
Prediction Not checkable as stated
Sohmers: People will continue buying NVIDIA for AI training
“We are betting that people are going to continue to train on NVIDIA for at least the foreseeable future, where, since we're able to, you know, and I'll say, I really hope others are able to be successful in, in providing competition against NVIDIA, but Given t…”
Thomas Sohmers Aug 18, 2025 ▶ 22:17 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Sep 8, 2025 bearish
Opinion
Taskaya: Custom ASICs do not make sense given low Nvidia GEMM overhead
“What is the overhead of an NVIDIA GAM instruction, right? It's like 16%. So like you're essentially buying a, Matrix multiplication machine. So, like, it doesn't really make sense to specialize it that much.”
Batuhan Taskaya Sep 8, 2025 ▶ 25:25 A Technical History of Generative Media
Oct 1, 2025 bearish
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Andrew Feldman Oct 1, 2025 ▶ 3:05 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Nov 3, 2025 neutral
Assertion Supported
AMD's Composable Kernel library offers functionality similar to Nvidia's CUTLASS
“And then on the AMD side, they have this composable kernel library that does something very similar.”
Quentin Anthony Nov 3, 2025 ▶ 12:07 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Nov 3, 2025 bullish
Opinion
AMD has completely caught up with Nvidia on AI software
“So they caught up on hardware. Now they've caught up on software. Not a lot of people have sort of discovered that they've caught up on software and we're kind of capitalizing on that.”
Quentin Anthony Nov 3, 2025 ▶ 5:10 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Nov 25, 2025 bearish
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Justin Johnson Nov 25, 2025 ▶ 13:01 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Feb 10, 2026 positive
Assertion Partly supported
Datology BeyondWeb 3B matches NVIDIA Nemotron 8B in 2.7x less training time
“As you can see that we achieved the same performance as the NVIDIA model in almost, like, 2.7 X, like, lesser time. And then much faster than anything that hugging face or pajama does. Very interestingly, our three B model is pretty much the same performance a…”
Pratyush Maini Feb 10, 2026 ▶ 21:35 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Feb 24, 2026 bullish
Opinion
O'Laughlin: TPU v7 is Google's peak TCO advantage over Nvidia
“I think Ironwood a V seven is the peak gap between on Between TCO, between NVIDIA and TPU, right?”
Doug O'Laughlin Feb 24, 2026 ▶ 1:35:07 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Feb 24, 2026 bearish
Prediction Not checkable as stated
O'Laughlin: TPU v8 will not compete as well against Nvidia Rubin
“V-A we just don't think will be as competitive to Ruben, and that's when your special window starts to close.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:38:59 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Feb 24, 2026 bullish
Opinion
O'Laughlin: Nvidia commands the best supply chain bar none
“I think on the infrastructure side, or sorry, on the supply chain side, bar none, NVIDIA is the best. They own the entire supply chain. They really do.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:39:55 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Feb 26, 2026 neutral
Insight
Patel: Nvidia must vastly outperform rivals to overcome vertical integration
“Google, Amazon, they get to vertically integrate, integrate, and vertical integration always saves tons of money. So he has to be better than everyone. By, not just like a little bit, by a ton. To justify his margins. Otherwise, the vertical integration of his…”
Dylan Patel Feb 26, 2026 ▶ 44:05 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Mar 8, 2026 neutral
Disclosure
NVIDIA mandates running OpenClaw in isolated Brev cloud VMs
“Internally people want to run this and we know we have to be really careful from the security implications. Do we let this run on the corporate network securities guidance was, Hey, run this on breath. It's in, you know, it's a VM. It's sitting in the cloud. …”
Nader Khalil Mar 8, 2026 ▶ 9:08 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 neutral
Assertion Supported
NVIDIA chip design begins three to five years before market release
“The design process starts like- Exactly. ...three to five years before the chip gets to the market.”
Kyle Kranen Mar 8, 2026 ▶ 22:39 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 bullish
Assertion Supported
NVIDIA announces Rubin CPX as a dedicated prefill-specific hardware accelerator
“And like with our future generations, generations of hardware, we actually announced like with Rubin, this new accelerator that is pre-fill specific. It's called Rubin CPX.”
Kyle Kranen Mar 8, 2026 ▶ 42:04 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 positive
Assertion Supported
NVIDIA Dynamo dynamically sizes and schedules Kubernetes prefill and decode workers
“Dynamo has a set of components that A, tell you how to scale. It tells you how many pre-fill workers and decoded workers it thinks you should have. And also provides a scheduling API for Kubernetes that allows you to actually represent and affect this scheduli…”
Kyle Kranen Mar 8, 2026 ▶ 43:57 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 positive
Disclosure
NVIDIA's Brev team plans to open-source business application command-line interfaces
“We're gonna open source all of this and like, yeah, all the, I mean, they're just, they're, yeah, CLIs for the business applications.”
Nader Khalil Mar 8, 2026 ▶ 1:08:58 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 positive
Insight
NVIDIA uses 'Speed of Light' framework to define minimum viable goals
“SOL is a term that NVIDIA is used to sort of like instigate a compelling event. You say, this is done. How do we get there? What is the minimum, as much as necessary, as little as possible thing that it takes for us to get exactly here? And it helps you just b…”
Kyle Kranen Mar 8, 2026 ▶ 15:28 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 bullish
Prediction Not checkable as stated
Algorithmic unhobblers could soon expand context windows to 100 million tokens
“I wouldn't be surprised if we do see the ability to, like, break through to, like, ten million, twenty million, a hundred million context through the, an unhobbler showing up.”
Kyle Kranen Mar 8, 2026 ▶ 52:37 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 neutral
Assertion Not checkable as stated
NVIDIA's build.nvidia.com was internally the company's largest inference deployment
“At one point, there's a website called build.nvd.com, and also for us, inference.nvd.com, that is, allows people to try models. It gives an API service, you can call the model with like a REST API, and, you know, you get a response. I ran the model site for th…”
Kyle Kranen Mar 8, 2026 ▶ 1:01:02 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Mar 8, 2026 positive
Assertion Supported
Jensen Huang prioritizes strategic investments in 'zero billion dollar markets'
“Jensen, He says, we're completely happy investing in zero billion dollar markets. We don't care if this creates revenue. It's important for us to know about this market. We think it will be important in the future. It can be zero billion dollars for a while.”
Kyle Kranen Mar 8, 2026 ▶ 25:14 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Apr 2, 2026 neutral
Assertion Not checkable as stated
Sun: NVIDIA pays heavily to purchase interactive simulation worlds for robotics
“In industry, like folks at NVIDIA are actually paying a lot of dollars to purchase these types of interactive worlds, whether it's for the sake of evaluation or training the robots or policies or models.”
Fan-yun Sun Apr 2, 2026 ▶ 2:41 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Apr 2, 2026 bullish
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Fan-yun Sun Apr 2, 2026 ▶ 2:56 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Apr 3, 2026 bullish
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Marc Andreessen Apr 3, 2026 ▶ 23:23 Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
Apr 22, 2026
Assertion Supported
Parakhin: Bing Sydney first launched in India using Megatron, not OpenAI
“The funny thing, I mean, the most interesting anecdote is that Sydney was first shipped in India for and it was not noticed for a long time. And first implementation of Sydney didn't even have open AI model under it. It was during Megatron. Microsoft and the N…”
Mikhail Parakhin Apr 22, 2026 ▶ 1:10:53 AI-Native Engineering: 100% adoption, 5x search throughput, unlimited tokens — Mikhail Parakhin
Jun 1, 2026 positive
Assertion Supported
Ethan He: Megatron MoE was first to train trillion-parameter MoEs at 40% MFU
“The Megatron MOEs was the first It was the first framework open source to be able to train these MOEs at very large scales, like a hundred billion parameters to even trillion parameters efficiently at like 40% MFU.”
Ethan He Jun 1, 2026 ▶ 1:41:50 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Disclosure
NVIDIA Cosmos trained on tens of trillions of visual tokens
“And if you look at a number of tokens we disclose that in cosmos, it's also like tens of trillions of tokens. On the visual tokens.”
Ethan He Jun 1, 2026 ▶ 37:55 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026
Disclosure
Ethan He: NVIDIA Cosmos required labelers to describe videos for blind reconstruction
“So that's in the protocol of Cosmos labeling. We required the objective we gave to the labelers was that you have to describe the video as detailed as possible, such that a blind person hears a blob of text, can reconstruct what the video is like from their he…”
Ethan He Jun 1, 2026 ▶ 13:39 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 positive
Insight
Ethan He: Video Foundation Models Follow Scaling Laws Like LLMs
“There, once I built the Cosmos one, I realized as this thing also has a scaling law similar to language model.”
Ethan He Jun 1, 2026 ▶ 3:11 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Assertion Not checkable as stated
Ethan He: NVIDIA spent about a year building the Cosmos model
“One thing I say, like, thanks to my experience at NVIDIA, because first time when we were building Cosmos together, we built it for about a year.”
Ethan He Jun 1, 2026 ▶ 5:13 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Assertion Supported
NVIDIA Cosmos uses 50,000 to 60,000 tokens for five seconds of video
“Yeah, for example, like in Cosmos, I think just five seconds of video is like a 50, 50 K or a 60 K number of tokens. So like, if you do 50 seconds as a 500 K tokens, if you do longer than that, easily explode.”
Ethan He Jun 1, 2026 ▶ 56:16 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 1, 2026 neutral
Disclosure
He: NVIDIA Cosmos runs in 4 to 8 steps, or 1 step for transfer
“In Cosmos, I believe we have like four steps and eight steps. If you do some simpler task, like image to image translation, it can even run in first step, that one step in, in Cosmos transfer.”
Ethan He Jun 1, 2026 ▶ 40:23 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Jun 18, 2026 positive
Assertion Open · timeframe Jun 2029
Midha: MatX chips adopt NVIDIA reference architecture to plug into existing sites
“When they decided to pick the standard for their data center, they picked the NVIDIA reference architecture. So the Matex chips just plug in to any site that has an NVIDIA bring up planned. And you know.”
Anjney Midha Jun 18, 2026 ▶ 31:28 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Jun 30, 2026 bullish
Prediction Not checkable as stated
Feinberg: Chipmakers will invest in life sciences as LLM alpha shrinks
“I do think that chip makers, including Nvidia, are going to want to get a lot more invested in life sciences because it will always be high in demand. And The amount of alpha left in pure LLM space is just getting a little questionable.”
Evan Feinberg Jun 30, 2026 ▶ 1:44:44 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Jul 10, 2026 neutral
Opinion
Swyx: Etched is not disrupting NVIDIA because inference requires dedicated ASICs
“I don't think they want to disrupt NVIDIA. I think the better framing is that all of inference is just so goddamn big that, of course, you're gonna have ASICs for inference.”
Shawn Wang Jul 10, 2026 ▶ 3:13 Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv
Jul 16, 2026 positive
Disclosure
Beam: Lila avoids pre-training from scratch, builds on open-weight models
“We have not decided to take on pre-training as well, just because the black magic that you have to do is, is insane, and we've been gifted, you know, something like a billion dollars worth of compute in the form of open-weight models. So we start with an open-…”
Andy Beam Jul 16, 2026 ▶ 1:21:02 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Aug 3, 2026 bearish
Opinion
AI ASIC startups are doomed as NVIDIA GPUs become domain-specialized
“How can you look at this trend and then still be bullish on companies that are coming up with Asics for AI.”
Ali Taha Aug 3, 2026 ▶ 1:04:48 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 bearish
Assertion Supported
TurboQuant is inefficient on high-bandwidth data center GPUs like NVIDIA B200
“TurboQuant would not be like, it would not be used. Like Nvidia made it clear that this is not a good optimization. And we've seen it firsthand where the overhead of doing dequantization, quantization of You know, in the kernel itself, the turbo-quant kernel, …”
Ali Taha Aug 3, 2026 ▶ 50:02 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 positive
Prediction Not checkable as stated
NVIDIA or model creators will always publish open quantized checkpoints
“There's always going to be like an open source quantized checkpoint. NVIDIA is going to push one out if no one else does. You usually the providers will have their own spec tech that they've trained as well. You don't need to train your own spec tech. You can …”
Ali Taha Aug 3, 2026 ▶ 41:22 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 neutral
Opinion
NVIDIA Dynamo is a developer toolkit, not an out-of-the-box performance optimizer
“I would think of Dynamo as less of a sort of out of box system and more of a toolkit for building with. So when we talk about doing KV aware routing, when we talk about doing KV out offloading, when we talk about doing PD disaggregation, Dynamo fundamentally i…”
Philip Kiely Aug 3, 2026 ▶ 42:20 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 bearish
Prediction Not checkable as stated
NVIDIA's Vera Rubin GPU architecture effectively renders mega kernel research obsolete
“The GPU is, is It's designed in such a way that it basically kills megakernels. You don't need to use megakernels that much anymore. So it seems like that entire research field goes into, like, won't be continued, but yeah.”
Ali Taha Aug 3, 2026 ▶ 59:20 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Aug 3, 2026 positive
Assertion Not checkable as stated
NVIDIA's Rubin is the first GPU architecture designed entirely for modern LLMs
“Ruben's honestly the first chip that was fully built in that world. And so you can see a lot of the understanding of the shape of the workload that this chip's going to be asked to do in the way it's designed.”
Philip Kiely Aug 3, 2026 ▶ 1:06:16 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Sep 2, 2026 neutral
Insight
Lie: 100 to 200 tokens per second is becoming the new batch mode
“What used to be considered fast at, like, 1000 100 or 200 tokens per second is quickly becoming the new batch mode. Right. Quickly becoming the new, like overnight is true.”
Sean Lie Sep 2, 2026 ▶ 32:29 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 bearish
Opinion
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Sean Lie Sep 2, 2026 ▶ 30:35 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sep 2, 2026 bullish
Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Sean Lie Sep 2, 2026 ▶ 16:30 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.