The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 80 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 2 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Sean Lie Sep 2, 2026 ▶ 16:30 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Opinion
AI ASIC startups are doomed as NVIDIA GPUs become domain-specialized
“How can you look at this trend and then still be bullish on companies that are coming up with Asics for AI.”
Ali Taha Aug 3, 2026 ▶ 1:04:48 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Marc Andreessen Apr 3, 2026 ▶ 23:23 Marc Andreessen introspects on Death of the Browser, Pi + OpenClaw, and Why "This Time Is Different"
Opinion
AMD has completely caught up with Nvidia on AI software
“So they caught up on hardware. Now they've caught up on software. Not a lot of people have sort of discovered that they've caught up on software and we're kind of capitalizing on that.”
Quentin Anthony Nov 3, 2025 ▶ 5:10 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Andrew Feldman Oct 1, 2025 ▶ 3:05 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Chris Lattner Jun 13, 2025 ▶ 12:25 The Shape of Compute (Chris Lattner of Modular)
Prediction Not checkable as stated
Conrad: Hyperscalers will probably lose significant money reselling Nvidia GPUs
“My intuition is that the hyperscalers are probably going to lose a lot of money, and they know they're going to lose a lot of money on reselling NVIDIA GPUs at least.”
Evan Conrad Apr 11, 2025 ▶ 6:28 SF Compute: Commoditizing Compute
Opinion
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Sean Lie Sep 2, 2026 ▶ 30:35 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Prediction Not checkable as stated
NVIDIA's Vera Rubin GPU architecture effectively renders mega kernel research obsolete
“The GPU is, is It's designed in such a way that it basically kills megakernels. You don't need to use megakernels that much anymore. So it seems like that entire research field goes into, like, won't be continued, but yeah.”
Ali Taha Aug 3, 2026 ▶ 59:20 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Prediction Not checkable as stated
Feinberg: Chipmakers will invest in life sciences as LLM alpha shrinks
“I do think that chip makers, including Nvidia, are going to want to get a lot more invested in life sciences because it will always be high in demand. And The amount of alpha left in pure LLM space is just getting a little questionable.”
Evan Feinberg Jun 30, 2026 ▶ 1:44:44 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Fan-yun Sun Apr 2, 2026 ▶ 2:56 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Prediction Not checkable as stated
Algorithmic unhobblers could soon expand context windows to 100 million tokens
“I wouldn't be surprised if we do see the ability to, like, break through to, like, ten million, twenty million, a hundred million context through the, an unhobbler showing up.”
Kyle Kranen Mar 8, 2026 ▶ 52:37 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Insight
Patel: Nvidia must vastly outperform rivals to overcome vertical integration
“Google, Amazon, they get to vertically integrate, integrate, and vertical integration always saves tons of money. So he has to be better than everyone. By, not just like a little bit, by a ton. To justify his margins. Otherwise, the vertical integration of his…”
Dylan Patel Feb 26, 2026 ▶ 44:05 Dylan Patel Explains the AI War While Cooking | In-Context Cooking
Opinion
O'Laughlin: TPU v7 is Google's peak TCO advantage over Nvidia
“I think Ironwood a V seven is the peak gap between on Between TCO, between NVIDIA and TPU, right?”
Doug O'Laughlin Feb 24, 2026 ▶ 1:35:07 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Prediction Not checkable as stated
O'Laughlin: TPU v8 will not compete as well against Nvidia Rubin
“V-A we just don't think will be as competitive to Ruben, and that's when your special window starts to close.”
Doug O'Laughlin Feb 24, 2026 ▶ 1:38:59 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Justin Johnson Nov 25, 2025 ▶ 13:01 After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs
Opinion
Taskaya: Custom ASICs do not make sense given low Nvidia GEMM overhead
“What is the overhead of an NVIDIA GAM instruction, right? It's like 16%. So like you're essentially buying a, Matrix multiplication machine. So, like, it doesn't really make sense to specialize it that much.”
Batuhan Taskaya Sep 8, 2025 ▶ 25:25 A Technical History of Generative Media
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Thomas Sohmers Aug 18, 2025 ▶ 14:53 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Thomas Sohmers Aug 18, 2025 ▶ 16:28 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Disclosure
Modular completely eliminates and replaces NVIDIA's CUDA stack
“In the case of Modular, we go literally, like, we only work at that level, because we get rid of all of CUDA, right? And so we've replaced the entire stack, and so we only do that.”
Chris Lattner Jun 13, 2025 ▶ 55:03 The Shape of Compute (Chris Lattner of Modular)
Opinion
AMD GPUs perform horribly on Windows for generative AI tasks
“AMD works horribly on Windows. Like, on Linux, it works fine. It's lower than the price equivalent NVIDIA GPU. But it works, like, you can use it, generate images, everything works. On Linux, on Windows, You might have a hard time”
comfyanonymous (Comfy) Jan 4, 2025 ▶ 34:09 AI Engineering for Art - with comfyanonymous
Assertion Not checkable as stated
Swyx: VC appetite for GPU-rich early-stage startups is completely gone
“The appetite for GPU rich startups, like the, you know, the funding plan is we will raise sixty million and we'll give 50 of that to Nvidia. That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like …”
Shawn Wang Jan 1, 2025 ▶ 38:52 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Sarah Chieng Dec 7, 2024 ▶ 3:06 [Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)
Insight
Ethan He: Peak pre-training learning rate works best for MoE upcycling
“We found that the best is to use the original highest peak learning rate from pre-training, which works the best.”
Ethan He Oct 29, 2024 ▶ 33:34 [Paper Club] Upcycling Large Language Models into Mixture of Experts
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Drew Houston Oct 18, 2024 ▶ 52:59 Building the Silicon Brain - Drew Houston of Dropbox
Opinion
Chintala: Nvidia's primary competitive moat is NVLink interconnect, not GPU silicon
“The mode that Nvidia has right now, I feel like, is that they're, they have the interconnect that no one else has. Like, AMD GPUs are pretty good. I'm sure there's very silicon that is not bad at all, but, like, the interconnect like, NVLink is uniquely awesom…”
Soumith Chintala Mar 6, 2024 ▶ 18:57 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Prediction Partly held up
Patel: Intel will release a chip surpassing Nvidia H100 within a quarter
“Intel bought that company from him, and then shut it down, and bought this other AI company, and now that company is kind of, ah, you know, got new chips. They're gonna release a better chip than the H 100, ah, within the next quarter or so, right?”
Dylan Patel Dec 5, 2023 ▶ 41:01 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Prediction Partly held up
Patel: Nvidia to ship next-gen chip in Q2/Q3 2024 with 3x LLM performance
“Nvidia's releasing a new chip, you know, in, you know, they're gonna announce it in March, and they're gonna release it, you know, and ship it, you know, Q-two, Q-three next year anyways, right? And that chip will probably be three or four times as good. Right…”
Dylan Patel Dec 5, 2023 ▶ 44:00 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Not checkable as stated
Patel: Microsoft's custom AI chip performs worse than Nvidia H100
“Microsoft's going to announce their chip soon. It's worse performance than the H-H-E-N-H-E-D but the cost effectiveness of it is, is better for Microsoft internally, just because they don't have to pay the Nvidia tax.”
Dylan Patel Dec 5, 2023 ▶ 48:53 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Opinion
Hotz: Google TPUs are the only successful non-Nvidia training chips
“The only company, there's one other company aside from Nvidia who's succeeded at all at making training chips... Mid journey is trained on TPU, right? Like a lot of startups do actually train on TPUs, and they're the only other successful training chip aside f…”
George Hotz Jun 20, 2023 ▶ 4:33 Ep 18: Petaflops to the People — with George Hotz of tinycorp
Insight
Lie: 100 to 200 tokens per second is becoming the new batch mode
“What used to be considered fast at, like, 1000 100 or 200 tokens per second is quickly becoming the new batch mode. Right. Quickly becoming the new, like overnight is true.”
Sean Lie Sep 2, 2026 ▶ 32:29 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Disclosure
Beam: Lila avoids pre-training from scratch, builds on open-weight models
“We have not decided to take on pre-training as well, just because the black magic that you have to do is, is insane, and we've been gifted, you know, something like a billion dollars worth of compute in the form of open-weight models. So we start with an open-…”
Andy Beam Jul 16, 2026 ▶ 1:21:02 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Opinion
Swyx: Etched is not disrupting NVIDIA because inference requires dedicated ASICs
“I don't think they want to disrupt NVIDIA. I think the better framing is that all of inference is just so goddamn big that, of course, you're gonna have ASICs for inference.”
Shawn Wang Jul 10, 2026 ▶ 3:13 Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv
Insight
NVIDIA uses 'Speed of Light' framework to define minimum viable goals
“SOL is a term that NVIDIA is used to sort of like instigate a compelling event. You say, this is done. How do we get there? What is the minimum, as much as necessary, as little as possible thing that it takes for us to get exactly here? And it helps you just b…”
Kyle Kranen Mar 8, 2026 ▶ 15:28 Agent Inference at the "Speed of Light" — How NVIDIA moves like a $4.3 Trillion Startup
Assertion Partly supported
Datology BeyondWeb 3B matches NVIDIA Nemotron 8B in 2.7x less training time
“As you can see that we achieved the same performance as the NVIDIA model in almost, like, 2.7 X, like, lesser time. And then much faster than anything that hugging face or pajama does. Very interestingly, our three B model is pretty much the same performance a…”
Pratyush Maini Feb 10, 2026 ▶ 21:35 ⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
Disclosure
Agrawal: Positron cannot beat NVIDIA on matrix-matrix performance per dollar today
“Thomas said that, you know, we are accelerating matrix vector. It doesn't mean that we can't do matrix, matrix. We can, and we can do it fairly well. It's just that the point becomes is like, are you really beating NVIDIA on it from a perf per dollar? And if y…”
Mitesh Agrawal Aug 18, 2025 ▶ 35:54 ⚡️Accelerators @ 3x NVIDIA H200 perf, Made in the USA - Thomas Sohmers + Mitesh Agrawal, Positron AI
Insight
Chintala: Hyperscalers build custom silicon to exploit vertical workload efficiencies
“Each large company has a sufficient enough set of verticalized workloads that have a pattern to them that, say, a more generic accelerator like an NVIDIA or an AMD GPU does not exploit. So there is some level of power efficiency that you're leaving on the tabl…”
Soumith Chintala Mar 6, 2024 ▶ 57:49 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Prediction Held up
Prakash predicts up to 5 million AI GPUs will sell in 2024
“There is four to five million GPUs that will be sold this year. NVIDIA and others.”
Vipul Ved Prakash Feb 8, 2024 ▶ 30:43 Building an open AI company - with Ce and Vipul of Together AI
Assertion Supported
Patel: Nvidia manufactured 400k H100s last quarter and will sell 530k this quarter
“There's 400 to 500,000 being, 400,000 manufactured last quarter, and like five 30,000 this quarter being sold, right, of H-Hundreds”
Dylan Patel Dec 5, 2023 ▶ 7:03 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Not checkable as stated
Patel: Nvidia and Google control over 80% of advanced AI silicon manufacturing capacity
“To simplify it, NVIDIA has a little bit more than half, and Google has, like, 30%, right, through Broadcom. So it's like, the total capacity for everyone else is much lower, and they're all sharing it”
Dylan Patel Dec 5, 2023 ▶ 42:01 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Prediction Held up
Patel: Nvidia will sell over 3 million GPUs in 2024
“NVIDIA is going to sell well over three million, you know, total GPUs next year. You know, over a million H 100 this year alone, right?”
Dylan Patel Dec 5, 2023 ▶ 7:46 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Supported
TurboQuant is inefficient on high-bandwidth data center GPUs like NVIDIA B200
“TurboQuant would not be like, it would not be used. Like Nvidia made it clear that this is not a good optimization. And we've seen it firsthand where the overhead of doing dequantization, quantization of You know, in the kernel itself, the turbo-quant kernel, …”
Ali Taha Aug 3, 2026 ▶ 50:02 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Prediction Not checkable as stated
NVIDIA or model creators will always publish open quantized checkpoints
“There's always going to be like an open source quantized checkpoint. NVIDIA is going to push one out if no one else does. You usually the providers will have their own spec tech that they've trained as well. You don't need to train your own spec tech. You can …”
Ali Taha Aug 3, 2026 ▶ 41:22 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Opinion
NVIDIA Dynamo is a developer toolkit, not an out-of-the-box performance optimizer
“I would think of Dynamo as less of a sort of out of box system and more of a toolkit for building with. So when we talk about doing KV aware routing, when we talk about doing KV out offloading, when we talk about doing PD disaggregation, Dynamo fundamentally i…”
Philip Kiely Aug 3, 2026 ▶ 42:20 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Not checkable as stated
NVIDIA's Rubin is the first GPU architecture designed entirely for modern LLMs
“Ruben's honestly the first chip that was fully built in that world. And so you can see a lot of the understanding of the shape of the workload that this chip's going to be asked to do in the way it's designed.”
Philip Kiely Aug 3, 2026 ▶ 1:06:16 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Assertion Open · timeframe Jun 2029
Midha: MatX chips adopt NVIDIA reference architecture to plug into existing sites
“When they decided to pick the standard for their data center, they picked the NVIDIA reference architecture. So the Matex chips just plug in to any site that has an NVIDIA bring up planned. And you know.”
Anjney Midha Jun 18, 2026 ▶ 31:28 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Assertion Supported
Ethan He: Megatron MoE was first to train trillion-parameter MoEs at 40% MFU
“The Megatron MOEs was the first It was the first framework open source to be able to train these MOEs at very large scales, like a hundred billion parameters to even trillion parameters efficiently at like 40% MFU.”
Ethan He Jun 1, 2026 ▶ 1:41:50 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Insight
Ethan He: Video Foundation Models Follow Scaling Laws Like LLMs
“There, once I built the Cosmos one, I realized as this thing also has a scaling law similar to language model.”
Ethan He Jun 1, 2026 ▶ 3:11 Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.