why aren't all 80 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 2 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Opinion
AI ASIC startups are doomed as NVIDIA GPUs become domain-specialized
“How can you look at this trend and then still be bullish on companies that are coming up with Asics for AI.”
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Opinion
AMD has completely caught up with Nvidia on AI software
“So they caught up on hardware. Now they've caught up on software. Not a lot of people have sort of discovered that they've caught up on software and we're kind of capitalizing on that.”
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Opinion
Lattner: Modular MAX is 'more open source' than vLLM
“This thing's more open source than VLM because VLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?”
Prediction Not checkable as stated
Conrad: Hyperscalers will probably lose significant money reselling Nvidia GPUs
“My intuition is that the hyperscalers are probably going to lose a lot of money, and they know they're going to lose a lot of money on reselling NVIDIA GPUs at least.”
Opinion
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Prediction Not checkable as stated
NVIDIA's Vera Rubin GPU architecture effectively renders mega kernel research obsolete
“The GPU is, is It's designed in such a way that it basically kills megakernels. You don't need to use megakernels that much anymore. So it seems like that entire research field goes into, like, won't be continued, but yeah.”
Prediction Not checkable as stated
Feinberg: Chipmakers will invest in life sciences as LLM alpha shrinks
“I do think that chip makers, including Nvidia, are going to want to get a lot more invested in life sciences because it will always be high in demand. And The amount of alpha left in pure LLM space is just getting a little questionable.”
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Prediction Not checkable as stated
Algorithmic unhobblers could soon expand context windows to 100 million tokens
“I wouldn't be surprised if we do see the ability to, like, break through to, like, ten million, twenty million, a hundred million context through the, an unhobbler showing up.”
Insight
Patel: Nvidia must vastly outperform rivals to overcome vertical integration
“Google, Amazon, they get to vertically integrate, integrate, and vertical integration always saves tons of money. So he has to be better than everyone. By, not just like a little bit, by a ton. To justify his margins. Otherwise, the vertical integration of his…”
Opinion
O'Laughlin: TPU v7 is Google's peak TCO advantage over Nvidia
“I think Ironwood a V seven is the peak gap between on Between TCO, between NVIDIA and TPU, right?”
Prediction Not checkable as stated
O'Laughlin: TPU v8 will not compete as well against Nvidia Rubin
“V-A we just don't think will be as competitive to Ruben, and that's when your special window starts to close.”
Assertion Contradicted
Johnson: Nvidia Blackwell offers roughly same performance per watt as Hopper
“Like, if you look at the numbers, like, even going from Hopper to Blackwell, like, the performance per watt is about the same. They mostly make the number of transistors go up, and they make the chip size go up, and they make the power usage go up. But even fr…”
Opinion
Taskaya: Custom ASICs do not make sense given low Nvidia GEMM overhead
“What is the overhead of an NVIDIA GAM instruction, right? It's like 16%. So like you're essentially buying a, Matrix multiplication machine. So, like, it doesn't really make sense to specialize it that much.”
Prediction Didn’t hold up
Sohmers: NVIDIA Blackwell memory bandwidth efficiency will be lower than Hopper
“All indications are, even though they, you know, more than doubled the theoretical memory bandwidth going from Hopper to Blackwell, the actual percentage of theoretical that you can achieve is, again, going to be less than the previous generation”
Assertion Open · timeframe Aug 2026
Sohmers: Positron hardware achieves 70% higher performance than NVIDIA at lower power
“So, you know, what that actually results in is like today, we're you know, able to achieve about you know, 70% higher performance than NVIDIA with the cards that we're shipping today. Significantly lower power and price point.”
Disclosure
Modular completely eliminates and replaces NVIDIA's CUDA stack
“In the case of Modular, we go literally, like, we only work at that level, because we get rid of all of CUDA, right? And so we've replaced the entire stack, and so we only do that.”
Opinion
AMD GPUs perform horribly on Windows for generative AI tasks
“AMD works horribly on Windows. Like, on Linux, it works fine. It's lower than the price equivalent NVIDIA GPU. But it works, like, you can use it, generate images, everything works. On Linux, on Windows, You might have a hard time”
Assertion Not checkable as stated
Swyx: VC appetite for GPU-rich early-stage startups is completely gone
“The appetite for GPU rich startups, like the, you know, the funding plan is we will raise sixty million and we'll give 50 of that to Nvidia. That is gone, right? Like no one's pitching that. This was literally the plan, the exact plan of like, I can name like …”
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Insight
Ethan He: Peak pre-training learning rate works best for MoE upcycling
“We found that the best is to use the original highest peak learning rate from pre-training, which works the best.”
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Opinion
Chintala: Nvidia's primary competitive moat is NVLink interconnect, not GPU silicon
“The mode that Nvidia has right now, I feel like, is that they're, they have the interconnect that no one else has. Like, AMD GPUs are pretty good. I'm sure there's very silicon that is not bad at all, but, like, the interconnect like, NVLink is uniquely awesom…”
Prediction Partly held up
Patel: Intel will release a chip surpassing Nvidia H100 within a quarter
“Intel bought that company from him, and then shut it down, and bought this other AI company, and now that company is kind of, ah, you know, got new chips. They're gonna release a better chip than the H 100, ah, within the next quarter or so, right?”
Prediction Partly held up
Patel: Nvidia to ship next-gen chip in Q2/Q3 2024 with 3x LLM performance
“Nvidia's releasing a new chip, you know, in, you know, they're gonna announce it in March, and they're gonna release it, you know, and ship it, you know, Q-two, Q-three next year anyways, right? And that chip will probably be three or four times as good. Right…”
Assertion Not checkable as stated
Patel: Microsoft's custom AI chip performs worse than Nvidia H100
“Microsoft's going to announce their chip soon. It's worse performance than the H-H-E-N-H-E-D but the cost effectiveness of it is, is better for Microsoft internally, just because they don't have to pay the Nvidia tax.”
Opinion
Hotz: Google TPUs are the only successful non-Nvidia training chips
“The only company, there's one other company aside from Nvidia who's succeeded at all at making training chips... Mid journey is trained on TPU, right? Like a lot of startups do actually train on TPUs, and they're the only other successful training chip aside f…”
Insight
Lie: 100 to 200 tokens per second is becoming the new batch mode
“What used to be considered fast at, like, 1000 100 or 200 tokens per second is quickly becoming the new batch mode. Right. Quickly becoming the new, like overnight is true.”
Disclosure
Beam: Lila avoids pre-training from scratch, builds on open-weight models
“We have not decided to take on pre-training as well, just because the black magic that you have to do is, is insane, and we've been gifted, you know, something like a billion dollars worth of compute in the form of open-weight models. So we start with an open-…”
Opinion
Swyx: Etched is not disrupting NVIDIA because inference requires dedicated ASICs
“I don't think they want to disrupt NVIDIA. I think the better framing is that all of inference is just so goddamn big that, of course, you're gonna have ASICs for inference.”
Insight
NVIDIA uses 'Speed of Light' framework to define minimum viable goals
“SOL is a term that NVIDIA is used to sort of like instigate a compelling event. You say, this is done. How do we get there? What is the minimum, as much as necessary, as little as possible thing that it takes for us to get exactly here? And it helps you just b…”
Assertion Partly supported
Datology BeyondWeb 3B matches NVIDIA Nemotron 8B in 2.7x less training time
“As you can see that we achieved the same performance as the NVIDIA model in almost, like, 2.7 X, like, lesser time. And then much faster than anything that hugging face or pajama does. Very interestingly, our three B model is pretty much the same performance a…”
Disclosure
Agrawal: Positron cannot beat NVIDIA on matrix-matrix performance per dollar today
“Thomas said that, you know, we are accelerating matrix vector. It doesn't mean that we can't do matrix, matrix. We can, and we can do it fairly well. It's just that the point becomes is like, are you really beating NVIDIA on it from a perf per dollar? And if y…”
Insight
Chintala: Hyperscalers build custom silicon to exploit vertical workload efficiencies
“Each large company has a sufficient enough set of verticalized workloads that have a pattern to them that, say, a more generic accelerator like an NVIDIA or an AMD GPU does not exploit. So there is some level of power efficiency that you're leaving on the tabl…”
Prediction Held up
Prakash predicts up to 5 million AI GPUs will sell in 2024
“There is four to five million GPUs that will be sold this year. NVIDIA and others.”
Assertion Supported
Patel: Nvidia manufactured 400k H100s last quarter and will sell 530k this quarter
“There's 400 to 500,000 being, 400,000 manufactured last quarter, and like five 30,000 this quarter being sold, right, of H-Hundreds”
Assertion Not checkable as stated
Patel: Nvidia and Google control over 80% of advanced AI silicon manufacturing capacity
“To simplify it, NVIDIA has a little bit more than half, and Google has, like, 30%, right, through Broadcom. So it's like, the total capacity for everyone else is much lower, and they're all sharing it”
Prediction Held up
Patel: Nvidia will sell over 3 million GPUs in 2024
“NVIDIA is going to sell well over three million, you know, total GPUs next year. You know, over a million H 100 this year alone, right?”
Assertion Supported
TurboQuant is inefficient on high-bandwidth data center GPUs like NVIDIA B200
“TurboQuant would not be like, it would not be used. Like Nvidia made it clear that this is not a good optimization. And we've seen it firsthand where the overhead of doing dequantization, quantization of You know, in the kernel itself, the turbo-quant kernel, …”
Prediction Not checkable as stated
NVIDIA or model creators will always publish open quantized checkpoints
“There's always going to be like an open source quantized checkpoint. NVIDIA is going to push one out if no one else does. You usually the providers will have their own spec tech that they've trained as well. You don't need to train your own spec tech. You can …”
Opinion
NVIDIA Dynamo is a developer toolkit, not an out-of-the-box performance optimizer
“I would think of Dynamo as less of a sort of out of box system and more of a toolkit for building with. So when we talk about doing KV aware routing, when we talk about doing KV out offloading, when we talk about doing PD disaggregation, Dynamo fundamentally i…”
Assertion Not checkable as stated
NVIDIA's Rubin is the first GPU architecture designed entirely for modern LLMs
“Ruben's honestly the first chip that was fully built in that world. And so you can see a lot of the understanding of the shape of the workload that this chip's going to be asked to do in the way it's designed.”
Assertion Open · timeframe Jun 2029
Midha: MatX chips adopt NVIDIA reference architecture to plug into existing sites
“When they decided to pick the standard for their data center, they picked the NVIDIA reference architecture. So the Matex chips just plug in to any site that has an NVIDIA bring up planned. And you know.”
Assertion Supported
Ethan He: Megatron MoE was first to train trillion-parameter MoEs at 40% MFU
“The Megatron MOEs was the first It was the first framework open source to be able to train these MOEs at very large scales, like a hundred billion parameters to even trillion parameters efficiently at like 40% MFU.”
Insight
Ethan He: Video Foundation Models Follow Scaling Laws Like LLMs
“There, once I built the Cosmos one, I realized as this thing also has a scaling law similar to language model.”