why aren't all 21 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Prediction Held up
Prakash predicts up to 5 million AI GPUs will sell in 2024
“There is four to five million GPUs that will be sold this year. NVIDIA and others.”
Assertion Supported
Patel: Nvidia manufactured 400k H100s last quarter and will sell 530k this quarter
“There's 400 to 500,000 being, 400,000 manufactured last quarter, and like five 30,000 this quarter being sold, right, of H-Hundreds”
Prediction Held up
Patel: Nvidia will sell over 3 million GPUs in 2024
“NVIDIA is going to sell well over three million, you know, total GPUs next year. You know, over a million H 100 this year alone, right?”
Assertion Supported
TurboQuant is inefficient on high-bandwidth data center GPUs like NVIDIA B200
“TurboQuant would not be like, it would not be used. Like Nvidia made it clear that this is not a good optimization. And we've seen it firsthand where the overhead of doing dequantization, quantization of You know, in the kernel itself, the turbo-quant kernel, …”
Assertion Supported
Ethan He: Megatron MoE was first to train trillion-parameter MoEs at 40% MFU
“The Megatron MOEs was the first It was the first framework open source to be able to train these MOEs at very large scales, like a hundred billion parameters to even trillion parameters efficiently at like 40% MFU.”
Assertion Supported
Parakhin: Bing Sydney first launched in India using Megatron, not OpenAI
“The funny thing, I mean, the most interesting anecdote is that Sydney was first shipped in India for and it was not noticed for a long time. And first implementation of Sydney didn't even have open AI model under it. It was during Megatron. Microsoft and the N…”
Assertion Supported
NVIDIA announces Rubin CPX as a dedicated prefill-specific hardware accelerator
“And like with our future generations, generations of hardware, we actually announced like with Rubin, this new accelerator that is pre-fill specific. It's called Rubin CPX.”
Assertion Supported
Jensen Huang prioritizes strategic investments in 'zero billion dollar markets'
“Jensen, He says, we're completely happy investing in zero billion dollar markets. We don't care if this creates revenue. It's important for us to know about this market. We think it will be important in the future. It can be zero billion dollars for a while.”
Assertion Supported
AMD's Composable Kernel library offers functionality similar to Nvidia's CUTLASS
“And then on the AMD side, they have this composable kernel library that does something very similar.”
Assertion Supported
Sohmers: NVIDIA's TF32 is actually a 19-bit precision format
“NVIDIA's TF-thirty-two number format is a nineteen-bit number format. They just call it thirty-two-bit.”
Assertion Supported
Ethan He: Hugging Face's sequential GEMM loop for Mixtral is inefficient
“Let's also look at the implementation of Mixtro eight by seven on Hagen-Phys transformer. You will soon notice the, in the expert operation there, You would iterate over all of the experts and compute each of the gem operations one by one. We found that this i…”
Assertion Supported
Fanelli: Singapore accounted for 15% of NVIDIA's Q3 2024 revenue
“Singapore was 15% of NVIDIA's revenue in Q three of 2024.”
Assertion Supported
Hotz: Tinygrad is about 5x slower than PyTorch on Nvidia GPUs
“The correctness for both forwards and backwards passes is there, but on Nvidia, it's about five X slower than PyTorch right now.”
Assertion Supported
NVIDIA Cosmos uses 50,000 to 60,000 tokens for five seconds of video
“Yeah, for example, like in Cosmos, I think just five seconds of video is like a 50, 50 K or a 60 K number of tokens. So like, if you do 50 seconds as a 500 K tokens, if you do longer than that, easily explode.”
Assertion Supported
NVIDIA chip design begins three to five years before market release
“The design process starts like- Exactly. ...three to five years before the chip gets to the market.”
Assertion Supported
NVIDIA Dynamo dynamically sizes and schedules Kubernetes prefill and decode workers
“Dynamo has a set of components that A, tell you how to scale. It tells you how many pre-fill workers and decoded workers it thinks you should have. And also provides a scheduling API for Kubernetes that allows you to actually represent and affect this scheduli…”
Assertion Supported
NVIDIA's 600M Parameter Parakeet Model Tops Speech Transcription Leaderboards
“And then Parakeet is Nvidia's new speech model, speech transcription model. That's number one on the leaderboards. And it's just like very enterprise tuned, like really, really rock solid, reliable at fairly small number of weights, like six hundred million pa…”
Assertion Supported
Ben Allal: NVIDIA generated 1.9 trillion synthetic tokens for Nemotron-CC
“This is a recent paper from NVIDIA, Mnemotron CC. They took things a bit further and they generated not a few billion tokens, but 1.9 trillion tokens, which is huge.”