Quantization, speculative decoding, and prefill-decode disaggregation each yield 2x speedups
“Going from BF-sixteen to NVFP four is, it's not quite a two X, right? It's like, I think it's about like 30, 30 to 40% from 16 to eight, and then another 30 to 40% multiplied from eight to four. So that doesn't quite get you a two X, but like roughly a two X. …”
TurboQuant is inefficient on high-bandwidth data center GPUs like NVIDIA B200
“TurboQuant would not be like, it would not be used. Like Nvidia made it clear that this is not a good optimization. And we've seen it firsthand where the overhead of doing dequantization, quantization of You know, in the kernel itself, the turbo-quant kernel, …”
Coogan: ByteDance securing 36,000 Nvidia Blackwell chips via Alawani Cloud
“ByteDance is working with a Southeast Asian company called Alawani Cloud on plans to use some 500 Blackwell computing systems totaling around 36,000 B-two hundred chips.”
O'Laughlin: Nvidia H100 and B200 GPU pricing has firmed up massively
“Like at this point H 100 pricing has massively firmed up. B 200 pricing definitely has super firmed up. And like, hey, there's clearly demand.”
Musk: Tesla Dojo 2 volume production expected by late 2025
“Dojo two is you know, should be, we should have dojo two in volume towards the end of next year. And that, that, that will be, we think, sort of comparable to sort of a B 200 type system, a training system.”