Dan Fu: DeepSeek-V3 was trained on ~2,000 H800s with 20% MFU
Dan Fu · The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI) · Jan 22, 2026 · at 17:26
Dan Fu, Assistant Professor at UCSD and VP of Kernels at Together AI, discusses hardware utilization efficiency during the training of open-source frontier models.
“If you look at the deep seek model, for instance, this is one of the best open source models we have out there today. It was trained at the end of 2024. On last generation, kind of nerfed GPUs, H 800 instead of H 100, the 800 is nerfed by all sorts of ways from NVIDIA to get around the expert restrictions at the time. And they were trained with, let's call it like about 2000 H 800 according to the report for, I think about a month. And when you compute how much, how long that took, when you see how much compute was actually available on the chip, you get something like a 20% effective chip utilization or something like that.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →