Dan Fu, VP of Kernels at Together AI, outlines the rapid deployment of NVIDIA Blackwell hardware clusters by frontier AI startups.
Opinion
Fu: Current LLMs meet the definition of AGI from 5-10 years ago
“By almost any definition anyone could have written down, let's say five years ago or 10 years ago, certainly when, you know, Tim, you and I started our PhD. We basically have the vision of AGI that, that we had back then. We have things that can write code. Th…”
Prediction Not checkable as stated
Fu: Next-generation models currently in training will achieve AGI
“You know, we maybe already have AGI or like some form of AGI. And if not, then certainly the next generation of models, the models that today are training already. If they're at all better than what we have today, then we're, we we've already hit something tha…”
Assertion Not checkable as stated
Fu: AI coding tools enable expert programmers to move 10x faster
“But if you give an expert programmer This set of tools, they can go 10, 10 times faster than they were able to go before.”
Assertion Partly supported
Dan Fu: DeepSeek-V3 was trained on ~2,000 H800s with 20% MFU
“If you look at the deep seek model, for instance, this is one of the best open source models we have out there today. It was trained at the end of 2024. On last generation, kind of nerfed GPUs, H 800 instead of H 100, the 800 is nerfed by all sorts of ways fro…”
Insight
Dan Fu: Deployed AI models lag cluster infrastructure by 1–2 years
“The models that we see today that we can play with today have been pre-trained on clusters that were built out a year or two ago. Because, you know, you need enough time to get the cluster running. You need enough time to do the large pre-training run. And the…”
Assertion Not checkable as stated
Fu: Hardware utilization during AI inference is under 5%
“At inference time, when the, when you have the model, when it's already been trained, already been post-trained, the hardware utilization is like less than five percent.”