Insight certainty 4/5 debate potential 2/5

Dan Fu: Deployed AI models lag cluster infrastructure by 1–2 years

Dan Fu · The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI) · Jan 22, 2026 · at 21:26

Dan Fu explains why visible AI model performance lags behind current GPU hardware capabilities.

0:00 / 0:19exact quote · 19.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The models that we see today that we can play with today have been pre-trained on clusters that were built out a year or two ago. Because, you know, you need enough time to get the cluster running. You need enough time to do the large pre-training run. And then you need enough time to really post train it, fine tune it, do all the RLHF and all that stuff.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dan Fu

Opinion
Fu: Current LLMs meet the definition of AGI from 5-10 years ago
“By almost any definition anyone could have written down, let's say five years ago or 10 years ago, certainly when, you know, Tim, you and I started our PhD. We basically have the vision of AGI that, that we had back then. We have things that can write code. Th…”
Dan Fu Jan 22, 2026 ▶ 4:14 The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
Prediction Not checkable as stated
Fu: Next-generation models currently in training will achieve AGI
“You know, we maybe already have AGI or like some form of AGI. And if not, then certainly the next generation of models, the models that today are training already. If they're at all better than what we have today, then we're, we we've already hit something tha…”
Dan Fu Jan 22, 2026 ▶ 5:17 The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
Assertion Not checkable as stated
Fu: AI coding tools enable expert programmers to move 10x faster
“But if you give an expert programmer This set of tools, they can go 10, 10 times faster than they were able to go before.”
Dan Fu Jan 22, 2026 ▶ 34:41 The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
Assertion Partly supported
Dan Fu: DeepSeek-V3 was trained on ~2,000 H800s with 20% MFU
“If you look at the deep seek model, for instance, this is one of the best open source models we have out there today. It was trained at the end of 2024. On last generation, kind of nerfed GPUs, H 800 instead of H 100, the 800 is nerfed by all sorts of ways fro…”
Dan Fu Jan 22, 2026 ▶ 17:26 The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
Assertion Not checkable as stated
Fu: Hardware utilization during AI inference is under 5%
“At inference time, when the, when you have the model, when it's already been trained, already been post-trained, the hardware utilization is like less than five percent.”
Dan Fu Jan 22, 2026 ▶ 55:13 The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
Assertion Not checkable as stated
Dan Fu: Chinese AI labs take more architectural risks
“I think you see a lot more risk taking out of the Chinese labs where you're trying to differentiate the next model of your next open source model.”
Dan Fu Jan 22, 2026 ▶ 1:03:19 The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.