Jan 22, 2026 · 1h 4m · mad

The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)

Dan Fu · 28m spoken Tim Dettmers · 24m spoken Matt Turck · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

This episode of The MAD Podcast features host Matt Turck in conversation with AI researchers Tim Dettmers and Dan Fu as they debate whether GPU scaling has hit physical hardware limits or if computational efficiency and AI agents will drive the next era of AGI development. Together, they explore hardware bottlenecks, agent management strategies, post-training workflows, and industry predictions heading into 2026.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 10.8% of the talking time here. How this is scored →

Matt as informed peer 2.8 Guest teaching 5.1 Guest disagreement 0.6 Matt pushing back 0.3
05100:0015:0030:0045:001:00:001:04–3:29 · Matt as informed peer 1/10 Guest Introductions and Dual Academic-Industry Backgrounds Matt welcomes guests Tim Dettmers and Dan Fu and prompts them to detail their dual academic-industry backgrounds. The guests explain specialized topics like model quantization and GPU kernel optimization in an approachable manner.3:29–7:16 · Matt as informed peer 2/10 Defining AGI and Evaluating Economic Productivity Matt asks both guests for a working definition of AGI. Dan and Tim gently reframe AGI away from sci-fi tropes toward tangible economic productivity and industrial impact.7:16–16:13 · Matt as informed peer 3/10 Tim Dettmers on Physical Limits and GPU Bottlenecks Tim critiques superficial AGI timelines from rationalist communities and delivers a dense technical overview of physical compute limits, geometric memory constraints, and 4-bit quantization ceilings. Matt follows along closely by quoting key statements from Tim's blog post.16:13–22:47 · Matt as informed peer 3/10 Dan Fu on Compute Growth and Lagging Model Capabilities Dan directly counters Tim's physical limit argument by pointing out low hardware utilization (20% MFU) and upcoming 100x compute gains from Blackwell clusters. Matt highlights Dan's core insight that current AI models are lagging indicators of hardware capabilities.22:47–29:48 · Matt as informed peer 5/10 Post-Training Workflows and Real-World AI Utility Matt introduces post-training workflows and skillfully synthesizes both guests' opposing essays around practical economic utility rather than abstract AGI definitions. Both guests agree, expanding on technological diffusion and self-driving autonomy inflection points.29:48–32:16 · Matt as informed peer 4/10 Multi-Chip Architectures, Custom ASICs, and Local Inference Matt demonstrates industry knowledge by bringing up alternative AI chip makers like Groq and Cerebras. Dan details low-level software abstraction challenges across AMD versus NVIDIA and distinct demands of training versus inference.32:16–39:11 · Matt as informed peer 4/10 The Agent Era and the Coding Singularity Matt prompts the guests on whether AI agents have reached an inflection point, citing Tim's writing. Dan shares how Cursor agents conquered complex C++ GPU kernel writing, while Tim explains why code execution serves as a universal interface for digital tasks.39:16–43:52 · Matt as informed peer 2/10 Non-Coders and Building Tools with AI Agents Matt turns to practical advice for non-programmers seeking to automate daily tasks. Tim shares a story about writing a custom video-slicing script in 20 minutes and outlines an industrial automation methodology for evaluating true task ROI.43:52–48:04 · Matt as informed peer 2/10 Managing AI Agents Like Junior Engineers Matt asks Dan for management principles when using AI agents in technical workflows. Dan compares agent interaction to onboarding junior interns and explains why domain expertise dramatically increases agent leverage.48:04–52:28 · Matt as informed peer 3/10 Onboarding Junior Engineers in the Agent Era Matt asks a sharp question about how junior engineers develop core domain expertise when entry-level tasks are automated. Dan shares Together AI's training approach, while Tim details the computer science education paradox.52:28–55:53 · Matt as informed peer 1/10 Current AI Projects at Ai2 and Together AI Matt asks both guests to share their current research projects. Tim announces an upcoming Ai2 release enabling local 32B models to adapt to private codebases at 100x lower synthetic data cost, while Dan discusses Together AI's work on inference efficiency.55:53–58:07 · Matt as informed peer 3/10 Deep Dive into Mega Kernels and Together Atlas Matt specifically asks Dan to explain Mega Kernels and Together Atlas. Dan explains how compiling an entire neural network into a single GPU kernel yields 2-3x speedups and how adaptive speculative decoding optimizes model response times over time.58:07–1:02:02 · Matt as informed peer 2/10 Predictions and Expectations for AI Progress Through 2026 Matt asks for specific 2026 AI predictions. Tim foresees frontier model stagnation alongside the rise of specialized 100B local models, while Dan highlights new hardware generations like NVIDIA Rubin and rapid multimodal progress.1:02:02–1:03:35 · Matt as informed peer 4/10 Post-Transformer Architectures and Alternative Models Matt demonstrates solid technical understanding by asking about alternative architectures like state-space models and JEPA. Dan explains how hybrid architectures and Chinese research labs are expanding model diversity beyond standard transformers.1:04–3:29 · Guest teaching 3/10 Guest Introductions and Dual Academic-Industry Backgrounds Matt welcomes guests Tim Dettmers and Dan Fu and prompts them to detail their dual academic-industry backgrounds. The guests explain specialized topics like model quantization and GPU kernel optimization in an approachable manner.3:29–7:16 · Guest teaching 4/10 Defining AGI and Evaluating Economic Productivity Matt asks both guests for a working definition of AGI. Dan and Tim gently reframe AGI away from sci-fi tropes toward tangible economic productivity and industrial impact.7:16–16:13 · Guest teaching 7/10 Tim Dettmers on Physical Limits and GPU Bottlenecks Tim critiques superficial AGI timelines from rationalist communities and delivers a dense technical overview of physical compute limits, geometric memory constraints, and 4-bit quantization ceilings. Matt follows along closely by quoting key statements from Tim's blog post.16:13–22:47 · Guest teaching 6/10 Dan Fu on Compute Growth and Lagging Model Capabilities Dan directly counters Tim's physical limit argument by pointing out low hardware utilization (20% MFU) and upcoming 100x compute gains from Blackwell clusters. Matt highlights Dan's core insight that current AI models are lagging indicators of hardware capabilities.22:47–29:48 · Guest teaching 4/10 Post-Training Workflows and Real-World AI Utility Matt introduces post-training workflows and skillfully synthesizes both guests' opposing essays around practical economic utility rather than abstract AGI definitions. Both guests agree, expanding on technological diffusion and self-driving autonomy inflection points.29:48–32:16 · Guest teaching 5/10 Multi-Chip Architectures, Custom ASICs, and Local Inference Matt demonstrates industry knowledge by bringing up alternative AI chip makers like Groq and Cerebras. Dan details low-level software abstraction challenges across AMD versus NVIDIA and distinct demands of training versus inference.32:16–39:11 · Guest teaching 5/10 The Agent Era and the Coding Singularity Matt prompts the guests on whether AI agents have reached an inflection point, citing Tim's writing. Dan shares how Cursor agents conquered complex C++ GPU kernel writing, while Tim explains why code execution serves as a universal interface for digital tasks.39:16–43:52 · Guest teaching 5/10 Non-Coders and Building Tools with AI Agents Matt turns to practical advice for non-programmers seeking to automate daily tasks. Tim shares a story about writing a custom video-slicing script in 20 minutes and outlines an industrial automation methodology for evaluating true task ROI.43:52–48:04 · Guest teaching 5/10 Managing AI Agents Like Junior Engineers Matt asks Dan for management principles when using AI agents in technical workflows. Dan compares agent interaction to onboarding junior interns and explains why domain expertise dramatically increases agent leverage.48:04–52:28 · Guest teaching 6/10 Onboarding Junior Engineers in the Agent Era Matt asks a sharp question about how junior engineers develop core domain expertise when entry-level tasks are automated. Dan shares Together AI's training approach, while Tim details the computer science education paradox.52:28–55:53 · Guest teaching 5/10 Current AI Projects at Ai2 and Together AI Matt asks both guests to share their current research projects. Tim announces an upcoming Ai2 release enabling local 32B models to adapt to private codebases at 100x lower synthetic data cost, while Dan discusses Together AI's work on inference efficiency.55:53–58:07 · Guest teaching 6/10 Deep Dive into Mega Kernels and Together Atlas Matt specifically asks Dan to explain Mega Kernels and Together Atlas. Dan explains how compiling an entire neural network into a single GPU kernel yields 2-3x speedups and how adaptive speculative decoding optimizes model response times over time.58:07–1:02:02 · Guest teaching 5/10 Predictions and Expectations for AI Progress Through 2026 Matt asks for specific 2026 AI predictions. Tim foresees frontier model stagnation alongside the rise of specialized 100B local models, while Dan highlights new hardware generations like NVIDIA Rubin and rapid multimodal progress.1:02:02–1:03:35 · Guest teaching 5/10 Post-Transformer Architectures and Alternative Models Matt demonstrates solid technical understanding by asking about alternative architectures like state-space models and JEPA. Dan explains how hybrid architectures and Chinese research labs are expanding model diversity beyond standard transformers.1:04–3:29 · Guest disagreement 0/10 Guest Introductions and Dual Academic-Industry Backgrounds Matt welcomes guests Tim Dettmers and Dan Fu and prompts them to detail their dual academic-industry backgrounds. The guests explain specialized topics like model quantization and GPU kernel optimization in an approachable manner.3:29–7:16 · Guest disagreement 1/10 Defining AGI and Evaluating Economic Productivity Matt asks both guests for a working definition of AGI. Dan and Tim gently reframe AGI away from sci-fi tropes toward tangible economic productivity and industrial impact.7:16–16:13 · Guest disagreement 2/10 Tim Dettmers on Physical Limits and GPU Bottlenecks Tim critiques superficial AGI timelines from rationalist communities and delivers a dense technical overview of physical compute limits, geometric memory constraints, and 4-bit quantization ceilings. Matt follows along closely by quoting key statements from Tim's blog post.16:13–22:47 · Guest disagreement 3/10 Dan Fu on Compute Growth and Lagging Model Capabilities Dan directly counters Tim's physical limit argument by pointing out low hardware utilization (20% MFU) and upcoming 100x compute gains from Blackwell clusters. Matt highlights Dan's core insight that current AI models are lagging indicators of hardware capabilities.22:47–29:48 · Guest disagreement 1/10 Post-Training Workflows and Real-World AI Utility Matt introduces post-training workflows and skillfully synthesizes both guests' opposing essays around practical economic utility rather than abstract AGI definitions. Both guests agree, expanding on technological diffusion and self-driving autonomy inflection points.29:48–32:16 · Guest disagreement 0/10 Multi-Chip Architectures, Custom ASICs, and Local Inference Matt demonstrates industry knowledge by bringing up alternative AI chip makers like Groq and Cerebras. Dan details low-level software abstraction challenges across AMD versus NVIDIA and distinct demands of training versus inference.32:16–39:11 · Guest disagreement 1/10 The Agent Era and the Coding Singularity Matt prompts the guests on whether AI agents have reached an inflection point, citing Tim's writing. Dan shares how Cursor agents conquered complex C++ GPU kernel writing, while Tim explains why code execution serves as a universal interface for digital tasks.39:16–43:52 · Guest disagreement 0/10 Non-Coders and Building Tools with AI Agents Matt turns to practical advice for non-programmers seeking to automate daily tasks. Tim shares a story about writing a custom video-slicing script in 20 minutes and outlines an industrial automation methodology for evaluating true task ROI.43:52–48:04 · Guest disagreement 0/10 Managing AI Agents Like Junior Engineers Matt asks Dan for management principles when using AI agents in technical workflows. Dan compares agent interaction to onboarding junior interns and explains why domain expertise dramatically increases agent leverage.48:04–52:28 · Guest disagreement 1/10 Onboarding Junior Engineers in the Agent Era Matt asks a sharp question about how junior engineers develop core domain expertise when entry-level tasks are automated. Dan shares Together AI's training approach, while Tim details the computer science education paradox.52:28–55:53 · Guest disagreement 0/10 Current AI Projects at Ai2 and Together AI Matt asks both guests to share their current research projects. Tim announces an upcoming Ai2 release enabling local 32B models to adapt to private codebases at 100x lower synthetic data cost, while Dan discusses Together AI's work on inference efficiency.55:53–58:07 · Guest disagreement 0/10 Deep Dive into Mega Kernels and Together Atlas Matt specifically asks Dan to explain Mega Kernels and Together Atlas. Dan explains how compiling an entire neural network into a single GPU kernel yields 2-3x speedups and how adaptive speculative decoding optimizes model response times over time.58:07–1:02:02 · Guest disagreement 0/10 Predictions and Expectations for AI Progress Through 2026 Matt asks for specific 2026 AI predictions. Tim foresees frontier model stagnation alongside the rise of specialized 100B local models, while Dan highlights new hardware generations like NVIDIA Rubin and rapid multimodal progress.1:02:02–1:03:35 · Guest disagreement 0/10 Post-Transformer Architectures and Alternative Models Matt demonstrates solid technical understanding by asking about alternative architectures like state-space models and JEPA. Dan explains how hybrid architectures and Chinese research labs are expanding model diversity beyond standard transformers.1:04–3:29 · Matt pushing back 0/10 Guest Introductions and Dual Academic-Industry Backgrounds Matt welcomes guests Tim Dettmers and Dan Fu and prompts them to detail their dual academic-industry backgrounds. The guests explain specialized topics like model quantization and GPU kernel optimization in an approachable manner.3:29–7:16 · Matt pushing back 0/10 Defining AGI and Evaluating Economic Productivity Matt asks both guests for a working definition of AGI. Dan and Tim gently reframe AGI away from sci-fi tropes toward tangible economic productivity and industrial impact.7:16–16:13 · Matt pushing back 1/10 Tim Dettmers on Physical Limits and GPU Bottlenecks Tim critiques superficial AGI timelines from rationalist communities and delivers a dense technical overview of physical compute limits, geometric memory constraints, and 4-bit quantization ceilings. Matt follows along closely by quoting key statements from Tim's blog post.16:13–22:47 · Matt pushing back 2/10 Dan Fu on Compute Growth and Lagging Model Capabilities Dan directly counters Tim's physical limit argument by pointing out low hardware utilization (20% MFU) and upcoming 100x compute gains from Blackwell clusters. Matt highlights Dan's core insight that current AI models are lagging indicators of hardware capabilities.22:47–29:48 · Matt pushing back 1/10 Post-Training Workflows and Real-World AI Utility Matt introduces post-training workflows and skillfully synthesizes both guests' opposing essays around practical economic utility rather than abstract AGI definitions. Both guests agree, expanding on technological diffusion and self-driving autonomy inflection points.29:48–32:16 · Matt pushing back 0/10 Multi-Chip Architectures, Custom ASICs, and Local Inference Matt demonstrates industry knowledge by bringing up alternative AI chip makers like Groq and Cerebras. Dan details low-level software abstraction challenges across AMD versus NVIDIA and distinct demands of training versus inference.32:16–39:11 · Matt pushing back 0/10 The Agent Era and the Coding Singularity Matt prompts the guests on whether AI agents have reached an inflection point, citing Tim's writing. Dan shares how Cursor agents conquered complex C++ GPU kernel writing, while Tim explains why code execution serves as a universal interface for digital tasks.39:16–43:52 · Matt pushing back 0/10 Non-Coders and Building Tools with AI Agents Matt turns to practical advice for non-programmers seeking to automate daily tasks. Tim shares a story about writing a custom video-slicing script in 20 minutes and outlines an industrial automation methodology for evaluating true task ROI.43:52–48:04 · Matt pushing back 0/10 Managing AI Agents Like Junior Engineers Matt asks Dan for management principles when using AI agents in technical workflows. Dan compares agent interaction to onboarding junior interns and explains why domain expertise dramatically increases agent leverage.48:04–52:28 · Matt pushing back 0/10 Onboarding Junior Engineers in the Agent Era Matt asks a sharp question about how junior engineers develop core domain expertise when entry-level tasks are automated. Dan shares Together AI's training approach, while Tim details the computer science education paradox.52:28–55:53 · Matt pushing back 0/10 Current AI Projects at Ai2 and Together AI Matt asks both guests to share their current research projects. Tim announces an upcoming Ai2 release enabling local 32B models to adapt to private codebases at 100x lower synthetic data cost, while Dan discusses Together AI's work on inference efficiency.55:53–58:07 · Matt pushing back 0/10 Deep Dive into Mega Kernels and Together Atlas Matt specifically asks Dan to explain Mega Kernels and Together Atlas. Dan explains how compiling an entire neural network into a single GPU kernel yields 2-3x speedups and how adaptive speculative decoding optimizes model response times over time.58:07–1:02:02 · Matt pushing back 0/10 Predictions and Expectations for AI Progress Through 2026 Matt asks for specific 2026 AI predictions. Tim foresees frontier model stagnation alongside the rise of specialized 100B local models, while Dan highlights new hardware generations like NVIDIA Rubin and rapid multimodal progress.1:02:02–1:03:35 · Matt pushing back 0/10 Post-Transformer Architectures and Alternative Models Matt demonstrates solid technical understanding by asking about alternative architectures like state-space models and JEPA. Dan explains how hybrid architectures and Chinese research labs are expanding model diversity beyond standard transformers.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 35.8% · guest 64.2%0:00 · Matt 35.8% · guest 64.2%3:00 · Matt 14.1% · guest 85.9%3:00 · Matt 14.1% · guest 85.9%6:00 · Matt 13.2% · guest 86.8%6:00 · Matt 13.2% · guest 86.8%9:00 · Matt 7.3% · guest 92.7%9:00 · Matt 7.3% · guest 92.7%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 2.1% · guest 97.9%15:00 · Matt 2.1% · guest 97.9%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 17% · guest 83%21:00 · Matt 17% · guest 83%24:00 · Matt 20% · guest 80%24:00 · Matt 20% · guest 80%27:00 · Matt 6.6% · guest 93.4%27:00 · Matt 6.6% · guest 93.4%30:00 · Matt 14.3% · guest 85.7%30:00 · Matt 14.3% · guest 85.7%33:00 · Matt 13.3% · guest 86.7%33:00 · Matt 13.3% · guest 86.7%36:00 · Matt 10.6% · guest 89.4%36:00 · Matt 10.6% · guest 89.4%39:00 · Matt 10.7% · guest 89.3%39:00 · Matt 10.7% · guest 89.3%42:00 · Matt 4.8% · guest 95.2%42:00 · Matt 4.8% · guest 95.2%45:00 · Matt 12.2% · guest 87.8%45:00 · Matt 12.2% · guest 87.8%48:00 · Matt 2.6% · guest 97.4%48:00 · Matt 2.6% · guest 97.4%51:00 · Matt 8.9% · guest 91.1%51:00 · Matt 8.9% · guest 91.1%54:00 · Matt 3.1% · guest 96.9%54:00 · Matt 3.1% · guest 96.9%57:00 · Matt 6.6% · guest 93.4%57:00 · Matt 6.6% · guest 93.4%1:00:00 · Matt 12.4% · guest 87.6%1:00:00 · Matt 12.4% · guest 87.6%1:03:00 · Matt 42% · guest 58%1:03:00 · Matt 42% · guest 58%
Sharpest disagreement ▶ 7:31 Tim rejects EA and rationalist AGI timelines

Tim explicitly criticizes public AGI predictions coming from effective altruism and rationalist circles as lazy thinking from living in an isolated bubble.

Hardest push from Matt ▶ 11:27 Matt directly challenges Tim on the end of GPU scaling

Matt refuses to let a bold blog quote slide, directly confronting Tim with his statement that GPUs will no longer improve meaningfully and demanding a technical explanation.

Biggest teaching moment ▶ 14:20 Tim explains geometric memory bottlenecks and four-bit quantization limits

Tim delivers an in-depth physics lecture explaining geometric DRAM access patterns, Von Neumann bottlenecks, and information-theoretic limits showing four-bit quantization is the end of precision scaling.

Matt holds his own ▶ 25:30 Matt unifies opposing guest essays around economic usefulness

Matt displays strong high-level synthesis by connecting the core points of both guests' opposing essays, showing how economic usefulness is the true common ground between their positions.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Guest Introductions and Dual Academic-Industry Backgrounds 1300 Matt welcomes guests Tim Dettmers and Dan Fu and prompts them to detail their dual academic-industry backgrounds. The guests explain specialized topics like model quantization and GPU kernel optimization in an approachable manner.
Defining AGI and Evaluating Economic Productivity 2410 Matt asks both guests for a working definition of AGI. Dan and Tim gently reframe AGI away from sci-fi tropes toward tangible economic productivity and industrial impact.
Tim Dettmers on Physical Limits and GPU Bottlenecks 3721 Tim critiques superficial AGI timelines from rationalist communities and delivers a dense technical overview of physical compute limits, geometric memory constraints, and 4-bit quantization ceilings. Matt follows along closely by quoting key statements from Tim's blog post.
Dan Fu on Compute Growth and Lagging Model Capabilities 3632 Dan directly counters Tim's physical limit argument by pointing out low hardware utilization (20% MFU) and upcoming 100x compute gains from Blackwell clusters. Matt highlights Dan's core insight that current AI models are lagging indicators of hardware capabilities.
Post-Training Workflows and Real-World AI Utility 5411 Matt introduces post-training workflows and skillfully synthesizes both guests' opposing essays around practical economic utility rather than abstract AGI definitions. Both guests agree, expanding on technological diffusion and self-driving autonomy inflection points.
Multi-Chip Architectures, Custom ASICs, and Local Inference 4500 Matt demonstrates industry knowledge by bringing up alternative AI chip makers like Groq and Cerebras. Dan details low-level software abstraction challenges across AMD versus NVIDIA and distinct demands of training versus inference.
The Agent Era and the Coding Singularity 4510 Matt prompts the guests on whether AI agents have reached an inflection point, citing Tim's writing. Dan shares how Cursor agents conquered complex C++ GPU kernel writing, while Tim explains why code execution serves as a universal interface for digital tasks.
Non-Coders and Building Tools with AI Agents 2500 Matt turns to practical advice for non-programmers seeking to automate daily tasks. Tim shares a story about writing a custom video-slicing script in 20 minutes and outlines an industrial automation methodology for evaluating true task ROI.
Managing AI Agents Like Junior Engineers 2500 Matt asks Dan for management principles when using AI agents in technical workflows. Dan compares agent interaction to onboarding junior interns and explains why domain expertise dramatically increases agent leverage.
Onboarding Junior Engineers in the Agent Era 3610 Matt asks a sharp question about how junior engineers develop core domain expertise when entry-level tasks are automated. Dan shares Together AI's training approach, while Tim details the computer science education paradox.
Current AI Projects at Ai2 and Together AI 1500 Matt asks both guests to share their current research projects. Tim announces an upcoming Ai2 release enabling local 32B models to adapt to private codebases at 100x lower synthetic data cost, while Dan discusses Together AI's work on inference efficiency.
Deep Dive into Mega Kernels and Together Atlas 3600 Matt specifically asks Dan to explain Mega Kernels and Together Atlas. Dan explains how compiling an entire neural network into a single GPU kernel yields 2-3x speedups and how adaptive speculative decoding optimizes model response times over time.
Predictions and Expectations for AI Progress Through 2026 2500 Matt asks for specific 2026 AI predictions. Tim foresees frontier model stagnation alongside the rise of specialized 100B local models, while Dan highlights new hardware generations like NVIDIA Rubin and rapid multimodal progress.
Post-Transformer Architectures and Alternative Models 4500 Matt demonstrates solid technical understanding by asking about alternative architectures like state-space models and JEPA. Dan explains how hybrid architectures and Chinese research labs are expanding model diversity beyond standard transformers.

Statements from this episode (22)

Opinion
Fu: Current LLMs meet the definition of AGI from 5-10 years ago
“By almost any definition anyone could have written down, let's say five years ago or 10 years ago, certainly when, you know, Tim, you and I started our PhD. We basically have the vision of AGI that, that we had back then. We have things that can write code. Th…”
Dan Fu Jan 22, 2026 ▶ 4:14
Prediction Not checkable as stated
Fu: Next-generation models currently in training will achieve AGI
“You know, we maybe already have AGI or like some form of AGI. And if not, then certainly the next generation of models, the models that today are training already. If they're at all better than what we have today, then we're, we we've already hit something tha…”
Dan Fu Jan 22, 2026 ▶ 5:17
Insight
Dettmers: AGI economic adoption may temporarily dip productivity before boosting it
“And similar to when computers were introduced, productivity increased. Not initially, productivity actually went down. And you need to do diffusion in the economy to pick it up again. We might see something like that with AGI more broadly, and starts with soft…”
Tim Dettmers Jan 22, 2026 ▶ 6:57
Assertion Contradicted
Dettmers: AI hardware has maxed out and won't get faster
“The hardware is maxed out. We have no new technology. We can make it easier to manufacture and a little bit cheaper, but not faster. And we have maxed out on the additional features.”
Tim Dettmers Jan 22, 2026 ▶ 15:46
Assertion Not checkable as stated
Dettmers: 4-bit precision is the end of quantization
“Four bit precision is the end of quantization.”
Tim Dettmers Jan 22, 2026 ▶ 16:03
Assertion Partly supported
Dan Fu: DeepSeek-V3 was trained on ~2,000 H800s with 20% MFU
“If you look at the deep seek model, for instance, this is one of the best open source models we have out there today. It was trained at the end of 2024. On last generation, kind of nerfed GPUs, H 800 instead of H 100, the 800 is nerfed by all sorts of ways fro…”
Dan Fu Jan 22, 2026 ▶ 17:26
Assertion Supported
Poolside and Reflection are building clusters with massive B200 GPU deployments
“They're companies like Poolside. They're building out tens of thousands of B-two hundred, GB-two hundred chips. You know, there's other folks like Reflection who are who are building out. Tens of thousands of B 200 chips.”
Dan Fu Jan 22, 2026 ▶ 19:13
Insight
Dan Fu: Deployed AI models lag cluster infrastructure by 1–2 years
“The models that we see today that we can play with today have been pre-trained on clusters that were built out a year or two ago. Because, you know, you need enough time to get the cluster running. You need enough time to do the large pre-training run. And the…”
Dan Fu Jan 22, 2026 ▶ 21:26
Prediction Not checkable as stated
Fu predicts increasing hardware diversity, particularly for AI model inference
“I'm sure NVIDIA will still do great and still grow beyond their five trillion dollar company or whatever it is at the time of recording. But I think you're going to see a lot more diversity especially around, I think inference of the model.”
Dan Fu Jan 22, 2026 ▶ 31:39
Assertion Not checkable as stated
Fu: AI coding tools enable expert programmers to move 10x faster
“But if you give an expert programmer This set of tools, they can go 10, 10 times faster than they were able to go before.”
Dan Fu Jan 22, 2026 ▶ 34:41
Insight
Dettmers: Coding agents serve as general-purpose AI agents for digital tasks
“Coding agents are general agents. Coding agents can write programs that solve other problems, and code is so general, if there's a digital problem, you could solve it for code, and coding agents make the thing so easy that now you can solve a variety of proble…”
Tim Dettmers Jan 22, 2026 ▶ 36:19
Insight
Dettmers: Automating calendar scheduling with AI agents yields minimal productivity gain
“An agent doesn't know that. And if you specify that to an agent, you could also just create the calendar invite, the meeting pipeline, and it doesn't increase productivity by much.”
Tim Dettmers Jan 22, 2026 ▶ 43:34
Insight
Dan Fu: Junior engineers using AI agents communicate better and level up faster
“When they are really gung ho about understanding and being able to use the AI agents, there's, they're able to communicate so much better than in the olden days. They're able to level up their level of understanding a lot faster.”
Dan Fu Jan 22, 2026 ▶ 49:34
Insight
Dettmers: Students using AI agents perform poorly on basic domain knowledge
“If we let people use agents, they perform very poorly on basic knowledge. And if we let people just do the basic knowledge, they don't know how to use agents and they can't compete. So they can't do useful work in the workforce nowadays.”
Tim Dettmers Jan 22, 2026 ▶ 51:22
Prediction Not checkable as stated
Dettmers: Humans will work on problems only AI agents understand
“In the future it's realistic that we work on problems that we don't understand, that agents understand, but we need to keep up in some way”
Tim Dettmers Jan 22, 2026 ▶ 52:07
Disclosure
Ai2 to release open-source coding agent with 100x cheaper training
“We will have a major release of an open source coding agent that has a couple of key features. For one, training is a hundred times cheaper. You need to generate synthetic data and you need to train on it. And so we have a method that's a hundred roughly a hun…”
Tim Dettmers Jan 22, 2026 ▶ 52:57
Assertion Not checkable as stated
Fu: Hardware utilization during AI inference is under 5%
“At inference time, when the, when you have the model, when it's already been trained, already been post-trained, the hardware utilization is like less than five percent.”
Dan Fu Jan 22, 2026 ▶ 55:13
Insight
Dan Fu: Proper speculative decoding yields 2x to 3x model speedups
“So if you do the speculative decoding right, you can get, again, two X, three X speed ups over, over, you know, just running a vanilla model.”
Dan Fu Jan 22, 2026 ▶ 57:34
Assertion Contradicted
Dettmers: No open-source system supports complex inference across eight-plus GPUs
“There's no open source system that can do that at the moment.”
Tim Dettmers Jan 22, 2026 ▶ 1:00:13
Prediction Not checkable as stated
Dettmers: Frontier AI performance will stagnate while smaller models improve
“Performance on the frontier will stagnate, but on the smaller level, we get more and more powerful models still, because you can distill from these large models into these small models.”
Tim Dettmers Jan 22, 2026 ▶ 1:00:38
Assertion Not checkable as stated
Dan Fu: Some top audio models use state space architectures
“So some of the best audio models in the world are at least partially based on state space models.”
Dan Fu Jan 22, 2026 ▶ 1:02:29
Assertion Not checkable as stated
Dan Fu: Chinese AI labs take more architectural risks
“I think you see a lot more risk taking out of the Chinese labs where you're trying to differentiate the next model of your next open source model.”
Dan Fu Jan 22, 2026 ▶ 1:03:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.