Insight certainty 4/5 debate potential 2/5

LLM inference prompts are processed locally on one to eight GPUs

Dr. Ben Lee · Will inference move to the edge? · Dec 18, 2025 · at 22:17

Computer science professor Dr. Ben Lee explains to Shail Khan why AI inference workloads are technically suited for distributed edge infrastructure rather than massive centralized compute clusters.

0:00 / 0:35exact quote · 35.5s
720p mp4 · rendered on demand · StarZero watermark
“When you send a prompt to for processing by a large language model, that prompt is probably handled by one GPU or maybe Eight GPUs inside a single machine. So, and the reason that is, is because the model sits in that machine, the data sits in that machine, and all of your prior conversations with that bot have, are sitting in that machine. And it's a very localized piece of compute that needs to be done. And you don't need tens or hundreds of GPUs to be coordinating to give you an answer back.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dr. Ben Lee

Prediction Not checkable as stated
Siting distributed edge compute will become easier than 1GW data centers
“I think finding capacity there may eventually become easier than finding the next thousand megawatts.”
Dr. Ben Lee Dec 18, 2025 ▶ 25:40 Will inference move to the edge?
Prediction Open · timeframe Dec 2035
80% of AI inference compute will move to the edge by 2035
“So I would say that we could be getting 80% of our compute done locally and leaving 20% of the heavy lifting or the more esoteric, the more corner case compute for the data center cloud. That is, of course, excluding the training. The training will Continue to…”
Dr. Ben Lee Dec 18, 2025 ▶ 41:54 Will inference move to the edge?
Prediction Not checkable as stated
Cyber-physical AI applications will require edge computing for low latency
“So I agree that there will be cases where we will need those really low latencies, and that is going to require edge computing much closer to the user, so we have much shorter internet delays, network delays.”
Dr. Ben Lee Dec 18, 2025 ▶ 16:08 Will inference move to the edge?
Prediction Not checkable as stated
If AI scaling slows, repurposed central GPUs will cannibalize edge data centers
“And if it turns out that maybe there are diminishing returns from training larger and larger models, or maybe we run out of data because we've exhausted all the data that's available on the internet. When those things happen, it may be that demand for these GP…”
Dr. Ben Lee Dec 18, 2025 ▶ 32:52 Will inference move to the edge?
Insight
Shifting inference to edge data centers will reduce efficiency and increase energy costs
“I think as you shrink the system down, you will get, you will lose an efficiency. You will be trying to build these 20 megawatt data centers and maybe footprints or facilities that weren't designed initially for those workloads. So yes, I think total energy co…”
Dr. Ben Lee Dec 18, 2025 ▶ 45:48 Will inference move to the edge?
Prediction Open · timeframe Dec 2030
Software agents, rather than human queries, will drive most AI inference workloads
“I think increasingly most of the inference workload will come from other software agents.”
Dr. Ben Lee Dec 18, 2025 ▶ 46:40 Will inference move to the edge?
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.