Dr. Ben Lee, a computer science professor, discusses AI inference architectures with host Shail Khan, contrasting traditional search engine latency demands with generative AI.
0:00 / 0:17exact quote · 17.5s
720p mp4 · rendered on demand · StarZero watermark
“What is interesting with generative AI is that we are being reconditioned to tolerate much longer delays. So if you use something like GPT or you use something like Claude or your favorite chatbot, oftentimes it's just sitting there thinking for seconds and seconds, maybe tens of seconds before it gets you the first token.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Dr. Ben Lee
PredictionNot checkable as stated
Siting distributed edge compute will become easier than 1GW data centers
“I think finding capacity there may eventually become easier than finding the next thousand megawatts.”
Dr. Ben LeeDec 18, 2025▶ 25:40Will inference move to the edge?
PredictionOpen · timeframe Dec 2035
80% of AI inference compute will move to the edge by 2035
“So I would say that we could be getting 80% of our compute done locally and leaving 20% of the heavy lifting or the more esoteric, the more corner case compute for the data center cloud. That is, of course, excluding the training. The training will Continue to…”
Dr. Ben LeeDec 18, 2025▶ 41:54Will inference move to the edge?
PredictionNot checkable as stated
Cyber-physical AI applications will require edge computing for low latency
“So I agree that there will be cases where we will need those really low latencies, and that is going to require edge computing much closer to the user, so we have much shorter internet delays, network delays.”
Dr. Ben LeeDec 18, 2025▶ 16:08Will inference move to the edge?
PredictionNot checkable as stated
If AI scaling slows, repurposed central GPUs will cannibalize edge data centers
“And if it turns out that maybe there are diminishing returns from training larger and larger models, or maybe we run out of data because we've exhausted all the data that's available on the internet. When those things happen, it may be that demand for these GP…”
Dr. Ben LeeDec 18, 2025▶ 32:52Will inference move to the edge?
Insight
Shifting inference to edge data centers will reduce efficiency and increase energy costs
“I think as you shrink the system down, you will get, you will lose an efficiency. You will be trying to build these 20 megawatt data centers and maybe footprints or facilities that weren't designed initially for those workloads. So yes, I think total energy co…”
Dr. Ben LeeDec 18, 2025▶ 45:48Will inference move to the edge?
PredictionOpen · timeframe Dec 2030
Software agents, rather than human queries, will drive most AI inference workloads
“I think increasingly most of the inference workload will come from other software agents.”
Dr. Ben LeeDec 18, 2025▶ 46:40Will inference move to the edge?
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.