Dec 18, 2025 · 48m · catalyst

Will inference move to the edge?

Dr. Ben Lee · 24m spoken Shayle Kann · 16m spoken
0:00 / 0:00

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Catalyst, host Shail Khan and computer science professor Dr. Ben Lee examine whether AI inference workloads will shift from centralized hyperscale data centers to distributed edge infrastructure and explore the profound technical, latency, and power grid implications of this transition.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Shayle holds 37.3% of the talking time here. How this is scored →

Shayle as informed peer 4.5 Guest teaching 4.7 Guest disagreement 0.1 Shayle pushing back 0.4
05100:0015:0030:0045:000:00–2:48 · Shayle as informed peer 0/10 Listener Survey Announcement and Gift Card Promotion Introductory promotional announcement and sponsor ads read by a voiceover. No dialogue occurs between the host and guest.2:54–5:13 · Shayle as informed peer 0/10 Framing the AI Compute and Grid Demand Dilemma Solo host monologue framing the central thesis of the episode regarding AI compute growth, power grid bottlenecks, and edge inference potential.5:14–9:46 · Shayle as informed peer 5/10 The Three Compute Tiers and Cloud Efficiency The host asks the guest to define compute categories and recalls historical AV edge computing discussions, while the guest clearly explains hyperscale PUE metrics and hardware-sharing efficiencies.9:47–12:49 · Shayle as informed peer 4/10 Model Training vs. Inference Energy Workloads The host asks whether training compute will ever decentralize, and the guest details Meta research showing the three-way energy split between preprocessing, training, and inference.12:50–16:21 · Shayle as informed peer 5/10 Edge Rationale, Latency Tolerances, and Cyber-Physical AI The host notes how users tolerate latency in reasoning models like Deep Research and asks about robotics, which the guest categorizes under cyber-physical AI requiring strict responsiveness guarantees.16:22–20:00 · Shayle as informed peer 6/10 Hardware Interconnects, Power Spikes, and DIDT Challenges The host displays solid technical insight by pointing out the bizarre practice of dummy workloads to mitigate power spikes, which the guest validates and terms the DIDT challenge.20:03–26:53 · Shayle as informed peer 6/10 Sponsor Break: Bloom Energy, Engie, and Energy Hub Following the sponsor read, the host poses a detailed siting thought experiment comparing one 1-GW data center site against one hundred 10-MW sites, which the guest enthusiastically endorses.26:54–29:19 · Shayle as informed peer 5/10 Geographic Clustering, Reliability, and Network Redundancy The host queries why historical clustering occurred in regions like Northern Virginia, and the guest elaborates on internet exchange points, tax incentives, and workload rollover redundancy.29:19–33:52 · Shayle as informed peer 5/10 CDN Precedents and Market Drivers for Edge AI The host presses on why small edge inference sites are not yet being actively built, prompting the guest to draw parallels to Content Delivery Networks and Points of Presence.33:54–40:24 · Shayle as informed peer 5/10 On-Device Inference: Privacy, Parameters, and Thermal Limits The host brings up Apple as an obvious driver for on-device inference, and the guest breaks down the severe memory, parameter reduction, and battery/thermal constraints on consumer devices.40:24–44:12 · Shayle as informed peer 6/10 The 2035 Prediction: The 80/20 Compute Distribution Rule The host presses the guest for a concrete 2035 projection; the guest introduces the 80/20 rule, which the host immediately drills into to pin down the exact edge vs. device split.44:13–48:10 · Shayle as informed peer 7/10 Systemic Energy Impacts and Autonomous Agentic Workloads The host insightfully deduces that decentralized edge inference will actually increase total grid power consumption due to degraded PUE, which the guest confirms before exploring autonomous agent workloads.0:00–2:48 · Guest teaching 0/10 Listener Survey Announcement and Gift Card Promotion Introductory promotional announcement and sponsor ads read by a voiceover. No dialogue occurs between the host and guest.2:54–5:13 · Guest teaching 0/10 Framing the AI Compute and Grid Demand Dilemma Solo host monologue framing the central thesis of the episode regarding AI compute growth, power grid bottlenecks, and edge inference potential.5:14–9:46 · Guest teaching 6/10 The Three Compute Tiers and Cloud Efficiency The host asks the guest to define compute categories and recalls historical AV edge computing discussions, while the guest clearly explains hyperscale PUE metrics and hardware-sharing efficiencies.9:47–12:49 · Guest teaching 6/10 Model Training vs. Inference Energy Workloads The host asks whether training compute will ever decentralize, and the guest details Meta research showing the three-way energy split between preprocessing, training, and inference.12:50–16:21 · Guest teaching 5/10 Edge Rationale, Latency Tolerances, and Cyber-Physical AI The host notes how users tolerate latency in reasoning models like Deep Research and asks about robotics, which the guest categorizes under cyber-physical AI requiring strict responsiveness guarantees.16:22–20:00 · Guest teaching 6/10 Hardware Interconnects, Power Spikes, and DIDT Challenges The host displays solid technical insight by pointing out the bizarre practice of dummy workloads to mitigate power spikes, which the guest validates and terms the DIDT challenge.20:03–26:53 · Guest teaching 4/10 Sponsor Break: Bloom Energy, Engie, and Energy Hub Following the sponsor read, the host poses a detailed siting thought experiment comparing one 1-GW data center site against one hundred 10-MW sites, which the guest enthusiastically endorses.26:54–29:19 · Guest teaching 5/10 Geographic Clustering, Reliability, and Network Redundancy The host queries why historical clustering occurred in regions like Northern Virginia, and the guest elaborates on internet exchange points, tax incentives, and workload rollover redundancy.29:19–33:52 · Guest teaching 6/10 CDN Precedents and Market Drivers for Edge AI The host presses on why small edge inference sites are not yet being actively built, prompting the guest to draw parallels to Content Delivery Networks and Points of Presence.33:54–40:24 · Guest teaching 6/10 On-Device Inference: Privacy, Parameters, and Thermal Limits The host brings up Apple as an obvious driver for on-device inference, and the guest breaks down the severe memory, parameter reduction, and battery/thermal constraints on consumer devices.40:24–44:12 · Guest teaching 6/10 The 2035 Prediction: The 80/20 Compute Distribution Rule The host presses the guest for a concrete 2035 projection; the guest introduces the 80/20 rule, which the host immediately drills into to pin down the exact edge vs. device split.44:13–48:10 · Guest teaching 6/10 Systemic Energy Impacts and Autonomous Agentic Workloads The host insightfully deduces that decentralized edge inference will actually increase total grid power consumption due to degraded PUE, which the guest confirms before exploring autonomous agent workloads.0:00–2:48 · Guest disagreement 0/10 Listener Survey Announcement and Gift Card Promotion Introductory promotional announcement and sponsor ads read by a voiceover. No dialogue occurs between the host and guest.2:54–5:13 · Guest disagreement 0/10 Framing the AI Compute and Grid Demand Dilemma Solo host monologue framing the central thesis of the episode regarding AI compute growth, power grid bottlenecks, and edge inference potential.5:14–9:46 · Guest disagreement 0/10 The Three Compute Tiers and Cloud Efficiency The host asks the guest to define compute categories and recalls historical AV edge computing discussions, while the guest clearly explains hyperscale PUE metrics and hardware-sharing efficiencies.9:47–12:49 · Guest disagreement 0/10 Model Training vs. Inference Energy Workloads The host asks whether training compute will ever decentralize, and the guest details Meta research showing the three-way energy split between preprocessing, training, and inference.12:50–16:21 · Guest disagreement 0/10 Edge Rationale, Latency Tolerances, and Cyber-Physical AI The host notes how users tolerate latency in reasoning models like Deep Research and asks about robotics, which the guest categorizes under cyber-physical AI requiring strict responsiveness guarantees.16:22–20:00 · Guest disagreement 0/10 Hardware Interconnects, Power Spikes, and DIDT Challenges The host displays solid technical insight by pointing out the bizarre practice of dummy workloads to mitigate power spikes, which the guest validates and terms the DIDT challenge.20:03–26:53 · Guest disagreement 0/10 Sponsor Break: Bloom Energy, Engie, and Energy Hub Following the sponsor read, the host poses a detailed siting thought experiment comparing one 1-GW data center site against one hundred 10-MW sites, which the guest enthusiastically endorses.26:54–29:19 · Guest disagreement 0/10 Geographic Clustering, Reliability, and Network Redundancy The host queries why historical clustering occurred in regions like Northern Virginia, and the guest elaborates on internet exchange points, tax incentives, and workload rollover redundancy.29:19–33:52 · Guest disagreement 1/10 CDN Precedents and Market Drivers for Edge AI The host presses on why small edge inference sites are not yet being actively built, prompting the guest to draw parallels to Content Delivery Networks and Points of Presence.33:54–40:24 · Guest disagreement 0/10 On-Device Inference: Privacy, Parameters, and Thermal Limits The host brings up Apple as an obvious driver for on-device inference, and the guest breaks down the severe memory, parameter reduction, and battery/thermal constraints on consumer devices.40:24–44:12 · Guest disagreement 0/10 The 2035 Prediction: The 80/20 Compute Distribution Rule The host presses the guest for a concrete 2035 projection; the guest introduces the 80/20 rule, which the host immediately drills into to pin down the exact edge vs. device split.44:13–48:10 · Guest disagreement 0/10 Systemic Energy Impacts and Autonomous Agentic Workloads The host insightfully deduces that decentralized edge inference will actually increase total grid power consumption due to degraded PUE, which the guest confirms before exploring autonomous agent workloads.0:00–2:48 · Shayle pushing back 0/10 Listener Survey Announcement and Gift Card Promotion Introductory promotional announcement and sponsor ads read by a voiceover. No dialogue occurs between the host and guest.2:54–5:13 · Shayle pushing back 0/10 Framing the AI Compute and Grid Demand Dilemma Solo host monologue framing the central thesis of the episode regarding AI compute growth, power grid bottlenecks, and edge inference potential.5:14–9:46 · Shayle pushing back 0/10 The Three Compute Tiers and Cloud Efficiency The host asks the guest to define compute categories and recalls historical AV edge computing discussions, while the guest clearly explains hyperscale PUE metrics and hardware-sharing efficiencies.9:47–12:49 · Shayle pushing back 0/10 Model Training vs. Inference Energy Workloads The host asks whether training compute will ever decentralize, and the guest details Meta research showing the three-way energy split between preprocessing, training, and inference.12:50–16:21 · Shayle pushing back 0/10 Edge Rationale, Latency Tolerances, and Cyber-Physical AI The host notes how users tolerate latency in reasoning models like Deep Research and asks about robotics, which the guest categorizes under cyber-physical AI requiring strict responsiveness guarantees.16:22–20:00 · Shayle pushing back 1/10 Hardware Interconnects, Power Spikes, and DIDT Challenges The host displays solid technical insight by pointing out the bizarre practice of dummy workloads to mitigate power spikes, which the guest validates and terms the DIDT challenge.20:03–26:53 · Shayle pushing back 1/10 Sponsor Break: Bloom Energy, Engie, and Energy Hub Following the sponsor read, the host poses a detailed siting thought experiment comparing one 1-GW data center site against one hundred 10-MW sites, which the guest enthusiastically endorses.26:54–29:19 · Shayle pushing back 0/10 Geographic Clustering, Reliability, and Network Redundancy The host queries why historical clustering occurred in regions like Northern Virginia, and the guest elaborates on internet exchange points, tax incentives, and workload rollover redundancy.29:19–33:52 · Shayle pushing back 1/10 CDN Precedents and Market Drivers for Edge AI The host presses on why small edge inference sites are not yet being actively built, prompting the guest to draw parallels to Content Delivery Networks and Points of Presence.33:54–40:24 · Shayle pushing back 0/10 On-Device Inference: Privacy, Parameters, and Thermal Limits The host brings up Apple as an obvious driver for on-device inference, and the guest breaks down the severe memory, parameter reduction, and battery/thermal constraints on consumer devices.40:24–44:12 · Shayle pushing back 2/10 The 2035 Prediction: The 80/20 Compute Distribution Rule The host presses the guest for a concrete 2035 projection; the guest introduces the 80/20 rule, which the host immediately drills into to pin down the exact edge vs. device split.44:13–48:10 · Shayle pushing back 0/10 Systemic Energy Impacts and Autonomous Agentic Workloads The host insightfully deduces that decentralized edge inference will actually increase total grid power consumption due to degraded PUE, which the guest confirms before exploring autonomous agent workloads.

speaking balance: gold is Shayle, purple is the guest (3 minute bins)

0:00 · Shayle 7% · guest 93%0:00 · Shayle 7% · guest 93%3:00 · Shayle 95.1% · guest 4.9%3:00 · Shayle 95.1% · guest 4.9%6:00 · Shayle 28.9% · guest 71.1%6:00 · Shayle 28.9% · guest 71.1%9:00 · Shayle 39.9% · guest 60.1%9:00 · Shayle 39.9% · guest 60.1%12:00 · Shayle 34.9% · guest 65.1%12:00 · Shayle 34.9% · guest 65.1%15:00 · Shayle 44.4% · guest 55.6%15:00 · Shayle 44.4% · guest 55.6%18:00 · Shayle 24.1% · guest 75.9%18:00 · Shayle 24.1% · guest 75.9%21:00 · Shayle 34% · guest 66%21:00 · Shayle 34% · guest 66%24:00 · Shayle 36.5% · guest 63.5%24:00 · Shayle 36.5% · guest 63.5%27:00 · Shayle 56.7% · guest 43.3%27:00 · Shayle 56.7% · guest 43.3%30:00 · Shayle 23.5% · guest 76.5%30:00 · Shayle 23.5% · guest 76.5%33:00 · Shayle 37.4% · guest 62.6%33:00 · Shayle 37.4% · guest 62.6%36:00 · Shayle 21.1% · guest 78.9%36:00 · Shayle 21.1% · guest 78.9%39:00 · Shayle 32.2% · guest 67.8%39:00 · Shayle 32.2% · guest 67.8%42:00 · Shayle 47.6% · guest 52.4%42:00 · Shayle 47.6% · guest 52.4%45:00 · Shayle 23.3% · guest 76.7%45:00 · Shayle 23.3% · guest 76.7%48:00 · Shayle 67.5% · guest 32.5%48:00 · Shayle 67.5% · guest 32.5%
Sharpest disagreement ▶ 32:37 Guest offers contrarian view on pipeline GPU saturation

The guest pushes back on assumptions of guaranteed edge construction by highlighting potential diminishing returns in model scaling and the likelihood of repurposing existing hyperscale capacity.

Hardest push from Shayle ▶ 42:10 Host demands clarification on the 80/20 local breakdown

The host refuses to accept the broad categorization of 'local compute' and actively pushes the guest to delineate between regional edge data centers and on-device hardware.

Biggest teaching moment ▶ 36:22 Guest explains parameter reduction and thermal barriers

The guest walks through the severe technical physics trade-offs of shrinking a 1-trillion parameter model down to a 7-billion parameter mobile model, detailing context memory and thermal constraints.

Shayle holds their own ▶ 44:50 Host connects edge distribution to aggregate energy increases

The host demonstrates sharp analytical command of energy systems by pointing out that edge data center adoption will sacrifice hyperscale PUE efficiency and raise aggregate electricity consumption.

the scores for every segment, with the reasoning behind each
ChapterTopicShayle as informed peerGuest teachingGuest disagreementShayle pushing backWhy
Listener Survey Announcement and Gift Card Promotion 0000 Introductory promotional announcement and sponsor ads read by a voiceover. No dialogue occurs between the host and guest.
Framing the AI Compute and Grid Demand Dilemma 0000 Solo host monologue framing the central thesis of the episode regarding AI compute growth, power grid bottlenecks, and edge inference potential.
The Three Compute Tiers and Cloud Efficiency 5600 The host asks the guest to define compute categories and recalls historical AV edge computing discussions, while the guest clearly explains hyperscale PUE metrics and hardware-sharing efficiencies.
Model Training vs. Inference Energy Workloads 4600 The host asks whether training compute will ever decentralize, and the guest details Meta research showing the three-way energy split between preprocessing, training, and inference.
Edge Rationale, Latency Tolerances, and Cyber-Physical AI 5500 The host notes how users tolerate latency in reasoning models like Deep Research and asks about robotics, which the guest categorizes under cyber-physical AI requiring strict responsiveness guarantees.
Hardware Interconnects, Power Spikes, and DIDT Challenges 6601 The host displays solid technical insight by pointing out the bizarre practice of dummy workloads to mitigate power spikes, which the guest validates and terms the DIDT challenge.
Sponsor Break: Bloom Energy, Engie, and Energy Hub 6401 Following the sponsor read, the host poses a detailed siting thought experiment comparing one 1-GW data center site against one hundred 10-MW sites, which the guest enthusiastically endorses.
Geographic Clustering, Reliability, and Network Redundancy 5500 The host queries why historical clustering occurred in regions like Northern Virginia, and the guest elaborates on internet exchange points, tax incentives, and workload rollover redundancy.
CDN Precedents and Market Drivers for Edge AI 5611 The host presses on why small edge inference sites are not yet being actively built, prompting the guest to draw parallels to Content Delivery Networks and Points of Presence.
On-Device Inference: Privacy, Parameters, and Thermal Limits 5600 The host brings up Apple as an obvious driver for on-device inference, and the guest breaks down the severe memory, parameter reduction, and battery/thermal constraints on consumer devices.
The 2035 Prediction: The 80/20 Compute Distribution Rule 6602 The host presses the guest for a concrete 2035 projection; the guest introduces the 80/20 rule, which the host immediately drills into to pin down the exact edge vs. device split.
Systemic Energy Impacts and Autonomous Agentic Workloads 7600 The host insightfully deduces that decentralized edge inference will actually increase total grid power consumption due to degraded PUE, which the guest confirms before exploring autonomous agent workloads.

Statements from this episode (16)

Assertion Partly supported
Effectively 100% of global AI compute currently occurs in centralized cloud
“It's an energy question because the answer today is effectively a hundred percent in that first category, cloud.”
Shayle Kann Dec 18, 2025 ▶ 3:18
Assertion Supported
Google achieves a data center power usage effectiveness near 1.1
“Google's PUE is close to 1.1, which is to say for every watt going to compute, there's an additional .1 watts going to the overheads of power delivery or cooling or whatever.”
Dr. Ben Lee Dec 18, 2025 ▶ 8:59
Assertion Supported
Meta study: AI energy costs split evenly across pre-processing, training, and inference
“There was a study we did when I was a visiting research scientist at Meta where we Found that energy costs for AI were roughly broken into three categories. There's a data pre-processing aspect as well, and that's about a third. The training is another third, …”
Dr. Ben Lee Dec 18, 2025 ▶ 11:52
Prediction Not checkable as stated
AI inference spending must surge to justify current adoption optimism
“If the optimism about AI is to be justified, you're going to have to see inference costs go way up because that will be an indicator that adoption has gone up in a fairly significant way, both among individual users, but also among companies and enterprise use…”
Dr. Ben Lee Dec 18, 2025 ▶ 12:26
Insight
Generative AI is reconditioning users to tolerate multi-second latency delays
“What is interesting with generative AI is that we are being reconditioned to tolerate much longer delays. So if you use something like GPT or you use something like Claude or your favorite chatbot, oftentimes it's just sitting there thinking for seconds and se…”
Dr. Ben Lee Dec 18, 2025 ▶ 14:19
Prediction Not checkable as stated
Cyber-physical AI applications will require edge computing for low latency
“So I agree that there will be cases where we will need those really low latencies, and that is going to require edge computing much closer to the user, so we have much shorter internet delays, network delays.”
Dr. Ben Lee Dec 18, 2025 ▶ 16:08
Insight
AI training compute cycles cause massive power fluctuations unlike inference workloads
“And some of the people in the energy space may know that there are massive energy fluctuations or power fluctuations we will see in data center usage when the GPUs go from this computational intensive phase where you're learning the model weights to this commu…”
Dr. Ben Lee Dec 18, 2025 ▶ 18:09
Assertion Supported
AI training data centers run dummy workloads to flatten dangerous power spikes
“What they do in large part, because those spikes are actually problematic, not just to the grid, but to the equipment inside the data center as well. So what they do at least sometimes to manage that is they create dummy workloads. So they keep the power profi…”
Shayle Kann Dec 18, 2025 ▶ 19:00
Insight
LLM inference prompts are processed locally on one to eight GPUs
“When you send a prompt to for processing by a large language model, that prompt is probably handled by one GPU or maybe Eight GPUs inside a single machine. So, and the reason that is, is because the model sits in that machine, the data sits in that machine, an…”
Dr. Ben Lee Dec 18, 2025 ▶ 22:17
Prediction Not checkable as stated
Siting distributed edge compute will become easier than 1GW data centers
“I think finding capacity there may eventually become easier than finding the next thousand megawatts.”
Dr. Ben Lee Dec 18, 2025 ▶ 25:40
Prediction Not checkable as stated
If AI scaling slows, repurposed central GPUs will cannibalize edge data centers
“And if it turns out that maybe there are diminishing returns from training larger and larger models, or maybe we run out of data because we've exhausted all the data that's available on the internet. When those things happen, it may be that demand for these GP…”
Dr. Ben Lee Dec 18, 2025 ▶ 32:52
Prediction Not checkable as stated
Competition on AI latency will force local GPU deployments in major markets
“I think the catch there will be if one of these model providers or one of these application developers makes performance a distinguishing feature of their offering, right? If they start competing on performance rather than on capability, then we're going to se…”
Dr. Ben Lee Dec 18, 2025 ▶ 33:26
Prediction Not checkable as stated
Apple is the most likely company to move AI inference on-device
“It's not hard to picture that, like, if somebody's gonna move a lot of this inference on device, it's gonna be Apple.”
Shayle Kann Dec 18, 2025 ▶ 35:46
Prediction Open · timeframe Dec 2035
80% of AI inference compute will move to the edge by 2035
“So I would say that we could be getting 80% of our compute done locally and leaving 20% of the heavy lifting or the more esoteric, the more corner case compute for the data center cloud. That is, of course, excluding the training. The training will Continue to…”
Dr. Ben Lee Dec 18, 2025 ▶ 41:54
Insight
Shifting inference to edge data centers will reduce efficiency and increase energy costs
“I think as you shrink the system down, you will get, you will lose an efficiency. You will be trying to build these 20 megawatt data centers and maybe footprints or facilities that weren't designed initially for those workloads. So yes, I think total energy co…”
Dr. Ben Lee Dec 18, 2025 ▶ 45:48
Prediction Open · timeframe Dec 2030
Software agents, rather than human queries, will drive most AI inference workloads
“I think increasingly most of the inference workload will come from other software agents.”
Dr. Ben Lee Dec 18, 2025 ▶ 46:40
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.