Aug 28, 2025 · 37m · catalyst

The mechanics of data center flexibility

Varun Sivaram · 21m spoken Shayle Kann · 9m spoken
0:00 / 0:00

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Host Shayle Kann and Emerald AI CEO Varun Sivaram examine how software-driven workload orchestration and silicon-level throttling can transform AI data centers into flexible grid assets. They discuss the evolution of commercial service level agreements and empirical field trials designed to resolve grid interconnection bottlenecks and accelerate clean energy integration.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Shayle holds 26.9% of the talking time here. How this is scored →

Shayle as informed peer 6.0 Guest teaching 4.6 Guest disagreement 0.9 Shayle pushing back 1.9
05100:0010:0020:0030:003:53–8:02 · Shayle as informed peer 6/10 Understanding AI Load Profiles and Grid Interconnection Planning Shayle frames the grid planning paradigm around 8,760 hours of peak capacity. Varun extends the explanation by pointing out that interconnection studies actually plan across ten-year horizons (87,600 hours) under compounding worst-case reliability assumptions.8:03–12:56 · Shayle as informed peer 7/10 Operational Headroom and the Physics of AI Power Density Shayle accurately summarizes how grid operators face a dilemma when unpredictable spiky loads enter the system. Varun respectfully counters Shayle's earlier comment that AI loads are not dissimilar to traditional loads, emphasizing unprecedented power density jumps from 5 kW to 132 kW per rack.13:00–15:05 · Shayle as informed peer 5/10 Comparing Training and Inference Workloads over Data Center Lifecycles Shayle asks whether training and inference create distinct grid signatures. Varun validates the question, describing the spiky, checkpoint-driven profile of training runs versus smoothed inference patterns, noting that data centers repurpose hardware across their lifecycles.15:05–17:32 · Shayle as informed peer 5/10 Mechanisms of Workload Flexibility and Computational Orchestration Shayle prompts Varun to explain the physical and computational mechanics behind demand response in AI facilities. Varun details how spatial and temporal flexibility can be harvested directly through workload orchestration.17:35–23:53 · Shayle as informed peer 6/10 Mid-roll Sponsor Break: Bloom Energy, Engie, and Energy Hub Following mid-roll ads, Shayle digs into temporal load shifting, asking whether it is as simple as delaying jobs. Varun unpacks the dual optimization problem, explaining finer interventions like dynamic auto-scaling and hardware clock frequency throttling.23:54–28:01 · Shayle as informed peer 6/10 Rethinking SLAs, Grid Headroom, and Long-Term Load Shifting Shayle contextualizes historic hyperscaler reluctance to break 24/7 uptime SLAs due to customer commitments. Varun explains that massive grid interconnection bottlenecks and new flexible SLA structures are forcing customer adaptation.28:02–32:58 · Shayle as informed peer 7/10 Empirical Proof: Phoenix Field Test and PowerFlex SLA Feasibility Shayle probes realistic demand reduction limits and insists that everything ultimately hinges on customer SLA contracts. Varun agrees and shares empirical findings from their Phoenix pilot with Oracle and EPRI demonstrating 25% to 40% load reductions.33:00–36:39 · Shayle as informed peer 6/10 Overcoming Utility Skepticism with Digital Twins and Commercial Demonstrations Shayle asks what concrete proof utilities will need to trust software-driven load shedding during interconnection studies. Varun details commercial pilot scaling, digital twin simulation modeling, and governor-level economic pressures pushing utilities to find solutions.3:53–8:02 · Guest teaching 5/10 Understanding AI Load Profiles and Grid Interconnection Planning Shayle frames the grid planning paradigm around 8,760 hours of peak capacity. Varun extends the explanation by pointing out that interconnection studies actually plan across ten-year horizons (87,600 hours) under compounding worst-case reliability assumptions.8:03–12:56 · Guest teaching 6/10 Operational Headroom and the Physics of AI Power Density Shayle accurately summarizes how grid operators face a dilemma when unpredictable spiky loads enter the system. Varun respectfully counters Shayle's earlier comment that AI loads are not dissimilar to traditional loads, emphasizing unprecedented power density jumps from 5 kW to 132 kW per rack.13:00–15:05 · Guest teaching 4/10 Comparing Training and Inference Workloads over Data Center Lifecycles Shayle asks whether training and inference create distinct grid signatures. Varun validates the question, describing the spiky, checkpoint-driven profile of training runs versus smoothed inference patterns, noting that data centers repurpose hardware across their lifecycles.15:05–17:32 · Guest teaching 4/10 Mechanisms of Workload Flexibility and Computational Orchestration Shayle prompts Varun to explain the physical and computational mechanics behind demand response in AI facilities. Varun details how spatial and temporal flexibility can be harvested directly through workload orchestration.17:35–23:53 · Guest teaching 5/10 Mid-roll Sponsor Break: Bloom Energy, Engie, and Energy Hub Following mid-roll ads, Shayle digs into temporal load shifting, asking whether it is as simple as delaying jobs. Varun unpacks the dual optimization problem, explaining finer interventions like dynamic auto-scaling and hardware clock frequency throttling.23:54–28:01 · Guest teaching 4/10 Rethinking SLAs, Grid Headroom, and Long-Term Load Shifting Shayle contextualizes historic hyperscaler reluctance to break 24/7 uptime SLAs due to customer commitments. Varun explains that massive grid interconnection bottlenecks and new flexible SLA structures are forcing customer adaptation.28:02–32:58 · Guest teaching 5/10 Empirical Proof: Phoenix Field Test and PowerFlex SLA Feasibility Shayle probes realistic demand reduction limits and insists that everything ultimately hinges on customer SLA contracts. Varun agrees and shares empirical findings from their Phoenix pilot with Oracle and EPRI demonstrating 25% to 40% load reductions.33:00–36:39 · Guest teaching 4/10 Overcoming Utility Skepticism with Digital Twins and Commercial Demonstrations Shayle asks what concrete proof utilities will need to trust software-driven load shedding during interconnection studies. Varun details commercial pilot scaling, digital twin simulation modeling, and governor-level economic pressures pushing utilities to find solutions.3:53–8:02 · Guest disagreement 1/10 Understanding AI Load Profiles and Grid Interconnection Planning Shayle frames the grid planning paradigm around 8,760 hours of peak capacity. Varun extends the explanation by pointing out that interconnection studies actually plan across ten-year horizons (87,600 hours) under compounding worst-case reliability assumptions.8:03–12:56 · Guest disagreement 3/10 Operational Headroom and the Physics of AI Power Density Shayle accurately summarizes how grid operators face a dilemma when unpredictable spiky loads enter the system. Varun respectfully counters Shayle's earlier comment that AI loads are not dissimilar to traditional loads, emphasizing unprecedented power density jumps from 5 kW to 132 kW per rack.13:00–15:05 · Guest disagreement 0/10 Comparing Training and Inference Workloads over Data Center Lifecycles Shayle asks whether training and inference create distinct grid signatures. Varun validates the question, describing the spiky, checkpoint-driven profile of training runs versus smoothed inference patterns, noting that data centers repurpose hardware across their lifecycles.15:05–17:32 · Guest disagreement 0/10 Mechanisms of Workload Flexibility and Computational Orchestration Shayle prompts Varun to explain the physical and computational mechanics behind demand response in AI facilities. Varun details how spatial and temporal flexibility can be harvested directly through workload orchestration.17:35–23:53 · Guest disagreement 1/10 Mid-roll Sponsor Break: Bloom Energy, Engie, and Energy Hub Following mid-roll ads, Shayle digs into temporal load shifting, asking whether it is as simple as delaying jobs. Varun unpacks the dual optimization problem, explaining finer interventions like dynamic auto-scaling and hardware clock frequency throttling.23:54–28:01 · Guest disagreement 1/10 Rethinking SLAs, Grid Headroom, and Long-Term Load Shifting Shayle contextualizes historic hyperscaler reluctance to break 24/7 uptime SLAs due to customer commitments. Varun explains that massive grid interconnection bottlenecks and new flexible SLA structures are forcing customer adaptation.28:02–32:58 · Guest disagreement 1/10 Empirical Proof: Phoenix Field Test and PowerFlex SLA Feasibility Shayle probes realistic demand reduction limits and insists that everything ultimately hinges on customer SLA contracts. Varun agrees and shares empirical findings from their Phoenix pilot with Oracle and EPRI demonstrating 25% to 40% load reductions.33:00–36:39 · Guest disagreement 0/10 Overcoming Utility Skepticism with Digital Twins and Commercial Demonstrations Shayle asks what concrete proof utilities will need to trust software-driven load shedding during interconnection studies. Varun details commercial pilot scaling, digital twin simulation modeling, and governor-level economic pressures pushing utilities to find solutions.3:53–8:02 · Shayle pushing back 2/10 Understanding AI Load Profiles and Grid Interconnection Planning Shayle frames the grid planning paradigm around 8,760 hours of peak capacity. Varun extends the explanation by pointing out that interconnection studies actually plan across ten-year horizons (87,600 hours) under compounding worst-case reliability assumptions.8:03–12:56 · Shayle pushing back 2/10 Operational Headroom and the Physics of AI Power Density Shayle accurately summarizes how grid operators face a dilemma when unpredictable spiky loads enter the system. Varun respectfully counters Shayle's earlier comment that AI loads are not dissimilar to traditional loads, emphasizing unprecedented power density jumps from 5 kW to 132 kW per rack.13:00–15:05 · Shayle pushing back 1/10 Comparing Training and Inference Workloads over Data Center Lifecycles Shayle asks whether training and inference create distinct grid signatures. Varun validates the question, describing the spiky, checkpoint-driven profile of training runs versus smoothed inference patterns, noting that data centers repurpose hardware across their lifecycles.15:05–17:32 · Shayle pushing back 1/10 Mechanisms of Workload Flexibility and Computational Orchestration Shayle prompts Varun to explain the physical and computational mechanics behind demand response in AI facilities. Varun details how spatial and temporal flexibility can be harvested directly through workload orchestration.17:35–23:53 · Shayle pushing back 2/10 Mid-roll Sponsor Break: Bloom Energy, Engie, and Energy Hub Following mid-roll ads, Shayle digs into temporal load shifting, asking whether it is as simple as delaying jobs. Varun unpacks the dual optimization problem, explaining finer interventions like dynamic auto-scaling and hardware clock frequency throttling.23:54–28:01 · Shayle pushing back 2/10 Rethinking SLAs, Grid Headroom, and Long-Term Load Shifting Shayle contextualizes historic hyperscaler reluctance to break 24/7 uptime SLAs due to customer commitments. Varun explains that massive grid interconnection bottlenecks and new flexible SLA structures are forcing customer adaptation.28:02–32:58 · Shayle pushing back 3/10 Empirical Proof: Phoenix Field Test and PowerFlex SLA Feasibility Shayle probes realistic demand reduction limits and insists that everything ultimately hinges on customer SLA contracts. Varun agrees and shares empirical findings from their Phoenix pilot with Oracle and EPRI demonstrating 25% to 40% load reductions.33:00–36:39 · Shayle pushing back 2/10 Overcoming Utility Skepticism with Digital Twins and Commercial Demonstrations Shayle asks what concrete proof utilities will need to trust software-driven load shedding during interconnection studies. Varun details commercial pilot scaling, digital twin simulation modeling, and governor-level economic pressures pushing utilities to find solutions.

speaking balance: gold is Shayle, purple is the guest (3 minute bins)

0:00 · Shayle 28.4% · guest 71.6%0:00 · Shayle 28.4% · guest 71.6%3:00 · Shayle 45.2% · guest 54.8%3:00 · Shayle 45.2% · guest 54.8%6:00 · Shayle 29% · guest 71%6:00 · Shayle 29% · guest 71%9:00 · Shayle 40% · guest 60%9:00 · Shayle 40% · guest 60%12:00 · Shayle 14.1% · guest 85.9%12:00 · Shayle 14.1% · guest 85.9%15:00 · Shayle 27.8% · guest 72.2%15:00 · Shayle 27.8% · guest 72.2%18:00 · Shayle 25.8% · guest 74.2%18:00 · Shayle 25.8% · guest 74.2%21:00 · Shayle 34.6% · guest 65.4%21:00 · Shayle 34.6% · guest 65.4%24:00 · Shayle 9.9% · guest 90.1%24:00 · Shayle 9.9% · guest 90.1%27:00 · Shayle 22.6% · guest 77.4%27:00 · Shayle 22.6% · guest 77.4%30:00 · Shayle 15.5% · guest 84.5%30:00 · Shayle 15.5% · guest 84.5%33:00 · Shayle 25.4% · guest 74.6%33:00 · Shayle 25.4% · guest 74.6%36:00 · Shayle 38.5% · guest 61.5%36:00 · Shayle 38.5% · guest 61.5%
Sharpest disagreement ▶ 11:01 Varun reframes Shayle's premise on load similarity

Varun explicitly pushes back on Shayle's assertion that data center loads are similar to previous large power loads, arguing that rapid exponential growth and power density make AI uniquely disruptive.

Hardest push from Shayle ▶ 30:41 Shayle zeroes in on SLA constraints

Shayle challenges the flexibility demonstration claims by pointing out that load reduction is only meaningful if it complies with strictly defined customer performance SLAs.

Biggest teaching moment ▶ 7:07 Varun corrects interconnection planning time horizons

When Shayle suggests utilities plan for 8,760 hours of peak demand, Varun points out that formal interconnection studies actually model ten-year compound worst-case scenarios spanning 87,600 hours.

Shayle holds their own ▶ 9:49 Shayle articulates the utility planning dilemma

Shayle demonstrates his grasp of utility regulation and grid operations, synthesizing why unpredictable load profiles force conservative grid operators to treat data centers as full nameplate demand.

the scores for every segment, with the reasoning behind each
ChapterTopicShayle as informed peerGuest teachingGuest disagreementShayle pushing backWhy
Understanding AI Load Profiles and Grid Interconnection Planning 6512 Shayle frames the grid planning paradigm around 8,760 hours of peak capacity. Varun extends the explanation by pointing out that interconnection studies actually plan across ten-year horizons (87,600 hours) under compounding worst-case reliability assumptions.
Operational Headroom and the Physics of AI Power Density 7632 Shayle accurately summarizes how grid operators face a dilemma when unpredictable spiky loads enter the system. Varun respectfully counters Shayle's earlier comment that AI loads are not dissimilar to traditional loads, emphasizing unprecedented power density jumps from 5 kW to 132 kW per rack.
Comparing Training and Inference Workloads over Data Center Lifecycles 5401 Shayle asks whether training and inference create distinct grid signatures. Varun validates the question, describing the spiky, checkpoint-driven profile of training runs versus smoothed inference patterns, noting that data centers repurpose hardware across their lifecycles.
Mechanisms of Workload Flexibility and Computational Orchestration 5401 Shayle prompts Varun to explain the physical and computational mechanics behind demand response in AI facilities. Varun details how spatial and temporal flexibility can be harvested directly through workload orchestration.
Mid-roll Sponsor Break: Bloom Energy, Engie, and Energy Hub 6512 Following mid-roll ads, Shayle digs into temporal load shifting, asking whether it is as simple as delaying jobs. Varun unpacks the dual optimization problem, explaining finer interventions like dynamic auto-scaling and hardware clock frequency throttling.
Rethinking SLAs, Grid Headroom, and Long-Term Load Shifting 6412 Shayle contextualizes historic hyperscaler reluctance to break 24/7 uptime SLAs due to customer commitments. Varun explains that massive grid interconnection bottlenecks and new flexible SLA structures are forcing customer adaptation.
Empirical Proof: Phoenix Field Test and PowerFlex SLA Feasibility 7513 Shayle probes realistic demand reduction limits and insists that everything ultimately hinges on customer SLA contracts. Varun agrees and shares empirical findings from their Phoenix pilot with Oracle and EPRI demonstrating 25% to 40% load reductions.
Overcoming Utility Skepticism with Digital Twins and Commercial Demonstrations 6402 Shayle asks what concrete proof utilities will need to trust software-driven load shedding during interconnection studies. Varun details commercial pilot scaling, digital twin simulation modeling, and governor-level economic pressures pushing utilities to find solutions.

Statements from this episode (15)

Assertion Supported
Kann: Data Centers Rarely Operate at Nameplate Peak Capacity
“First, because data centers aren't actually operating at nameplate peak most of the time anyway.”
Shayle Kann Aug 28, 2025 ▶ 2:59
Assertion Supported
Kann: Google partnered with TVA and Michigan Power for demand response
“Google actually made a big announcement about doing this at their data centers just a few weeks ago. They've announced that they've partnered with two utilities, Michigan Power and TVA, to introduce demand response via workload flexibility in their data center…”
Shayle Kann Aug 28, 2025 ▶ 3:11
Assertion Supported
Sivaram: Modern AI data centers convert 90% of power into compute
“Historically, a data center might lose 33% of the power or use it, 33% of that power for non-IT or information technology uses, and the remaining 66 or 67% goes into actual computations. Nowadays, with the increasingly customized design of these AI factories a…”
Varun Sivaram Aug 28, 2025 ▶ 5:14
Assertion Supported
Sivaram: Grid Interconnection Studies Test Data Centers Against 10-Year Worst-Case Peaks
“When you're running a This interconnection study to determine, can this data center connect to my system? You're saying in the next seven or 10 years, in an absolute worst case scenario, so not just 87 60, but 87 60 times 10 87,600 hours when a transmission li…”
Varun Sivaram Aug 28, 2025 ▶ 7:25
Assertion Contradicted
Sivaram: Data center power demand has more than doubled annually
“The power demand from data centers has more than doubled. Every year, the last several years, and that trend shows no sign of abating.”
Varun Sivaram Aug 28, 2025 ▶ 11:24
Assertion Supported
Sivaram: Computing demand is more than quadrupling every year
“Compute demand is more than quadrupling every year, a four X increase every year.”
Varun Sivaram Aug 28, 2025 ▶ 11:53
Prediction Open · timeframe Aug 2030
Sivaram: AI data center racks hit 132 kilowatts, heading toward 1 megawatt
“Today, I just was in a data center in Silicon Valley seeing a brand new deployment of NVIDIA GB 200, the Blackwell generation. The rack is 132 kilowatts. It's liquid cooled, and we're headed toward one megawatt rack.”
Varun Sivaram Aug 28, 2025 ▶ 12:23
Insight
Sivaram: AI data center load profiles drastically change within a year
“A data center will not do a single thing for its lifetime, right? A massive data center, for example, may initially be configured and specified to train a large language model, and then you'll finish training the large language model, and then you'll do other …”
Varun Sivaram Aug 28, 2025 ▶ 14:19
Assertion Supported
Sivaram: Data centers cannot regularly run backup diesel generators under air permits
“If you have a lot of backup generation, you might fire up the backup generation. Often you're not allowed to because your diesel generator will violate its air permit if you use it regularly.”
Varun Sivaram Aug 28, 2025 ▶ 15:56
Opinion
Kann: Spatial load shifting fits hyperscalers, temporal flexibility fits anyone
“That feels to me like it is more readily available to the hyperscalers who have lots and lots of data centers probably within one region than it is to Others. The temporal one, in theory, available to anybody.”
Shayle Kann Aug 28, 2025 ▶ 19:59
Assertion Supported
Sivaram: Google shifted video indexing to nighttime to lower carbon load
“Google also, by the way, has exploited temporal flexibility. There was a Paper, a post they put out a couple years ago, a friend of mine, Varun Mehra, wrote it about moving video indexing operations to nighttime in order to reduce load during periods, as you m…”
Varun Sivaram Aug 28, 2025 ▶ 20:29
Prediction Not checkable as stated
Sivaram: 50-100 GW of AI Data Centers Cannot Be Built Without Flexibility
“We've got 50 to a hundred gigawatts of latent AI demand in the pipeline. It's just not gonna get built unless you have this capability of flexibility.”
Varun Sivaram Aug 28, 2025 ▶ 23:54
Assertion Not checkable as stated
Sivaram: Hundreds of AI Customers Willing to Accept Power Availability Adjustments
“Given the necessity now, I think there's a range of AI customers, and we've talked to hundreds, who are willing to tolerate small levels of changed power availability.”
Varun Sivaram Aug 28, 2025 ▶ 24:53
Prediction Not checkable as stated
Sivaram: AI data centers could consume 25% of US electricity by 2035
“Data centers, which today are about four percent of American energy consumption, AI data centers are about five gigawatts of load, grow to 12% by the end of the decade. AI data centers could be anywhere up to 50 or even more gigawatts, to 25% of American load …”
Varun Sivaram Aug 28, 2025 ▶ 27:10
Assertion Supported
Frankel: Only 10% of Databricks Cluster Workloads Are Non-Preemptible
“It was surprising to me, by the way, to hear that he anticipated that just 10% Of the workloads on a representative Databricks cluster were non-preemptible. In other words, they absolutely could not be paused or delayed in any way.”
Varun Sivaram Aug 28, 2025 ▶ 29:30
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.