Oct 29, 2025 · 32m · a16z

Building the Real-World Infrastructure for AI, with Google, Cisco & a16z

Jeetu Patel · 12m spoken Amin Vahdat · 11m spoken Martin Casado · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Hosted by a16z's Martin Casado, Google's Amin Vahdat and Cisco's Jeetu Patel discuss the massive scale, physical constraints, and architectural innovations driving the AI infrastructure explosion. The panel explores chip specialization, distributed networking, internal enterprise adoption, and strategic advice for tech founders.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 4.5 Guest teaching 4.1 Guest disagreement 1.8 The host pushing back 3.0
05100:0010:0020:0030:000:00–3:09 · The host as informed peer 3/10 Highlight Reel: The Scale of the AI Infrastructure Shift The host sets the stage by framing the historical scale of infrastructure cycles and asking guests to compare current demand to past shifts. The guests emphasize the unprecedented speed and scale, comparing it to the internet build-out combined with the Manhattan Project.3:09–5:51 · The host as informed peer 5/10 CapEx Spend Cycles and Power & Supply Constraints The host digs into internal CapEx signals and specifically queries how internal demand curves align with hardware depreciation cycles. Amin explains that 7-8 year old TPUs maintain 100% utilization while physical land and power carry 25-40 year depreciation schedules.5:51–8:08 · The host as informed peer 2/10 Enterprise Readiness and Distributed Data Center Architecture Jeetu outlines the disparity between enterprise readiness and hyperscaler build-outs, detailing how localized power constraints force the creation of scale-across architectures across distances up to 900 km. The host primarily listens to the extended technical breakdown.8:08–15:56 · The host as informed peer 6/10 Re-Inventing Computing: Hardware-Software Co-Design & Specialized Processors The host prompts a debate on whether Nvidia GPUs represent a return to mainframes versus scale-out systems. Amin directly challenges the host's mainframe framing, explaining that workloads dynamically allocate across scale-out software layers.15:56–20:32 · The host as informed peer 5/10 The Next Generation of Scale-Up, Scale-Out, and Scale-Across Networking The host invites discussion on networking evolution, leading Amin to share how power utilities experience massive spikes during bursty network communications. Jeetu adds sharp critique regarding Broadcom's near-monopoly and the necessity for silicon diversity.20:32–23:42 · The host as informed peer 6/10 Inference Systems, RL Constraints, and Intelligence Efficiency The host drills into inference bottlenecks and directly pushes back on Amin's efficiency framing, correcting him that 'intelligence per dollar' is a business model metric rather than a pure hardware capability metric, which Amin accepts.23:42–28:47 · The host as informed peer 4/10 Internal AI Adoption: Code Migrations, Productivity, and Organizational Mindset Amin reveals that migrating Google's codebase off Bigtable was calculated at seven staff millennia, explaining why the effort was abandoned. Jeetu discusses internal tool usage across 25,000 engineers and the necessity of shifting engineer mindsets.28:47–32:21 · The host as informed peer 5/10 Advice for Startups and 12-Month Technological Outlook The host asks for 12-month predictions while explicitly interjecting to ban the generic answer that 'models will get better'. Jeetu advises founders against building thin wrappers, promoting deeper integration and Cisco's technological momentum.0:00–3:09 · Guest teaching 3/10 Highlight Reel: The Scale of the AI Infrastructure Shift The host sets the stage by framing the historical scale of infrastructure cycles and asking guests to compare current demand to past shifts. The guests emphasize the unprecedented speed and scale, comparing it to the internet build-out combined with the Manhattan Project.3:09–5:51 · Guest teaching 4/10 CapEx Spend Cycles and Power & Supply Constraints The host digs into internal CapEx signals and specifically queries how internal demand curves align with hardware depreciation cycles. Amin explains that 7-8 year old TPUs maintain 100% utilization while physical land and power carry 25-40 year depreciation schedules.5:51–8:08 · Guest teaching 5/10 Enterprise Readiness and Distributed Data Center Architecture Jeetu outlines the disparity between enterprise readiness and hyperscaler build-outs, detailing how localized power constraints force the creation of scale-across architectures across distances up to 900 km. The host primarily listens to the extended technical breakdown.8:08–15:56 · Guest teaching 4/10 Re-Inventing Computing: Hardware-Software Co-Design & Specialized Processors The host prompts a debate on whether Nvidia GPUs represent a return to mainframes versus scale-out systems. Amin directly challenges the host's mainframe framing, explaining that workloads dynamically allocate across scale-out software layers.15:56–20:32 · Guest teaching 5/10 The Next Generation of Scale-Up, Scale-Out, and Scale-Across Networking The host invites discussion on networking evolution, leading Amin to share how power utilities experience massive spikes during bursty network communications. Jeetu adds sharp critique regarding Broadcom's near-monopoly and the necessity for silicon diversity.20:32–23:42 · Guest teaching 4/10 Inference Systems, RL Constraints, and Intelligence Efficiency The host drills into inference bottlenecks and directly pushes back on Amin's efficiency framing, correcting him that 'intelligence per dollar' is a business model metric rather than a pure hardware capability metric, which Amin accepts.23:42–28:47 · Guest teaching 5/10 Internal AI Adoption: Code Migrations, Productivity, and Organizational Mindset Amin reveals that migrating Google's codebase off Bigtable was calculated at seven staff millennia, explaining why the effort was abandoned. Jeetu discusses internal tool usage across 25,000 engineers and the necessity of shifting engineer mindsets.28:47–32:21 · Guest teaching 3/10 Advice for Startups and 12-Month Technological Outlook The host asks for 12-month predictions while explicitly interjecting to ban the generic answer that 'models will get better'. Jeetu advises founders against building thin wrappers, promoting deeper integration and Cisco's technological momentum.0:00–3:09 · Guest disagreement 1/10 Highlight Reel: The Scale of the AI Infrastructure Shift The host sets the stage by framing the historical scale of infrastructure cycles and asking guests to compare current demand to past shifts. The guests emphasize the unprecedented speed and scale, comparing it to the internet build-out combined with the Manhattan Project.3:09–5:51 · Guest disagreement 1/10 CapEx Spend Cycles and Power & Supply Constraints The host digs into internal CapEx signals and specifically queries how internal demand curves align with hardware depreciation cycles. Amin explains that 7-8 year old TPUs maintain 100% utilization while physical land and power carry 25-40 year depreciation schedules.5:51–8:08 · Guest disagreement 1/10 Enterprise Readiness and Distributed Data Center Architecture Jeetu outlines the disparity between enterprise readiness and hyperscaler build-outs, detailing how localized power constraints force the creation of scale-across architectures across distances up to 900 km. The host primarily listens to the extended technical breakdown.8:08–15:56 · Guest disagreement 3/10 Re-Inventing Computing: Hardware-Software Co-Design & Specialized Processors The host prompts a debate on whether Nvidia GPUs represent a return to mainframes versus scale-out systems. Amin directly challenges the host's mainframe framing, explaining that workloads dynamically allocate across scale-out software layers.15:56–20:32 · Guest disagreement 3/10 The Next Generation of Scale-Up, Scale-Out, and Scale-Across Networking The host invites discussion on networking evolution, leading Amin to share how power utilities experience massive spikes during bursty network communications. Jeetu adds sharp critique regarding Broadcom's near-monopoly and the necessity for silicon diversity.20:32–23:42 · Guest disagreement 2/10 Inference Systems, RL Constraints, and Intelligence Efficiency The host drills into inference bottlenecks and directly pushes back on Amin's efficiency framing, correcting him that 'intelligence per dollar' is a business model metric rather than a pure hardware capability metric, which Amin accepts.23:42–28:47 · Guest disagreement 1/10 Internal AI Adoption: Code Migrations, Productivity, and Organizational Mindset Amin reveals that migrating Google's codebase off Bigtable was calculated at seven staff millennia, explaining why the effort was abandoned. Jeetu discusses internal tool usage across 25,000 engineers and the necessity of shifting engineer mindsets.28:47–32:21 · Guest disagreement 2/10 Advice for Startups and 12-Month Technological Outlook The host asks for 12-month predictions while explicitly interjecting to ban the generic answer that 'models will get better'. Jeetu advises founders against building thin wrappers, promoting deeper integration and Cisco's technological momentum.0:00–3:09 · The host pushing back 1/10 Highlight Reel: The Scale of the AI Infrastructure Shift The host sets the stage by framing the historical scale of infrastructure cycles and asking guests to compare current demand to past shifts. The guests emphasize the unprecedented speed and scale, comparing it to the internet build-out combined with the Manhattan Project.3:09–5:51 · The host pushing back 3/10 CapEx Spend Cycles and Power & Supply Constraints The host digs into internal CapEx signals and specifically queries how internal demand curves align with hardware depreciation cycles. Amin explains that 7-8 year old TPUs maintain 100% utilization while physical land and power carry 25-40 year depreciation schedules.5:51–8:08 · The host pushing back 1/10 Enterprise Readiness and Distributed Data Center Architecture Jeetu outlines the disparity between enterprise readiness and hyperscaler build-outs, detailing how localized power constraints force the creation of scale-across architectures across distances up to 900 km. The host primarily listens to the extended technical breakdown.8:08–15:56 · The host pushing back 4/10 Re-Inventing Computing: Hardware-Software Co-Design & Specialized Processors The host prompts a debate on whether Nvidia GPUs represent a return to mainframes versus scale-out systems. Amin directly challenges the host's mainframe framing, explaining that workloads dynamically allocate across scale-out software layers.15:56–20:32 · The host pushing back 2/10 The Next Generation of Scale-Up, Scale-Out, and Scale-Across Networking The host invites discussion on networking evolution, leading Amin to share how power utilities experience massive spikes during bursty network communications. Jeetu adds sharp critique regarding Broadcom's near-monopoly and the necessity for silicon diversity.20:32–23:42 · The host pushing back 6/10 Inference Systems, RL Constraints, and Intelligence Efficiency The host drills into inference bottlenecks and directly pushes back on Amin's efficiency framing, correcting him that 'intelligence per dollar' is a business model metric rather than a pure hardware capability metric, which Amin accepts.23:42–28:47 · The host pushing back 2/10 Internal AI Adoption: Code Migrations, Productivity, and Organizational Mindset Amin reveals that migrating Google's codebase off Bigtable was calculated at seven staff millennia, explaining why the effort was abandoned. Jeetu discusses internal tool usage across 25,000 engineers and the necessity of shifting engineer mindsets.28:47–32:21 · The host pushing back 5/10 Advice for Startups and 12-Month Technological Outlook The host asks for 12-month predictions while explicitly interjecting to ban the generic answer that 'models will get better'. Jeetu advises founders against building thin wrappers, promoting deeper integration and Cisco's technological momentum.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 8:42 Rejecting Mainframe Analogy

Amin directly refutes the host's premise that modern GPU clusters represent a return to mainframes, clarifying that modern infrastructure relies on software scale-out across flexible pools.

Hardest push from the host ▶ 29:25 Banning Generic AI Predictions

Martin interrupts Amin to explicitly forbid him from using the cliché answer 'models will get better', forcing a deeper technical response.

Biggest teaching moment ▶ 25:00 Seven Staff Millennia Lesson

Amin astounds the host by revealing that migrating legacy Bigtable to Spanner was estimated at seven staff millennia, forcing Google to abandon the migration.

The host holds their own ▶ 23:27 Correcting Efficiency Metrics

Martin demonstrates sharp expertise by challenging Amin's metric framing, pointing out that intelligence per dollar is a business metric rather than a pure processor capability, which Amin concedes.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Highlight Reel: The Scale of the AI Infrastructure Shift 3311 The host sets the stage by framing the historical scale of infrastructure cycles and asking guests to compare current demand to past shifts. The guests emphasize the unprecedented speed and scale, comparing it to the internet build-out combined with the Manhattan Project.
CapEx Spend Cycles and Power & Supply Constraints 5413 The host digs into internal CapEx signals and specifically queries how internal demand curves align with hardware depreciation cycles. Amin explains that 7-8 year old TPUs maintain 100% utilization while physical land and power carry 25-40 year depreciation schedules.
Enterprise Readiness and Distributed Data Center Architecture 2511 Jeetu outlines the disparity between enterprise readiness and hyperscaler build-outs, detailing how localized power constraints force the creation of scale-across architectures across distances up to 900 km. The host primarily listens to the extended technical breakdown.
Re-Inventing Computing: Hardware-Software Co-Design & Specialized Processors 6434 The host prompts a debate on whether Nvidia GPUs represent a return to mainframes versus scale-out systems. Amin directly challenges the host's mainframe framing, explaining that workloads dynamically allocate across scale-out software layers.
The Next Generation of Scale-Up, Scale-Out, and Scale-Across Networking 5532 The host invites discussion on networking evolution, leading Amin to share how power utilities experience massive spikes during bursty network communications. Jeetu adds sharp critique regarding Broadcom's near-monopoly and the necessity for silicon diversity.
Inference Systems, RL Constraints, and Intelligence Efficiency 6426 The host drills into inference bottlenecks and directly pushes back on Amin's efficiency framing, correcting him that 'intelligence per dollar' is a business model metric rather than a pure hardware capability metric, which Amin accepts.
Internal AI Adoption: Code Migrations, Productivity, and Organizational Mindset 4512 Amin reveals that migrating Google's codebase off Bigtable was calculated at seven staff millennia, explaining why the effort was abandoned. Jeetu discusses internal tool usage across 25,000 engineers and the necessity of shifting engineer mindsets.
Advice for Startups and 12-Month Technological Outlook 5325 The host asks for 12-month predictions while explicitly interjecting to ban the generic answer that 'models will get better'. Jeetu advises founders against building thin wrappers, promoting deeper integration and Cisco's technological momentum.

Statements from this episode (22)

Opinion
Patel: AI build-out combines internet, space race, and Manhattan Project
“This is, like, the combination of the build-out of the internet, the space race, and the Manhattan Project all put into one, where there's a geopolitical implication of it, there's an economic implication, there's a national security implication, and then ther…”
Jeetu Patel Oct 29, 2025 ▶ 0:03
Assertion Not checkable as stated
Vahdat: AI infrastructure build-out is 100x larger than 1990s internet boom
“The internet in the late nineties, early 2000 was big, and we felt like, oh my gosh, can't believe the Build out the rate. This makes it, I mean, 10 X is an understatement. It's a hundred X what the internet was.”
Amin Vahdat Oct 29, 2025 ▶ 0:24
Prediction Not checkable as stated
Patel: Current projections grossly underestimate required AI infrastructure build-out
“I think we're grossly underestimating the build-out. I think there's gonna be much more needed than what we are. Putting the you know, projections towards.”
Jeetu Patel Oct 29, 2025 ▶ 3:12
Assertion Not checkable as stated
Vahdat: Google's seven- and eight-year-old TPUs maintain 100% utilization
“Our seven and eight year old TPUs have a hundred percent utilization.”
Amin Vahdat Oct 29, 2025 ▶ 4:04
Prediction Not checkable as stated
Vahdat: AI hardware supply will lag behind demand in short term
“So one worry I have is that the supply isn't actually going to catch up to the demand as quickly as we'd all like.”
Amin Vahdat Oct 29, 2025 ▶ 5:03
Prediction Not checkable as stated
Vahdat: AI infrastructure deployment bottlenecks will persist for 3 to 5 years
“Like, in other words, literally, you all have some money, you can't spend it all as fast as you want. I think that's going to extend for three, four, five years.”
Amin Vahdat Oct 29, 2025 ▶ 5:22
Assertion Not checkable as stated
Patel: New AI data centers are built near available power sources
“Because there's not enough power singularly in one location, data centers are being built where the power is available rather than power being brought to where the data centers are.”
Jeetu Patel Oct 29, 2025 ▶ 6:51
Disclosure
Patel: Cisco silicon enables logical data centers spanning 900 kilometers
“We just launched a new piece of silicon as well as a new chip and a system for scale across networking, where you might have two data centers that act as a logical data center that could be up to eight, 900 kilometers apart.”
Jeetu Patel Oct 29, 2025 ▶ 7:44
Prediction Not checkable as stated
Vahdat: The entire computing stack will be unrecognizable by 2030
“And five years from now, whatever the computing stack is from the hardware to the software, right, it's going to be unrecognizable.”
Amin Vahdat Oct 29, 2025 ▶ 10:07
Assertion Supported
Vahdat: Google TPUs are 10x to 100x more energy efficient than CPUs
“TPU, I'll use that example again because I know it best for certain computation, is somewhere between 10 and a hundred times more efficient per watt, and it's this watt that really matters than a CPU.”
Amin Vahdat Oct 29, 2025 ▶ 12:52
Assertion Not checkable as stated
Vahdat: Bringing a specialized chip to production takes at least 2.5 years
“For the best teams in the world, really from concept to in live in production, the speed of light is two and a half years.”
Amin Vahdat Oct 29, 2025 ▶ 13:35
Assertion Supported
Patel: China is limited to manufacturing 7nm chips, not 2nm
“China actually doesn't make two nanometer chips, they make, you know, seven nanometer chips.”
Jeetu Patel Oct 29, 2025 ▶ 14:17
Assertion Supported
Vahdat: Gigawatt-scale AI workloads create power fluctuations visible to utilities
“We've written about this power utilities notice when we're doing network communication relative to computation at the scale of 1000 of megawatts. Right, like massive demand for power, stop all of a sudden and do some network communication, and then burst back …”
Amin Vahdat Oct 29, 2025 ▶ 17:16
Prediction Not checkable as stated
Patel: AI networking infrastructure will evolve toward dedicated inference-native designs
“And so I also feel like the way that networking will evolve is Rather than it being a training infrastructure that then gets applied to inferencing, you might have inferencing native infrastructure that gets built over time.”
Jeetu Patel Oct 29, 2025 ▶ 19:25
Opinion
Patel: Relying solely on Broadcom silicon creates a predatory monopoly
“Is, if you're just a rapper around Broadcom, Then you've got a monopoly that's gonna be a very predatory one. And so, one of the big reasons where Cisco is super relevant is you don't just have a Broadcom world with people just wrapping Broadcom, that their sy…”
Jeetu Patel Oct 29, 2025 ▶ 19:59
Assertion Not checkable as stated
Vahdat: Google is driving 10x to 100x reductions in AI inference costs
“What I would say though is that maybe something people don't realize is that we're actually driving massive reductions in the cost of inference. I mean, 10 X's and a hundred X's.”
Amin Vahdat Oct 29, 2025 ▶ 22:01
Assertion Not checkable as stated
Vahdat: Google estimated Bigtable to Spanner migration required 7,000 staff-years
“The estimate for doing that migration for Google was seven staffed millennia.”
Amin Vahdat Oct 29, 2025 ▶ 24:49
Prediction Not checkable as stated
Patel: Cisco targets 2x to 3x productivity boost for 25,000 engineers
“So, we've got 25,000 engineers. I'm hoping that we can get at least two or three X productivity within a very short amount of time within the next year.”
Jeetu Patel Oct 29, 2025 ▶ 27:56
Opinion
Patel: ChatGPT's initial competitive analysis beats human product marketers
“I think the first ChatGPT take on competitive is always better than what my, any product marketing person comes up by themselves.”
Jeetu Patel Oct 29, 2025 ▶ 28:33
Prediction Not checkable as stated
Vahdat: AI agents executing long tasks will be transformative within 12 months
“The agents that get built on top of them, And the frameworks for making that happen are also getting scary good. So the ability to have things go quite right for quite long over the coming 12 months is gonna be transformative.”
Amin Vahdat Oct 29, 2025 ▶ 29:35
Insight
Patel: AI startups building thin wrappers on third-party models won't endure
“I'd say the big shift and what I would urge startups to do is don't build thin wrappers around models that are other people's models. I think the combination of a model working very closely with the product and the model getting better as there's feedback in t…”
Jeetu Patel Oct 29, 2025 ▶ 30:00
Prediction Not checkable as stated
Vahdat: Multimodal AI tools will become transformative productivity drivers within 12 months
“I think that what's going to happen in the next 12 months is the same thing is going to be happening with input and output of images and video to these models. And to the extent that even for images, Imagine them as productivity and educational tools, not just…”
Amin Vahdat Oct 29, 2025 ▶ 31:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.