Sep 2, 2026 · 44m · latent-space

The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO

Sean Lie · 33m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space, Cerebras CTO Sean Lie discusses the breakthrough transition from traditional batch processing to ultra-fast inference reaching up to 10,000 tokens per second. He delves into Cerebras' wafer-scale architecture, the company's deep collaboration with OpenAI, and the physical engineering principles reshaping the semiconductor landscape.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.2 Guest teaching 4.0 Guest disagreement 1.3 The hosts pushing back 1.1
05100:0015:0030:000:00–3:39 · The hosts as informed peer 4/10 Rethinking Inference: From Batch Processing to Ultra-Fast Generation The hosts set up the episode after Hot Chips and ask broad opening questions about current semiconductor industry momentum. Sean Lie gives an enthusiastic overview of the rapid innovation across the entire stack.3:40–6:50 · The hosts as informed peer 4/10 Unveiling the CS-4: Breakthrough Speeds for Agentic Workflows The hosts prompt Sean to break down the newly announced CS-4 architecture. Sean details the modular Nexus platform, rack-level power delivery, and the implications of 4,000+ tokens per second for agentic loops.6:51–10:10 · The hosts as informed peer 5/10 Evolution of Wafer-Scale: Proving Viability and Partnering with OpenAI The host highlights how wafer-scale skepticism shifted over time, prompting Sean to recount the journey from building tech demos to production deployments powering OpenAI's ultra-fast models.10:11–15:45 · The hosts as informed peer 5/10 Architecting CS-5: The Road to 10,000 Tokens Per Second Hosts ask about CS-5 capabilities and capacity allocation with OpenAI. The host probes why OpenAI would make ultrafast inference public rather than keeping it proprietary, which Sean analyzes from business and mission perspectives.15:46–20:11 · The hosts as informed peer 6/10 Deconstructing OpenAI's Jalapeno Chip and AI-Driven EDA Tooling The conversation shifts to OpenAI's Jalapeno chip revealed at Hot Chips. Sean praises their AI-assisted EDA methodology and discusses potential pre-fill and decode disaggregation pipelines pairing Jalapeno with CS-5.20:11–25:52 · The hosts as informed peer 6/10 SRAM Limitations and Critiquing Grok's Architectural Strategy Sean offers a critique of Grok's architecture, noting that without wafer-scale integration, small SRAM capacity per chip forces them onto smaller 30B parameter models or impractical thousands-chip clusters for frontier scale.25:53–32:00 · The hosts as informed peer 6/10 Heterogeneous Disaggregation and the Untapped Potential of Co-Design Sean elaborates on heterogeneous disaggregation in multi-megawatt data centers and emphasizes how existing frontier models are heavily over-optimized for specific Nvidia GPU topologies rather than non-Nvidia architectures.32:00–40:40 · The hosts as informed peer 7/10 Competing on the Throughput Treadmill: Nvidia, AMD, and Custom Silicon Sean critiques the industry throughput treadmill and addresses Etched's claims, expressing skepticism about distributed storage without off-chip innovation while praising D-Matrix and 3D DRAM stacking packaging.40:41–43:04 · The hosts as informed peer 5/10 Semiconductor Geopolitics and the Rise of the Chinese AI Ecosystem The hosts raise Chinese foundation models running on native hardware like Huawei Ascend. Sean acknowledges the strategic challenge posed by Chinese dominance in open-weights models and calls for national-level semiconductor policy.43:04–43:56 · The hosts as informed peer 4/10 Cerebras IPO Milestone and Looking Ahead to Ultra-Fast Inference Hosts congratulate Sean on Cerebras's public listing milestone and close out the interview with playful demands for ultrafast inference rack allocations.0:00–3:39 · Guest teaching 2/10 Rethinking Inference: From Batch Processing to Ultra-Fast Generation The hosts set up the episode after Hot Chips and ask broad opening questions about current semiconductor industry momentum. Sean Lie gives an enthusiastic overview of the rapid innovation across the entire stack.3:40–6:50 · Guest teaching 4/10 Unveiling the CS-4: Breakthrough Speeds for Agentic Workflows The hosts prompt Sean to break down the newly announced CS-4 architecture. Sean details the modular Nexus platform, rack-level power delivery, and the implications of 4,000+ tokens per second for agentic loops.6:51–10:10 · Guest teaching 3/10 Evolution of Wafer-Scale: Proving Viability and Partnering with OpenAI The host highlights how wafer-scale skepticism shifted over time, prompting Sean to recount the journey from building tech demos to production deployments powering OpenAI's ultra-fast models.10:11–15:45 · Guest teaching 4/10 Architecting CS-5: The Road to 10,000 Tokens Per Second Hosts ask about CS-5 capabilities and capacity allocation with OpenAI. The host probes why OpenAI would make ultrafast inference public rather than keeping it proprietary, which Sean analyzes from business and mission perspectives.15:46–20:11 · Guest teaching 5/10 Deconstructing OpenAI's Jalapeno Chip and AI-Driven EDA Tooling The conversation shifts to OpenAI's Jalapeno chip revealed at Hot Chips. Sean praises their AI-assisted EDA methodology and discusses potential pre-fill and decode disaggregation pipelines pairing Jalapeno with CS-5.20:11–25:52 · Guest teaching 6/10 SRAM Limitations and Critiquing Grok's Architectural Strategy Sean offers a critique of Grok's architecture, noting that without wafer-scale integration, small SRAM capacity per chip forces them onto smaller 30B parameter models or impractical thousands-chip clusters for frontier scale.25:53–32:00 · Guest teaching 5/10 Heterogeneous Disaggregation and the Untapped Potential of Co-Design Sean elaborates on heterogeneous disaggregation in multi-megawatt data centers and emphasizes how existing frontier models are heavily over-optimized for specific Nvidia GPU topologies rather than non-Nvidia architectures.32:00–40:40 · Guest teaching 5/10 Competing on the Throughput Treadmill: Nvidia, AMD, and Custom Silicon Sean critiques the industry throughput treadmill and addresses Etched's claims, expressing skepticism about distributed storage without off-chip innovation while praising D-Matrix and 3D DRAM stacking packaging.40:41–43:04 · Guest teaching 4/10 Semiconductor Geopolitics and the Rise of the Chinese AI Ecosystem The hosts raise Chinese foundation models running on native hardware like Huawei Ascend. Sean acknowledges the strategic challenge posed by Chinese dominance in open-weights models and calls for national-level semiconductor policy.43:04–43:56 · Guest teaching 2/10 Cerebras IPO Milestone and Looking Ahead to Ultra-Fast Inference Hosts congratulate Sean on Cerebras's public listing milestone and close out the interview with playful demands for ultrafast inference rack allocations.0:00–3:39 · Guest disagreement 1/10 Rethinking Inference: From Batch Processing to Ultra-Fast Generation The hosts set up the episode after Hot Chips and ask broad opening questions about current semiconductor industry momentum. Sean Lie gives an enthusiastic overview of the rapid innovation across the entire stack.3:40–6:50 · Guest disagreement 0/10 Unveiling the CS-4: Breakthrough Speeds for Agentic Workflows The hosts prompt Sean to break down the newly announced CS-4 architecture. Sean details the modular Nexus platform, rack-level power delivery, and the implications of 4,000+ tokens per second for agentic loops.6:51–10:10 · Guest disagreement 1/10 Evolution of Wafer-Scale: Proving Viability and Partnering with OpenAI The host highlights how wafer-scale skepticism shifted over time, prompting Sean to recount the journey from building tech demos to production deployments powering OpenAI's ultra-fast models.10:11–15:45 · Guest disagreement 1/10 Architecting CS-5: The Road to 10,000 Tokens Per Second Hosts ask about CS-5 capabilities and capacity allocation with OpenAI. The host probes why OpenAI would make ultrafast inference public rather than keeping it proprietary, which Sean analyzes from business and mission perspectives.15:46–20:11 · Guest disagreement 1/10 Deconstructing OpenAI's Jalapeno Chip and AI-Driven EDA Tooling The conversation shifts to OpenAI's Jalapeno chip revealed at Hot Chips. Sean praises their AI-assisted EDA methodology and discusses potential pre-fill and decode disaggregation pipelines pairing Jalapeno with CS-5.20:11–25:52 · Guest disagreement 4/10 SRAM Limitations and Critiquing Grok's Architectural Strategy Sean offers a critique of Grok's architecture, noting that without wafer-scale integration, small SRAM capacity per chip forces them onto smaller 30B parameter models or impractical thousands-chip clusters for frontier scale.25:53–32:00 · Guest disagreement 1/10 Heterogeneous Disaggregation and the Untapped Potential of Co-Design Sean elaborates on heterogeneous disaggregation in multi-megawatt data centers and emphasizes how existing frontier models are heavily over-optimized for specific Nvidia GPU topologies rather than non-Nvidia architectures.32:00–40:40 · Guest disagreement 3/10 Competing on the Throughput Treadmill: Nvidia, AMD, and Custom Silicon Sean critiques the industry throughput treadmill and addresses Etched's claims, expressing skepticism about distributed storage without off-chip innovation while praising D-Matrix and 3D DRAM stacking packaging.40:41–43:04 · Guest disagreement 1/10 Semiconductor Geopolitics and the Rise of the Chinese AI Ecosystem The hosts raise Chinese foundation models running on native hardware like Huawei Ascend. Sean acknowledges the strategic challenge posed by Chinese dominance in open-weights models and calls for national-level semiconductor policy.43:04–43:56 · Guest disagreement 0/10 Cerebras IPO Milestone and Looking Ahead to Ultra-Fast Inference Hosts congratulate Sean on Cerebras's public listing milestone and close out the interview with playful demands for ultrafast inference rack allocations.0:00–3:39 · The hosts pushing back 1/10 Rethinking Inference: From Batch Processing to Ultra-Fast Generation The hosts set up the episode after Hot Chips and ask broad opening questions about current semiconductor industry momentum. Sean Lie gives an enthusiastic overview of the rapid innovation across the entire stack.3:40–6:50 · The hosts pushing back 0/10 Unveiling the CS-4: Breakthrough Speeds for Agentic Workflows The hosts prompt Sean to break down the newly announced CS-4 architecture. Sean details the modular Nexus platform, rack-level power delivery, and the implications of 4,000+ tokens per second for agentic loops.6:51–10:10 · The hosts pushing back 1/10 Evolution of Wafer-Scale: Proving Viability and Partnering with OpenAI The host highlights how wafer-scale skepticism shifted over time, prompting Sean to recount the journey from building tech demos to production deployments powering OpenAI's ultra-fast models.10:11–15:45 · The hosts pushing back 1/10 Architecting CS-5: The Road to 10,000 Tokens Per Second Hosts ask about CS-5 capabilities and capacity allocation with OpenAI. The host probes why OpenAI would make ultrafast inference public rather than keeping it proprietary, which Sean analyzes from business and mission perspectives.15:46–20:11 · The hosts pushing back 2/10 Deconstructing OpenAI's Jalapeno Chip and AI-Driven EDA Tooling The conversation shifts to OpenAI's Jalapeno chip revealed at Hot Chips. Sean praises their AI-assisted EDA methodology and discusses potential pre-fill and decode disaggregation pipelines pairing Jalapeno with CS-5.20:11–25:52 · The hosts pushing back 2/10 SRAM Limitations and Critiquing Grok's Architectural Strategy Sean offers a critique of Grok's architecture, noting that without wafer-scale integration, small SRAM capacity per chip forces them onto smaller 30B parameter models or impractical thousands-chip clusters for frontier scale.25:53–32:00 · The hosts pushing back 1/10 Heterogeneous Disaggregation and the Untapped Potential of Co-Design Sean elaborates on heterogeneous disaggregation in multi-megawatt data centers and emphasizes how existing frontier models are heavily over-optimized for specific Nvidia GPU topologies rather than non-Nvidia architectures.32:00–40:40 · The hosts pushing back 2/10 Competing on the Throughput Treadmill: Nvidia, AMD, and Custom Silicon Sean critiques the industry throughput treadmill and addresses Etched's claims, expressing skepticism about distributed storage without off-chip innovation while praising D-Matrix and 3D DRAM stacking packaging.40:41–43:04 · The hosts pushing back 1/10 Semiconductor Geopolitics and the Rise of the Chinese AI Ecosystem The hosts raise Chinese foundation models running on native hardware like Huawei Ascend. Sean acknowledges the strategic challenge posed by Chinese dominance in open-weights models and calls for national-level semiconductor policy.43:04–43:56 · The hosts pushing back 0/10 Cerebras IPO Milestone and Looking Ahead to Ultra-Fast Inference Hosts congratulate Sean on Cerebras's public listing milestone and close out the interview with playful demands for ultrafast inference rack allocations.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 22:15 Sean critiques Grok's memory scaling limits

Sean bluntly points out that Grok's non-wafer SRAM architecture cannot practically run frontier trillion-parameter models, highlighting that their launch numbers were limited to a 31B model.

Hardest push from the hosts ▶ 14:53 Host challenges OpenAI's decision to expose ultra-fast capacity

The host questions the business logic of OpenAI exposing their proprietary fast inference capability to enterprise customers instead of keeping it as an internal moat.

Biggest teaching moment ▶ 28:10 Sean reframes data center modularity for chip architects

Sean dismantles the assumption that disaggregation restricts data center design, explaining that architects view entire multi-gigawatt data centers like a single modular chip.

The host holds their own ▶ 19:07 Host highlights Jalapeno's lack of pre-fill/decode specialization

The host demonstrates sharp technical knowledge by immediately pointing out that OpenAI did not specialize the Jalapeno architecture for prefill/decode disaggregation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Rethinking Inference: From Batch Processing to Ultra-Fast Generation 4211 The hosts set up the episode after Hot Chips and ask broad opening questions about current semiconductor industry momentum. Sean Lie gives an enthusiastic overview of the rapid innovation across the entire stack.
Unveiling the CS-4: Breakthrough Speeds for Agentic Workflows 4400 The hosts prompt Sean to break down the newly announced CS-4 architecture. Sean details the modular Nexus platform, rack-level power delivery, and the implications of 4,000+ tokens per second for agentic loops.
Evolution of Wafer-Scale: Proving Viability and Partnering with OpenAI 5311 The host highlights how wafer-scale skepticism shifted over time, prompting Sean to recount the journey from building tech demos to production deployments powering OpenAI's ultra-fast models.
Architecting CS-5: The Road to 10,000 Tokens Per Second 5411 Hosts ask about CS-5 capabilities and capacity allocation with OpenAI. The host probes why OpenAI would make ultrafast inference public rather than keeping it proprietary, which Sean analyzes from business and mission perspectives.
Deconstructing OpenAI's Jalapeno Chip and AI-Driven EDA Tooling 6512 The conversation shifts to OpenAI's Jalapeno chip revealed at Hot Chips. Sean praises their AI-assisted EDA methodology and discusses potential pre-fill and decode disaggregation pipelines pairing Jalapeno with CS-5.
SRAM Limitations and Critiquing Grok's Architectural Strategy 6642 Sean offers a critique of Grok's architecture, noting that without wafer-scale integration, small SRAM capacity per chip forces them onto smaller 30B parameter models or impractical thousands-chip clusters for frontier scale.
Heterogeneous Disaggregation and the Untapped Potential of Co-Design 6511 Sean elaborates on heterogeneous disaggregation in multi-megawatt data centers and emphasizes how existing frontier models are heavily over-optimized for specific Nvidia GPU topologies rather than non-Nvidia architectures.
Competing on the Throughput Treadmill: Nvidia, AMD, and Custom Silicon 7532 Sean critiques the industry throughput treadmill and addresses Etched's claims, expressing skepticism about distributed storage without off-chip innovation while praising D-Matrix and 3D DRAM stacking packaging.
Semiconductor Geopolitics and the Rise of the Chinese AI Ecosystem 5411 The hosts raise Chinese foundation models running on native hardware like Huawei Ascend. Sean acknowledges the strategic challenge posed by Chinese dominance in open-weights models and calls for national-level semiconductor policy.
Cerebras IPO Milestone and Looking Ahead to Ultra-Fast Inference 4200 Hosts congratulate Sean on Cerebras's public listing milestone and close out the interview with playful demands for ultrafast inference rack allocations.

Statements from this episode (19)

Assertion Partly supported
Lie: Cerebras CS-4 doubles wafer power and bandwidth while halving latency
“We've designed this a modular platform that provides twice the amount of power to the wafer than we have in our previous generation. Twice the amount of interconnect bandwidth, half the latency.”
Sean Lie Sep 2, 2026 ▶ 4:22
Assertion Supported
Lie: Cerebras demoed GPT running at over 4,400 TPS at Hot Chips
“We here in this demo that we gave at hot chips we're showing GPT OSS running at over 4000 400 TPS, which is just mind blowing.”
Sean Lie Sep 2, 2026 ▶ 5:16
Insight
Lie: Ultra-fast token generation enables more capable AI agent reasoning
“If you're running your model at over 4000 tokens per second. Now the, you can do, you know, more agentic loops. You can do more reasoning. Ultimately you get significantly more capable, more intelligent agents.”
Sean Lie Sep 2, 2026 ▶ 6:16
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37
Prediction Open · timeframe Dec 2027
Sean Lie: Next Cerebras Chip Will Run Frontier AI at 5,000 TPS
“And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, We'll be pushing the performance even further. So we just saw a two X improvement this year with CS four. We're going to push it e…”
Sean Lie Sep 2, 2026 ▶ 11:01
Disclosure
Lie: Cerebras is currently sold out of all hardware capacity
“Right now we are basically, you know, sold out Of everything that we're building, right?”
Sean Lie Sep 2, 2026 ▶ 12:24
Assertion Not checkable as stated
Lie: OpenAI uses Cerebras hardware internally for incident response and research
“So right now internally, they're using it for a lot of really critical use cases where the speed really, really matters. Like they're using it in like, Their incidents response teams, right? When there's an outage in their service, for example, every single se…”
Sean Lie Sep 2, 2026 ▶ 13:29
Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Sean Lie Sep 2, 2026 ▶ 16:30
Disclosure
Sean Lie: OpenAI is Cerebras's biggest customer
“OpenAI is our biggest customer.”
Sean Lie Sep 2, 2026 ▶ 18:26
Disclosure
Sean Lie: Cerebras uses OpenAI's internal AI tools for chip design
“We're also collaborating very closely with OpenAI, right, to use their tools to help us also continue to push what's possible in our chip design, in our software, and all that.”
Sean Lie Sep 2, 2026 ▶ 19:59
Assertion Supported
Lie: Trillion-parameter models require thousands of Groq LPUs for weights
“To run a frontier level model, like, let's say, a few trillion parameters, you need thousands and thousands of Grok LPUs just to hold the weights, right?”
Sean Lie Sep 2, 2026 ▶ 23:17
Prediction Not checkable as stated
Lie: Groq will be forced to focus on significantly smaller models
“I think what's, what's going to end up happening is they're going to end up focusing on significantly smaller models. You know, if you have that limitation in your architecture, then I think that's what ends up happening.”
Sean Lie Sep 2, 2026 ▶ 23:42
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Sean Lie Sep 2, 2026 ▶ 23:58
Opinion
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Sean Lie Sep 2, 2026 ▶ 30:35
Insight
Lie: 100 to 200 tokens per second is becoming the new batch mode
“What used to be considered fast at, like, 1000 100 or 200 tokens per second is quickly becoming the new batch mode. Right. Quickly becoming the new, like overnight is true.”
Sean Lie Sep 2, 2026 ▶ 32:29
Opinion
Lie: AMD, Trainium, and TPU are all trying to build a better Nvidia Rubin
“AMD, Tranium, in many ways TPU, like all of these in my mind are all trying to build a better Reuben, right? And there's a huge amount of value in that.”
Sean Lie Sep 2, 2026 ▶ 33:21
Opinion
Sean Lie: Etched is not building anything better than a traditional GPU
“I think my reaction when I see pictures like this is that it's very impressive graphics design. But I also don't see them building anything beyond just, or trying to build something better than just, you know, a traditional GPU, right? You know, they've made c…”
Sean Lie Sep 2, 2026 ▶ 34:23
Assertion Not checkable as stated
Sean Lie: Cerebras has solved yield at scale and 3D packaging
“And so, you know, we already have solved yield at scale. For example, we've already solved how you can actually package in a three-dimensional way. And that's, You know, problems that Samsung, that D-Matrix, and everybody else are also gonna have to solve over…”
Sean Lie Sep 2, 2026 ▶ 40:25
Assertion Not checkable as stated
Sean Lie: 95% of High-Quality Open Models Come From Chinese Labs
“The open source model market is a hundred percent Chinese, right? . A hundred percent, but. Almost. Okay. 95%, right? Most of the big models, most of the big open models that are, you know, high quality are coming from the Chinese labs.”
Sean Lie Sep 2, 2026 ▶ 41:22
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.