Sep 2, 2026 · 44m · latent-space
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of Latent Space, Cerebras CTO Sean Lie discusses the breakthrough transition from traditional batch processing to ultra-fast inference reaching up to 10,000 tokens per second. He delves into Cerebras' wafer-scale architecture, the company's deep collaboration with OpenAI, and the physical engineering principles reshaping the semiconductor landscape.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Sean bluntly points out that Grok's non-wafer SRAM architecture cannot practically run frontier trillion-parameter models, highlighting that their launch numbers were limited to a 31B model.
Hardest push from the hosts ▶ 14:53 Host challenges OpenAI's decision to expose ultra-fast capacityThe host questions the business logic of OpenAI exposing their proprietary fast inference capability to enterprise customers instead of keeping it as an internal moat.
Biggest teaching moment ▶ 28:10 Sean reframes data center modularity for chip architectsSean dismantles the assumption that disaggregation restricts data center design, explaining that architects view entire multi-gigawatt data centers like a single modular chip.
The host holds their own ▶ 19:07 Host highlights Jalapeno's lack of pre-fill/decode specializationThe host demonstrates sharp technical knowledge by immediately pointing out that OpenAI did not specialize the Jalapeno architecture for prefill/decode disaggregation.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Rethinking Inference: From Batch Processing to Ultra-Fast Generation | 4 | 2 | 1 | 1 | The hosts set up the episode after Hot Chips and ask broad opening questions about current semiconductor industry momentum. Sean Lie gives an enthusiastic overview of the rapid innovation across the entire stack. | |
| Unveiling the CS-4: Breakthrough Speeds for Agentic Workflows | 4 | 4 | 0 | 0 | The hosts prompt Sean to break down the newly announced CS-4 architecture. Sean details the modular Nexus platform, rack-level power delivery, and the implications of 4,000+ tokens per second for agentic loops. | |
| Evolution of Wafer-Scale: Proving Viability and Partnering with OpenAI | 5 | 3 | 1 | 1 | The host highlights how wafer-scale skepticism shifted over time, prompting Sean to recount the journey from building tech demos to production deployments powering OpenAI's ultra-fast models. | |
| Architecting CS-5: The Road to 10,000 Tokens Per Second | 5 | 4 | 1 | 1 | Hosts ask about CS-5 capabilities and capacity allocation with OpenAI. The host probes why OpenAI would make ultrafast inference public rather than keeping it proprietary, which Sean analyzes from business and mission perspectives. | |
| Deconstructing OpenAI's Jalapeno Chip and AI-Driven EDA Tooling | 6 | 5 | 1 | 2 | The conversation shifts to OpenAI's Jalapeno chip revealed at Hot Chips. Sean praises their AI-assisted EDA methodology and discusses potential pre-fill and decode disaggregation pipelines pairing Jalapeno with CS-5. | |
| SRAM Limitations and Critiquing Grok's Architectural Strategy | 6 | 6 | 4 | 2 | Sean offers a critique of Grok's architecture, noting that without wafer-scale integration, small SRAM capacity per chip forces them onto smaller 30B parameter models or impractical thousands-chip clusters for frontier scale. | |
| Heterogeneous Disaggregation and the Untapped Potential of Co-Design | 6 | 5 | 1 | 1 | Sean elaborates on heterogeneous disaggregation in multi-megawatt data centers and emphasizes how existing frontier models are heavily over-optimized for specific Nvidia GPU topologies rather than non-Nvidia architectures. | |
| Competing on the Throughput Treadmill: Nvidia, AMD, and Custom Silicon | 7 | 5 | 3 | 2 | Sean critiques the industry throughput treadmill and addresses Etched's claims, expressing skepticism about distributed storage without off-chip innovation while praising D-Matrix and 3D DRAM stacking packaging. | |
| Semiconductor Geopolitics and the Rise of the Chinese AI Ecosystem | 5 | 4 | 1 | 1 | The hosts raise Chinese foundation models running on native hardware like Huawei Ascend. Sean acknowledges the strategic challenge posed by Chinese dominance in open-weights models and calls for national-level semiconductor policy. | |
| Cerebras IPO Milestone and Looking Ahead to Ultra-Fast Inference | 4 | 2 | 0 | 0 | Hosts congratulate Sean on Cerebras's public listing milestone and close out the interview with playful demands for ultrafast inference rack allocations. |