Sean Lie: Next Cerebras Chip Will Run Frontier AI at 5,000 TPS
Sean Lie · The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO · Sep 2, 2026 · at 11:01
Cerebras CTO Sean Lie discusses the performance roadmap for Cerebras's next-generation wafer-scale chip and modular Nexus rack platform.
“And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, We'll be pushing the performance even further. So we just saw a two X improvement this year with CS four. We're going to push it even further with another two X improvement in performance. And what this ultimately means is you'll be able to run, you know, medium sized models like GPT-OSS or JAMA at speeds up to 10,000 TPS. And Even frontier level models like Kimi or DeepSeq and GPT-Five-Six-Soul up to 5000 TPS.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →