Assertion certainty 3/5 debate potential 4/5

Sean Lie: 95% of High-Quality Open Models Come From Chinese Labs

Sean Lie · The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO · Sep 2, 2026 · at 41:22

Cerebras CTO Sean Lie discusses the semiconductor and AI model landscape between the US and China.

0:00 / 0:17exact quote · 17.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The open source model market is a hundred percent Chinese, right? . A hundred percent, but. Almost. Okay. 95%, right? Most of the big models, most of the big open models that are, you know, high quality are coming from the Chinese labs.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Sean Lie

Opinion
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Sean Lie Sep 2, 2026 ▶ 16:30 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Prediction Not checkable as stated
Lie: Groq will be forced to focus on significantly smaller models
“I think what's, what's going to end up happening is they're going to end up focusing on significantly smaller models. You know, if you have that limitation in your architecture, then I think that's what ends up happening.”
Sean Lie Sep 2, 2026 ▶ 23:42 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Opinion
Sean Lie: Etched is not building anything better than a traditional GPU
“I think my reaction when I see pictures like this is that it's very impressive graphics design. But I also don't see them building anything beyond just, or trying to build something better than just, you know, a traditional GPU, right? You know, they've made c…”
Sean Lie Sep 2, 2026 ▶ 34:23 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Prediction Open · timeframe Dec 2027
Sean Lie: Next Cerebras Chip Will Run Frontier AI at 5,000 TPS
“And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, We'll be pushing the performance even further. So we just saw a two X improvement this year with CS four. We're going to push it e…”
Sean Lie Sep 2, 2026 ▶ 11:01 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Sean Lie Sep 2, 2026 ▶ 23:58 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.