People, every show

Sean Lie

Co-founder & CTO, Cerebras Systems. On 1 show, 1 appearance. The Shows tab opens the full record on each.

founderexecutiveengineer

Sean Lie co-founded Cerebras Systems in 2015, where he leads hardware and system architecture for the company's Wafer-Scale Engine processors and AI systems. Previously, he served as Lead Hardware Architect at SeaMicro and became an AMD Fellow after SeaMicro's acquisition.

1shows
1appearances
19statements
5resolved
4supported
0contradicted
80%fully supported
2said about them ↓

Everything Sean Lie said on any show that made the record, most notable first. Each card names its show and opens the statement there.

OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Sean Lie Sep 2, 2026 ▶ 16:30 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Prediction Not checkable as stated
Lie: Groq will be forced to focus on significantly smaller models
“I think what's, what's going to end up happening is they're going to end up focusing on significantly smaller models. You know, if you have that limitation in your architecture, then I think that's what ends up happening.”
Sean Lie Sep 2, 2026 ▶ 23:42 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Sean Lie: Etched is not building anything better than a traditional GPU
“I think my reaction when I see pictures like this is that it's very impressive graphics design. But I also don't see them building anything beyond just, or trying to build something better than just, you know, a traditional GPU, right? You know, they've made c…”
Sean Lie Sep 2, 2026 ▶ 34:23 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Not checkable as stated
Sean Lie: 95% of High-Quality Open Models Come From Chinese Labs
“The open source model market is a hundred percent Chinese, right? . A hundred percent, but. Almost. Okay. 95%, right? Most of the big models, most of the big open models that are, you know, high quality are coming from the Chinese labs.”
Sean Lie Sep 2, 2026 ▶ 41:22 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie Sep 2, 2026 ▶ 8:37 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Prediction Open · timeframe Dec 2027
Sean Lie: Next Cerebras Chip Will Run Frontier AI at 5,000 TPS
“And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, We'll be pushing the performance even further. So we just saw a two X improvement this year with CS four. We're going to push it e…”
Sean Lie Sep 2, 2026 ▶ 11:01 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Sean Lie Sep 2, 2026 ▶ 23:58 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Sean Lie Sep 2, 2026 ▶ 30:35 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Not checkable as stated
Sean Lie: Cerebras has solved yield at scale and 3D packaging
“And so, you know, we already have solved yield at scale. For example, we've already solved how you can actually package in a three-dimensional way. And that's, You know, problems that Samsung, that D-Matrix, and everybody else are also gonna have to solve over…”
Sean Lie Sep 2, 2026 ▶ 40:25 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Supported
Lie: Cerebras demoed GPT running at over 4,400 TPS at Hot Chips
“We here in this demo that we gave at hot chips we're showing GPT OSS running at over 4000 400 TPS, which is just mind blowing.”
Sean Lie Sep 2, 2026 ▶ 5:16 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Lie: Ultra-fast token generation enables more capable AI agent reasoning
“If you're running your model at over 4000 tokens per second. Now the, you can do, you know, more agentic loops. You can do more reasoning. Ultimately you get significantly more capable, more intelligent agents.”
Sean Lie Sep 2, 2026 ▶ 6:16 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Disclosure
Sean Lie: OpenAI is Cerebras's biggest customer
“OpenAI is our biggest customer.”
Sean Lie Sep 2, 2026 ▶ 18:26 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Supported
Lie: Trillion-parameter models require thousands of Groq LPUs for weights
“To run a frontier level model, like, let's say, a few trillion parameters, you need thousands and thousands of Grok LPUs just to hold the weights, right?”
Sean Lie Sep 2, 2026 ▶ 23:17 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Lie: 100 to 200 tokens per second is becoming the new batch mode
“What used to be considered fast at, like, 1000 100 or 200 tokens per second is quickly becoming the new batch mode. Right. Quickly becoming the new, like overnight is true.”
Sean Lie Sep 2, 2026 ▶ 32:29 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Lie: AMD, Trainium, and TPU are all trying to build a better Nvidia Rubin
“AMD, Tranium, in many ways TPU, like all of these in my mind are all trying to build a better Reuben, right? And there's a huge amount of value in that.”
Sean Lie Sep 2, 2026 ▶ 33:21 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Partly supported
Lie: Cerebras CS-4 doubles wafer power and bandwidth while halving latency
“We've designed this a modular platform that provides twice the amount of power to the wafer than we have in our previous generation. Twice the amount of interconnect bandwidth, half the latency.”
Sean Lie Sep 2, 2026 ▶ 4:22 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Disclosure
Lie: Cerebras is currently sold out of all hardware capacity
“Right now we are basically, you know, sold out Of everything that we're building, right?”
Sean Lie Sep 2, 2026 ▶ 12:24 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Assertion Not checkable as stated
Lie: OpenAI uses Cerebras hardware internally for incident response and research
“So right now internally, they're using it for a lot of really critical use cases where the speed really, really matters. Like they're using it in like, Their incidents response teams, right? When there's an outage in their service, for example, every single se…”
Sean Lie Sep 2, 2026 ▶ 13:29 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
LATENT SPACE Disclosure
Sean Lie: Cerebras uses OpenAI's internal AI tools for chip design
“We're also collaborating very closely with OpenAI, right, to use their tools to help us also continue to push what's possible in our chip design, in our software, and all that.”
Sean Lie Sep 2, 2026 ▶ 19:59 The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO

The other half of the tape: Sean Lie's own voice is left out of every number here. Other people bring the name up 2 times in 1 episode across the shows. every mention, with the transcript →

Who brings them up most Andrew Feldman 2

Every mention by year

tap a year for its mentions
0011212025episodesmentions
0112025episodes it came up in
0010.5212025episodesmentions per episode

Latent Space 2

2025 2 mentions in 1 episode

One line per show, most statements first. The link opens Sean's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER Co-founder & CTO, Cerebras Systems 1 19 80% 4/5 full record on Latent Space →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.