Everything Sean Lie said on any show that made the record, most notable first. Each card names its show and opens the statement there.
OpenAI built a significantly better GPU than Nvidia with Jalapeno chip
“I think that, like, they, you know, they pushed a lot on the performance and the fact that they're significantly better than, you know, better performance than the GPU than NVIDIA. But what I see is that they've built a significantly better GPU. And that in it…”
Lie: Groq will be forced to focus on significantly smaller models
“I think what's, what's going to end up happening is they're going to end up focusing on significantly smaller models. You know, if you have that limitation in your architecture, then I think that's what ends up happening.”
Sean Lie: Etched is not building anything better than a traditional GPU
“I think my reaction when I see pictures like this is that it's very impressive graphics design. But I also don't see them building anything beyond just, or trying to build something better than just, you know, a traditional GPU, right? You know, they've made c…”
Sean Lie: 95% of High-Quality Open Models Come From Chinese Labs
“The open source model market is a hundred percent Chinese, right? . A hundred percent, but. Almost. Okay. 95%, right? Most of the big models, most of the big open models that are, you know, high quality are coming from the Chinese labs.”
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Sean Lie: Next Cerebras Chip Will Run Frontier AI at 5,000 TPS
“And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, We'll be pushing the performance even further. So we just saw a two X improvement this year with CS four. We're going to push it e…”
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Sean Lie: Cerebras has solved yield at scale and 3D packaging
“And so, you know, we already have solved yield at scale. For example, we've already solved how you can actually package in a three-dimensional way. And that's, You know, problems that Samsung, that D-Matrix, and everybody else are also gonna have to solve over…”
Lie: Cerebras demoed GPT running at over 4,400 TPS at Hot Chips
“We here in this demo that we gave at hot chips we're showing GPT OSS running at over 4000 400 TPS, which is just mind blowing.”
Lie: Ultra-fast token generation enables more capable AI agent reasoning
“If you're running your model at over 4000 tokens per second. Now the, you can do, you know, more agentic loops. You can do more reasoning. Ultimately you get significantly more capable, more intelligent agents.”
Sean Lie: OpenAI is Cerebras's biggest customer
“OpenAI is our biggest customer.”
Lie: Trillion-parameter models require thousands of Groq LPUs for weights
“To run a frontier level model, like, let's say, a few trillion parameters, you need thousands and thousands of Grok LPUs just to hold the weights, right?”
Lie: 100 to 200 tokens per second is becoming the new batch mode
“What used to be considered fast at, like, 1000 100 or 200 tokens per second is quickly becoming the new batch mode. Right. Quickly becoming the new, like overnight is true.”
Lie: AMD, Trainium, and TPU are all trying to build a better Nvidia Rubin
“AMD, Tranium, in many ways TPU, like all of these in my mind are all trying to build a better Reuben, right? And there's a huge amount of value in that.”
Lie: Cerebras CS-4 doubles wafer power and bandwidth while halving latency
“We've designed this a modular platform that provides twice the amount of power to the wafer than we have in our previous generation. Twice the amount of interconnect bandwidth, half the latency.”
Lie: Cerebras is currently sold out of all hardware capacity
“Right now we are basically, you know, sold out Of everything that we're building, right?”
Lie: OpenAI uses Cerebras hardware internally for incident response and research
“So right now internally, they're using it for a lot of really critical use cases where the speed really, really matters. Like they're using it in like, Their incidents response teams, right? When there's an outage in their service, for example, every single se…”
Sean Lie: Cerebras uses OpenAI's internal AI tools for chip design
“We're also collaborating very closely with OpenAI, right, to use their tools to help us also continue to push what's possible in our chip design, in our software, and all that.”