why aren't all 40 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Lie: Groq will be forced to focus on significantly smaller models
“I think what's, what's going to end up happening is they're going to end up focusing on significantly smaller models. You know, if you have that limitation in your architecture, then I think that's what ends up happening.”
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Opinion
Agrawal: Cerebras and Groq offer cloud APIs because their software struggles
“You know, if you look at Cerebrus and Grok and others, they've really tried to do this, their cloud kind of portal. And for us, you know, it shows two things. One is the difficulty of software for third party to implement that, that that's why they're kind of …”
Prediction Open · timeframe Dec 2027
Sean Lie: Next Cerebras Chip Will Run Frontier AI at 5,000 TPS
“And in particular, we designed it together with our next generation wafer chip that will be coming next year. And with that chip, We'll be pushing the performance even further. So we just saw a two X improvement this year with CS four. We're going to push it e…”
Assertion Supported
Lie: Cerebras chips have 100x more memory than Groq LPUs
“One of our chips has, You know, order a hundred times more memory than one of their chips, right? So you got two orders of magnitude difference in scale kind of for free, right?”
Opinion
Untapped AI hardware opportunity is co-designing models for non-Nvidia architectures
“What I think is the most untapped opportunity right now frankly for Cerebrus, but frankly for the entire non-NVIDIA environment right? And the non-environ sorry community is that, you know, we're all running models that were designed for NVIDIA GPUs, right? An…”
Assertion Not checkable as stated
Sean Lie: Cerebras has solved yield at scale and 3D packaging
“And so, you know, we already have solved yield at scale. For example, we've already solved how you can actually package in a three-dimensional way. And that's, You know, problems that Samsung, that D-Matrix, and everybody else are also gonna have to solve over…”
Assertion Supported
Lie: Cerebras runs OpenAI's flagship model 14x faster than GPUs
“We're running you know, frontier level, one of the most intelligent models, right? OpenAI's largest, most capable, most intelligent model right now at 14 times faster than their normal, you know, GPU speeds.”
Opinion
O'Laughlin: Prior to Cerebras and Groq, AI accelerator startups were failures
“The reason why my hit rate for every AI accelerator trip is, like, very, like, I just don't believe in them is because, like, where are they? Until Cerebrus and Grok, honestly, they were all considered failures, and even then, we're like, what are they gonna d…”
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Assertion Supported
Houston: Groq and Cerebras outperform Nvidia on latency
“There's also, like, non-NVIDIA stacks, like the Grok, or Cerebris, or some of these custom silicon companies that are super interesting, and all, and outperformed the NVIDIA stack in terms of latency and things like that.”
Insight
Lie: 100 to 200 tokens per second is becoming the new batch mode
“What used to be considered fast at, like, 1000 100 or 200 tokens per second is quickly becoming the new batch mode. Right. Quickly becoming the new, like overnight is true.”
Assertion Supported
Lie: Cerebras demoed GPT running at over 4,400 TPS at Hot Chips
“We here in this demo that we gave at hot chips we're showing GPT OSS running at over 4000 400 TPS, which is just mind blowing.”
Assertion Supported
Lie: Trillion-parameter models require thousands of Groq LPUs for weights
“To run a frontier level model, like, let's say, a few trillion parameters, you need thousands and thousands of Grok LPUs just to hold the weights, right?”
Disclosure
Sean Lie: OpenAI is Cerebras's biggest customer
“OpenAI is our biggest customer.”
Insight
Lie: Ultra-fast token generation enables more capable AI agent reasoning
“If you're running your model at over 4000 tokens per second. Now the, you can do, you know, more agentic loops. You can do more reasoning. Ultimately you get significantly more capable, more intelligent agents.”
Opinion
Lie: AMD, Trainium, and TPU are all trying to build a better Nvidia Rubin
“AMD, Tranium, in many ways TPU, like all of these in my mind are all trying to build a better Reuben, right? And there's a huge amount of value in that.”
Assertion Supported
Feldman: Cerebras raised $1.1B at an $8.1B valuation
“So we announced a 1.1 billion dollar fundraise that we had completed. It was done at an 8.1 billion dollar post money valuation, and it was led by Fidelity and Atreides management.”
Insight
Feldman: Multi-chip SRAM architectures make speculative decoding extremely difficult
“It limits the things you can do. It makes all sorts of cool AI techniques like speculative decode extremely difficult. Whereas if you have a giant chip, you might only need a handful.”
Assertion Contradicted
No competing AI framework disaggregates model storage from compute like Cerebras
“Like basically not, no one is doing anything close to where you're disaggregating. Model storage from compute. And none of these examples above do that either.”
Assertion Supported
GPUs cannot handle unstructured sparsity as efficiently as Cerebras hardware
“So both cerebris and GPUs can handle structured sparsity But GPUs are not designed to handle unstructured sparsity, whereas what I've just mentioned before is able to handle this unstructured sparsity.”
Disclosure
Cerebras avoids model parallelism in production due to communication overhead
“There's a lot of communication overhead with model parallelism. You have to share activation tensors, and that is why in this paper and, you know, in production, Cerebra's focus on data parallelism. So all of this is mentioned in the paper as well, but model p…”
Assertion Not checkable as stated
Lie: OpenAI uses Cerebras hardware internally for incident response and research
“So right now internally, they're using it for a lot of really critical use cases where the speed really, really matters. Like they're using it in like, Their incidents response teams, right? When there's an outage in their service, for example, every single se…”
Disclosure
Sean Lie: Cerebras uses OpenAI's internal AI tools for chip design
“We're also collaborating very closely with OpenAI, right, to use their tools to help us also continue to push what's possible in our chip design, in our software, and all that.”
Disclosure
Lie: Cerebras is currently sold out of all hardware capacity
“Right now we are basically, you know, sold out Of everything that we're building, right?”
Assertion Partly supported
Lie: Cerebras CS-4 doubles wafer power and bandwidth while halving latency
“We've designed this a modular platform that provides twice the amount of power to the wafer than we have in our previous generation. Twice the amount of interconnect bandwidth, half the latency.”
Assertion Supported
Feldman: Sam Altman and Ilya Sutskever invested in Cerebras' early rounds
“In 2016, we met with Sam Altman and Ilya Suskovard at OpenAI and they were an idea and we were PowerPoint, right? That's amazing. And what AI was doing was identifying cats in pictures. And I think they ended up investing in us, both of them and many of their …”
Assertion Not checkable as stated
Feldman: Non-compute bottlenecks masked Cerebras' 20-25x speedup for a hyperscaler
“We were doing work with, we're doing work with one of the hyperscalers and for a product they have, and they said, look, we used your system and it didn't make us that much faster. And so we said, well, that's a surprise because we're 20, 25 times faster on th…”
Insight
Feldman: Chip architecture begins by deciding what not to be good at
“One of the hardest things in computer architecture, and one of the first things you do is you decide what you're not going to be good at. I'm going to build chip. What am I not going to be good at? We're not going to be good at general purpose compute. We're n…”
Insight
Feldman: Pre-IPO startups should target investors primarily focused on public markets
“I think in later stages, as you get close to IPO, you're looking for a very different type of investor. You're looking for an investor who Primarily does public markets.”
Disclosure
Feldman: Cerebras serves Mistral AI's Le Chat assistant
“We serve Le Chat from Mistral.”
Assertion Supported
Cerebras weight streaming prunes up to 90% of data without accuracy loss
“And so memory, and so as MemoryX streams weights through SwarmX, it eliminates zero and near zero values, and so this reduces bandwidth requirements significantly, pruning up to 90% of the data while maintaining accuracy.”
Assertion Supported
Cerebras WSE-3 per-core SRAM eliminates central memory bandwidth bottlenecks
“So what Cerebrus has done for the wafer scale engine three is that instead of storing all these weights and values, weights and values off chip, Cerebrus stores everything on chip in SRAM. So every single one of the cores on the wafer scale engine three has it…”
Assertion Supported
Cerebras MemoryX scales to 2.4 petabytes to support 120-trillion-parameter AI models
“And you know, it scales from four terabytes to 2.4 petabytes, Supports models with up to 120 trillion parameters and then it utilizes DRAM and flash storage.”
Assertion Supported
Cerebras WSE-3 features 900,000 cores, 44GB SRAM, and 4 trillion transistors
“And so the wafer scale engine three, as I mentioned, 900,000 cores, 44 gigabytes of SRAM, four trillion transistors, and I do add a note here that the paper focuses on wafer scale engine two, and so the wafer scale engine three is, you know, just an upgraded v…”
Disclosure
Feldman: Cerebras trains models for Mayo Clinic, GSK, and US military
“We do a great deal of work with large enterprises in training, with Mayo Clinic, with GlaxoSmithKline, with the US military, with the Department of Energy, with our customers in the Middle East. We've trained leading models, language models in, in Arabic, in C…”
Assertion Supported
Cerebras weight streaming is exclusively for training, while inference runs on SRAM
“The memory X and swarm X, this whole waste streaming system is just used for is just used for training. So for inference, you just using the SRAM, you know, at 44 gigabytes on the chip, and then you can network multiple chips together to support larger models.”
Assertion Supported
Cerebras streams weights from MemoryX and computes updates externally
“And instead of storing all the weights that the compute units need on
[737] Sarah Chieng: On the compute unit, it's storing it externally in an external memory service. In this case, it's called memory X. And during training, these weights are streamed from me…”