Opinion
NVIDIA is extorting customers amid AI chip shortages, says Cerebras CEO
“I think they're now in a situation where they're extorting customers. They're extremely expensive. They're unable to ship. And that has, among other things, opened the door for many of us who have alternatives.”
Assertion Partly supported
Feldman: Cerebras is 15 to 20 times faster than GPUs at inference
“And right now we're the fastest at inference, not by a little bit, but by a lot. 1518, 20 X faster than GPUs.”
Prediction Not checkable as stated
Feldman: The market for slow AI inference will become zero
“How big is the market for slow search? It's zero. How big is the market for dial-up internet? It's zero. That's how big the market for slow inference will be.”
Assertion Supported
Feldman: Cerebras is the only pure-play public AI company
“We would be the first and only, for a period of time, AI pure play. We are the only company that, that you can, a hundred percent of the revenue, this exact market. There's no gaming, there's no graphics, there's no PC, that, this is it.”
Assertion Supported
Feldman: Cerebras signed a $20B+ OpenAI deal in 4.5 weeks
“For a, 20 plus billion dollar deal to do it in four and a half weeks was exceptional.”
Assertion Supported
Feldman: Cerebras clusters scale linearly up to 64 nodes
“The cluster we build keeps the parameters off chip in a parameter store, and it streams them in, and the result of this architecture is that, that we run strictly data parallel, which means even in a 64 node cluster, you run the exact same configuration on eac…”
Assertion Supported
Feldman: Cerebras can run a trillion-parameter model on a single system
“And that was an idea that came from supercomputing that we knew really well, that we could organize this so you could run an arbitrarily large, now a trillion parameter network on a single system.”
Prediction Not checkable as stated
AI shift to single-shot learning would doom NVIDIA and Cerebras hardware
“If we go to a type of model that doesn't require very much data, if we go to single-shot learning, right, NVIDIA's totally out of luck, right? Us too, everybody.”
Assertion Not checkable as stated
Feldman: Nvidia missed its own demand forecast ahead of AI crunch
“It's not just that Wall Street missed what NVIDIA would sell. NVIDIA missed it. They missed the forecast.”
Prediction Not checkable as stated
Feldman: Commercial AI will settle on 3B to 13B parameter models
“And so we're going to be down at three and at six billion and at thirteen billion, because that's, I can get pretty good, pretty darn good, and not break the bank with free inference.”
Disclosure
Feldman: Cerebras signed an agreement to deploy in AWS data centers
“And then in March, we signed an agreement with AWS, where we will be deployed in their data centers going forward”
Prediction Open · timeframe May 2029
Feldman: Cerebras will sell tens of thousands of third-generation systems
“It was like, you know, the first gen we might have sold a dozen. The second gen we probably sold 300, and now we're still going to sell tens of thousands in the third gen.”
Disclosure
Feldman: Cerebras spends $25k to $30k per engineer on AI tokens
“I would say that, that, you know, eight months ago, we weren't spending a thousand dollars in engineer on tokens, and we're probably at 25 or 30,000 right now, and it's ripping.”
Opinion
Feldman: Chinese open-source AI techniques force closed labs to innovate
“The open source community has sort of kept the interest alive and kept the flame going. And I think that, that the and pushed the closed source guys. I think the sort of techniques that we saw by some of the Chinese Like, whoa, we gotta stay ahead of that, rig…”
Assertion Not checkable as stated
Feldman: Sun Microsystems had 70 people dedicated to gaming benchmarks
“When our CTO was at Sun, they had a team of 70 whose job it was to gain benchmarks.”
Assertion Not checkable as stated
Feldman: Cerebras eliminates memory bandwidth bottlenecks using on-wafer SRAM
“We keep a huge amount of SRAM on the wafer. All right, and so there are no memory bandwidth problems ever. That also allows us to harvest sparsity, which is something that others really struggle with.”
Assertion Not checkable as stated
Feldman: Redistributing AI training takes one keystroke on Cerebras vs GPUs
“In March, we put seven GPT models in the open source community. Everybody else was putting one. Why? Because it's really hard to redistribute work across a GPU cluster. For us, it's one keystroke.”
Assertion Not checkable as stated
Feldman: Cerebras converged a model in 3.5 days after 60-day GPU failure
“We had a situation where they were trying to train on a GPU cluster and they were at 60 days and it wasn't converging and We stood it up, and three and a half days later, their model converged”
Prediction Not checkable as stated
Feldman: Chip market will use separate silicon for training and inference
“Now, whether you will have different silicon for training and for inference, I think you will.”
Insight
Feldman: Chip architects must solve hard problems generally, not guess layers
“The trick in architecture is to solve hard problems in a general way, so you don't have to rely on product management to sort of guess what, what's the next cool layer type, right?”
Assertion Supported
Running inference on an eight-GPU H100 system costs half a million dollars
“I mean, people using. Eight h, 100 to do inference on a big model. I mean, that's. Half a million dollars.”
Assertion Not checkable as stated
Feldman: Cerebras burned $8M monthly for two years before chips worked
“We had a period between about 2017, middle of 2017 and middle of 2019 where we couldn't build it. We were spending about eight million a month.”
Insight
Feldman: New chip architectures must start in supercomputing where speed trumps software maturity
“I think there's a path that has been laid down by new computer architectures, and often you begin in the supercomputer world, because those guys love speed, and they don't care if your software is immature, and so we sort of ran the table there.”
Assertion Supported
Feldman: UAE sovereign AI firm G42 placed a $1B order with Cerebras
“And we won a sovereign, a G-forty-two and they became a strategic partner and close friends and they placed a billion dollar order on us,”