Cerebras CEO Andrew Feldman recounts the timeline of closing an unprecedented AI inference hardware contract with OpenAI.
Opinion
NVIDIA is extorting customers amid AI chip shortages, says Cerebras CEO
“I think they're now in a situation where they're extorting customers. They're extremely expensive. They're unable to ship. And that has, among other things, opened the door for many of us who have alternatives.”
Assertion Partly supported
Feldman: Cerebras is 15 to 20 times faster than GPUs at inference
“And right now we're the fastest at inference, not by a little bit, but by a lot. 1518, 20 X faster than GPUs.”
Prediction Not checkable as stated
Feldman: The market for slow AI inference will become zero
“How big is the market for slow search? It's zero. How big is the market for dial-up internet? It's zero. That's how big the market for slow inference will be.”
Assertion Supported
Feldman: Cerebras is the only pure-play public AI company
“We would be the first and only, for a period of time, AI pure play. We are the only company that, that you can, a hundred percent of the revenue, this exact market. There's no gaming, there's no graphics, there's no PC, that, this is it.”
Assertion Supported
Feldman: Cerebras clusters scale linearly up to 64 nodes
“The cluster we build keeps the parameters off chip in a parameter store, and it streams them in, and the result of this architecture is that, that we run strictly data parallel, which means even in a 64 node cluster, you run the exact same configuration on eac…”
Assertion Supported
Feldman: Cerebras can run a trillion-parameter model on a single system
“And that was an idea that came from supercomputing that we knew really well, that we could organize this so you could run an arbitrarily large, now a trillion parameter network on a single system.”