Andrew Feldman

Co-Founder and CEO, Cerebras Systems · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutive@andrewdfeldman ↗cerebras.ai ↗Wikipedia ↗

Andrew Feldman leads Cerebras Systems, an AI hardware and cloud company known for creating the Wafer-Scale Engine for AI training and inference. Previously, he was the co-founder and CEO of SeaMicro, an energy-efficient microserver startup acquired by AMD in 2012 for $357 million.

16statements → 8claims → 4claims resolved → 50%fully supported → 4.06/5average certainty → 2.62/5average debate potential → 4.2/5argument clarity · the sources → 1said about them ↓

2 supported 1 partly supported 1 contradicted 4 not checkable as stated how the 8 claims stand · each chip opens the sources

8 assertions · 2 opinions · 2 insights · 4 disclosures · every statement was checked. The predictions and assertions are the 8 claims: statements the public record can support or contradict. 4 are resolved, and 4 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Andrew argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Feldman: Cerebras moves weights to compute ~2,500x faster than standard GPUs
“And so the speed of moving waits to compute is about two and a half thousand times faster here than on a Wilben GP.”
Andrew Feldman Jul 23, 2026 ▶ 44:43 Cerebras CEO: Why GPUs Can't Do Fast Inference

Their most notable contradicted claim

Assertion Contradicted
Feldman: Cerebras sales were 10x higher than Groq's at acquisition
“And we were the fastest at it, and the largest, and, you know, our sales were more than 10 times the Grox, and they paid twenty billion dollars for the number two collector.”
Andrew Feldman Jul 23, 2026 ▶ 9:01 Cerebras CEO: Why GPUs Can't Do Fast Inference

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
none yet certainty 3
50% certainty 4
100% certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Argument clarity: do they answer the question? how? →

4.2 / 5 directness 4.3 · coherence 4.4 · precision 4 · compression 3.8

redirected or did not address 1 of 12 assessed questions (8%). Watch them ▸

This is a score against a rubric. It is not a rank. Every host question → answer exchange is scored with names hidden on directness, coherence, precision and compression, 1–5 each, on meaning alone: disfluencies are ignored, and only raw unedited episodes count. This is the score that measures thought. Every scored exchange, scores shown → · The rubric and its checks →

How they sound: speaking style how? →

228 words/min while actually speaking · 26.8 um and uh per 1k words

Measured by listening to the audio itself: 8,195 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Andrew Feldman said on the MAD Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Feldman: Nvidia CUDA lost 70% of frontier AI model training market share
“I think two years ago every state of the art model was trained in a Cuda flow. And right now, Gemini is trained without Cuda. Anthropical is trained without Cuda. Open AI as strange as could. So in a one or two year period, they lost 70% share. Of training mod…”
Andrew Feldman Jul 23, 2026 ▶ 1:01:20 Cerebras CEO: Why GPUs Can't Do Fast Inference
Opinion
Feldman: AI is causing irreparable damage to the SaaS business model
“I think the business of dashboarding and the business, the AI doesn't the damages doing the SAS is, I think, unreparable. You could ask your AI, build me a tool like Salesforce. 30 seconds later, you have a working tool.”
Andrew Feldman Jul 23, 2026 ▶ 1:10:46 Cerebras CEO: Why GPUs Can't Do Fast Inference
Assertion Contradicted
Feldman: Cerebras sales were 10x higher than Groq's at acquisition
“And we were the fastest at it, and the largest, and, you know, our sales were more than 10 times the Grox, and they paid twenty billion dollars for the number two collector.”
Andrew Feldman Jul 23, 2026 ▶ 9:01 Cerebras CEO: Why GPUs Can't Do Fast Inference
Disclosure
Feldman: Cerebras signed an OpenAI compute deal worth over $20 billion
“Remember, we did a huge deal. This is probably the largest deals in Silicon Valley history north of twenty billion dollars.”
Andrew Feldman Jul 23, 2026 ▶ 9:49 Cerebras CEO: Why GPUs Can't Do Fast Inference
Opinion
Feldman: Chinese open-source AI models trail GPT, Anthropic, and Gemini
“They are behind in chips. But their approach was at the next level is open source models where they're producing some extraordinary models. Not as good as GPT or Anthropic or Google's Gemini, but very good.”
Andrew Feldman Jul 23, 2026 ▶ 12:59 Cerebras CEO: Why GPUs Can't Do Fast Inference
Assertion Not checkable as stated
Feldman: Agentic AI workflows are driving CPU demand through the roof
“And so, as we do more and more AI work, and more and more agentic work, we're making more and more calls to CPUs, and therefore the demand for CPUs is through the roof.”
Andrew Feldman Jul 23, 2026 ▶ 25:16 Cerebras CEO: Why GPUs Can't Do Fast Inference
Insight
Feldman: AI inference is bottlenecked by data movement, causing GPU slowness
“In inference in AI, it's the exact opposite. You move a huge amount of data, all the weights, from memory to compute, and you need one calculation to generate the next word. And then you have to do it again. So all the time is dominated by the movement of data…”
Andrew Feldman Jul 23, 2026 ▶ 32:16 Cerebras CEO: Why GPUs Can't Do Fast Inference
Assertion Supported
Feldman: Cerebras moves weights to compute ~2,500x faster than standard GPUs
“And so the speed of moving waits to compute is about two and a half thousand times faster here than on a Wilben GP.”
Andrew Feldman Jul 23, 2026 ▶ 44:43 Cerebras CEO: Why GPUs Can't Do Fast Inference
Assertion Not checkable as stated
Feldman: Leading AI labs paused video generation development due to compute costs
“Obviously, what follows that Is video, because a video is just a collection of images. But that takes an enormous amount of compute right now. And that's one of the reasons it's been sort of set aside by the leading labs. So unbelievably computation intensive.”
Andrew Feldman Jul 23, 2026 ▶ 53:29 Cerebras CEO: Why GPUs Can't Do Fast Inference
Disclosure
Feldman: Cerebras signed a 760-megawatt multi-year compute deal with OpenAI
“The deal is 760 megawatts, 250 megawatts in 26 on a multi-year lease. An additional 250 megawatts in 27, on a multi-year lease, and an additional in 28, a multi-year lease.”
Andrew Feldman Jul 23, 2026 ▶ 56:26 Cerebras CEO: Why GPUs Can't Do Fast Inference
Insight
Feldman: Tokens per second per user is the right AI speed metric
“The right metric is tokens per second per user. That that's how fast you get the first token all the way through the last token in, in your response.”
Andrew Feldman Jul 23, 2026 ▶ 2:40 Cerebras CEO: Why GPUs Can't Do Fast Inference
Disclosure
Feldman: Cerebras burned $8M monthly for 18 months before building successful chips
“And we had a, an 18 month period where we were spending eight million a month and we couldn't build them.”
Andrew Feldman Jul 23, 2026 ▶ 34:19 Cerebras CEO: Why GPUs Can't Do Fast Inference
Assertion Not checkable as stated
Feldman: GPUs suffer from high failure rates and infant mortality
“The JPs have a huge failure rate, so I'm sure you guys have spoken about this. Infant mortality is enormous, and they fail all the time.”
Andrew Feldman Jul 23, 2026 ▶ 40:08 Cerebras CEO: Why GPUs Can't Do Fast Inference
Assertion Supported
Feldman: Cerebras built a 46,000 square millimeter wafer-scale chip
“The biggest chip that had ever been built before us was 800 square millimeters. 840 to be exact. And this is 46,000.”
Andrew Feldman Jul 23, 2026 ▶ 33:41 Cerebras CEO: Why GPUs Can't Do Fast Inference
Assertion Partly supported
Feldman: Cerebras completed the largest semiconductor IPO in history on May 14th
“And so when we rang the bell and we went public on May 14th this year and the largest semiconductor IPO in history and we did something unusual.”
Andrew Feldman Jul 23, 2026 ▶ 37:56 Cerebras CEO: Why GPUs Can't Do Fast Inference
Disclosure
Feldman: Cerebras serves second-tier AI labs for model training
“We do RL and we do traditional training too. Not for the largest models, for the largest lab, but for the next tier.”
Andrew Feldman Jul 23, 2026 ▶ 45:43 Cerebras CEO: Why GPUs Can't Do Fast Inference

The other half of the tape: Andrew Feldman's own voice is left out of every number here. Other people bring the name up 1 time in 1 episode on the MAD Podcast. every mention, with the transcript →

Who brings them up most Matt Turck 1

Every mention by year

tap a year for its mentions
0011112026episodesmentions
0112026episodes it came up in
000.50.5112026episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
Cerebras CEO: Why GPUs Can't Do Fast Inference Jul 23, 2026 48m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.