Jul 23, 2026 · 1h 12m · mad

Cerebras CEO: Why GPUs Can't Do Fast Inference

Andrew Feldman · 48m spoken Matt Turck · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Cerebras Systems Co-Founder and CEO Andrew Feldman joins Matt Turck on The MAD Podcast to discuss the architectural limits of traditional GPUs, the engineering breakthroughs behind Cerebras' massive wafer-scale engine, and why ultra-fast AI inference speed is essential for the future of reasoning models and agentic workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 20.3% of the talking time here. How this is scored →

Matt as informed peer 4.5 Guest teaching 5.6 Guest disagreement 1.9 Matt pushing back 2.5
05100:0015:0030:0045:001:00:001:31–4:36 · Matt as informed peer 3/10 The AI Speed Revolution and Tokens Per Second Host sets the stage with clarifying questions on speed metrics and offers an apt UX analogy regarding broadband. Guest explains tokens per second per user and why latency matters for agentic workflows.4:36–11:29 · Matt as informed peer 5/10 Navigating the AI Chip Landscape and Multi-Silicon Ecosystem Host demonstrates strong industry knowledge by citing specific silicon names like Google TPU, AWS Trainium, and OpenAI's Jalapeno project with Broadcom. Guest outlines the ASIC landscape and highlights NVIDIA's acquisition of Grok as validation of Cerebras's vision.11:29–15:01 · Matt as informed peer 4/10 Geopolitics, China's AI Strategy, and Edge vs. Cloud Compute Host prompts discussion on China's AI stack, citing Huawei and DeepSeek. Guest educates on geopolitical constraints, comparing French nuclear power and Chinese grid infrastructure against US data center power bottlenecks.15:01–25:37 · Matt as informed peer 5/10 AI Market Dynamics and Three Hidden Hardware Shortages Host presents a strong devil's advocate pushback on market valuation crashes and potential artificial demand driven by VC-funded labs. Guest counters the bubble hypothesis by walking through three specific hardware supply bottlenecks: HBM DRAM, TSMC CoWoS packaging, and 3nm foundry capacity.25:37–39:05 · Matt as informed peer 4/10 The Cerebras Origin Story and Wafer-Scale Innovation Host guides Guest through Cerebras's origin story and its 10-year journey, probing into technical burn rates. Guest delivers a masterclass on why SRAM on wafer-scale chips eliminates the data movement bottleneck that plagues traditional GPUs.39:05–45:02 · Matt as informed peer 5/10 Technical Deep Dive: Why Wafer-Scale Beats GPUs at Inference Host asks Guest to explain token generation mechanisms simply and probes into pre-fill versus decode phases. Guest explains the sequential nature of decode, illustrating that moving model weights is equivalent to shuffling 100 HD movies per generated word.45:02–53:50 · Matt as informed peer 5/10 Reasoning Models, Verification, and Multimodal Performance Host asks informed questions covering RL training, model-data parallelism, verification guardrails, and multimodality. Guest details why distributed compute on GPUs requires complex tensor parallelism, whereas Cerebras's large memory footprint simplifies training flow.53:50–1:12:20 · Matt as informed peer 5/10 Cloud Strategy, Data Center Expansion, and Eradicating the CUDA Moat Host brings up NVIDIA's CUDA moat as a potential barrier to entry. Guest directly rejects the premise, pointing out that state-of-the-art models like Gemini and Anthropic are trained without CUDA and that switching in the cloud takes eight keystrokes.1:31–4:36 · Guest teaching 4/10 The AI Speed Revolution and Tokens Per Second Host sets the stage with clarifying questions on speed metrics and offers an apt UX analogy regarding broadband. Guest explains tokens per second per user and why latency matters for agentic workflows.4:36–11:29 · Guest teaching 5/10 Navigating the AI Chip Landscape and Multi-Silicon Ecosystem Host demonstrates strong industry knowledge by citing specific silicon names like Google TPU, AWS Trainium, and OpenAI's Jalapeno project with Broadcom. Guest outlines the ASIC landscape and highlights NVIDIA's acquisition of Grok as validation of Cerebras's vision.11:29–15:01 · Guest teaching 5/10 Geopolitics, China's AI Strategy, and Edge vs. Cloud Compute Host prompts discussion on China's AI stack, citing Huawei and DeepSeek. Guest educates on geopolitical constraints, comparing French nuclear power and Chinese grid infrastructure against US data center power bottlenecks.15:01–25:37 · Guest teaching 6/10 AI Market Dynamics and Three Hidden Hardware Shortages Host presents a strong devil's advocate pushback on market valuation crashes and potential artificial demand driven by VC-funded labs. Guest counters the bubble hypothesis by walking through three specific hardware supply bottlenecks: HBM DRAM, TSMC CoWoS packaging, and 3nm foundry capacity.25:37–39:05 · Guest teaching 6/10 The Cerebras Origin Story and Wafer-Scale Innovation Host guides Guest through Cerebras's origin story and its 10-year journey, probing into technical burn rates. Guest delivers a masterclass on why SRAM on wafer-scale chips eliminates the data movement bottleneck that plagues traditional GPUs.39:05–45:02 · Guest teaching 7/10 Technical Deep Dive: Why Wafer-Scale Beats GPUs at Inference Host asks Guest to explain token generation mechanisms simply and probes into pre-fill versus decode phases. Guest explains the sequential nature of decode, illustrating that moving model weights is equivalent to shuffling 100 HD movies per generated word.45:02–53:50 · Guest teaching 6/10 Reasoning Models, Verification, and Multimodal Performance Host asks informed questions covering RL training, model-data parallelism, verification guardrails, and multimodality. Guest details why distributed compute on GPUs requires complex tensor parallelism, whereas Cerebras's large memory footprint simplifies training flow.53:50–1:12:20 · Guest teaching 6/10 Cloud Strategy, Data Center Expansion, and Eradicating the CUDA Moat Host brings up NVIDIA's CUDA moat as a potential barrier to entry. Guest directly rejects the premise, pointing out that state-of-the-art models like Gemini and Anthropic are trained without CUDA and that switching in the cloud takes eight keystrokes.1:31–4:36 · Guest disagreement 1/10 The AI Speed Revolution and Tokens Per Second Host sets the stage with clarifying questions on speed metrics and offers an apt UX analogy regarding broadband. Guest explains tokens per second per user and why latency matters for agentic workflows.4:36–11:29 · Guest disagreement 2/10 Navigating the AI Chip Landscape and Multi-Silicon Ecosystem Host demonstrates strong industry knowledge by citing specific silicon names like Google TPU, AWS Trainium, and OpenAI's Jalapeno project with Broadcom. Guest outlines the ASIC landscape and highlights NVIDIA's acquisition of Grok as validation of Cerebras's vision.11:29–15:01 · Guest disagreement 2/10 Geopolitics, China's AI Strategy, and Edge vs. Cloud Compute Host prompts discussion on China's AI stack, citing Huawei and DeepSeek. Guest educates on geopolitical constraints, comparing French nuclear power and Chinese grid infrastructure against US data center power bottlenecks.15:01–25:37 · Guest disagreement 3/10 AI Market Dynamics and Three Hidden Hardware Shortages Host presents a strong devil's advocate pushback on market valuation crashes and potential artificial demand driven by VC-funded labs. Guest counters the bubble hypothesis by walking through three specific hardware supply bottlenecks: HBM DRAM, TSMC CoWoS packaging, and 3nm foundry capacity.25:37–39:05 · Guest disagreement 1/10 The Cerebras Origin Story and Wafer-Scale Innovation Host guides Guest through Cerebras's origin story and its 10-year journey, probing into technical burn rates. Guest delivers a masterclass on why SRAM on wafer-scale chips eliminates the data movement bottleneck that plagues traditional GPUs.39:05–45:02 · Guest disagreement 1/10 Technical Deep Dive: Why Wafer-Scale Beats GPUs at Inference Host asks Guest to explain token generation mechanisms simply and probes into pre-fill versus decode phases. Guest explains the sequential nature of decode, illustrating that moving model weights is equivalent to shuffling 100 HD movies per generated word.45:02–53:50 · Guest disagreement 1/10 Reasoning Models, Verification, and Multimodal Performance Host asks informed questions covering RL training, model-data parallelism, verification guardrails, and multimodality. Guest details why distributed compute on GPUs requires complex tensor parallelism, whereas Cerebras's large memory footprint simplifies training flow.53:50–1:12:20 · Guest disagreement 4/10 Cloud Strategy, Data Center Expansion, and Eradicating the CUDA Moat Host brings up NVIDIA's CUDA moat as a potential barrier to entry. Guest directly rejects the premise, pointing out that state-of-the-art models like Gemini and Anthropic are trained without CUDA and that switching in the cloud takes eight keystrokes.1:31–4:36 · Matt pushing back 1/10 The AI Speed Revolution and Tokens Per Second Host sets the stage with clarifying questions on speed metrics and offers an apt UX analogy regarding broadband. Guest explains tokens per second per user and why latency matters for agentic workflows.4:36–11:29 · Matt pushing back 2/10 Navigating the AI Chip Landscape and Multi-Silicon Ecosystem Host demonstrates strong industry knowledge by citing specific silicon names like Google TPU, AWS Trainium, and OpenAI's Jalapeno project with Broadcom. Guest outlines the ASIC landscape and highlights NVIDIA's acquisition of Grok as validation of Cerebras's vision.11:29–15:01 · Matt pushing back 2/10 Geopolitics, China's AI Strategy, and Edge vs. Cloud Compute Host prompts discussion on China's AI stack, citing Huawei and DeepSeek. Guest educates on geopolitical constraints, comparing French nuclear power and Chinese grid infrastructure against US data center power bottlenecks.15:01–25:37 · Matt pushing back 5/10 AI Market Dynamics and Three Hidden Hardware Shortages Host presents a strong devil's advocate pushback on market valuation crashes and potential artificial demand driven by VC-funded labs. Guest counters the bubble hypothesis by walking through three specific hardware supply bottlenecks: HBM DRAM, TSMC CoWoS packaging, and 3nm foundry capacity.25:37–39:05 · Matt pushing back 2/10 The Cerebras Origin Story and Wafer-Scale Innovation Host guides Guest through Cerebras's origin story and its 10-year journey, probing into technical burn rates. Guest delivers a masterclass on why SRAM on wafer-scale chips eliminates the data movement bottleneck that plagues traditional GPUs.39:05–45:02 · Matt pushing back 2/10 Technical Deep Dive: Why Wafer-Scale Beats GPUs at Inference Host asks Guest to explain token generation mechanisms simply and probes into pre-fill versus decode phases. Guest explains the sequential nature of decode, illustrating that moving model weights is equivalent to shuffling 100 HD movies per generated word.45:02–53:50 · Matt pushing back 2/10 Reasoning Models, Verification, and Multimodal Performance Host asks informed questions covering RL training, model-data parallelism, verification guardrails, and multimodality. Guest details why distributed compute on GPUs requires complex tensor parallelism, whereas Cerebras's large memory footprint simplifies training flow.53:50–1:12:20 · Matt pushing back 4/10 Cloud Strategy, Data Center Expansion, and Eradicating the CUDA Moat Host brings up NVIDIA's CUDA moat as a potential barrier to entry. Guest directly rejects the premise, pointing out that state-of-the-art models like Gemini and Anthropic are trained without CUDA and that switching in the cloud takes eight keystrokes.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 38.1% · guest 61.9%0:00 · Matt 38.1% · guest 61.9%3:00 · Matt 30.5% · guest 69.5%3:00 · Matt 30.5% · guest 69.5%6:00 · Matt 16.2% · guest 83.8%6:00 · Matt 16.2% · guest 83.8%9:00 · Matt 20.3% · guest 79.7%9:00 · Matt 20.3% · guest 79.7%12:00 · Matt 25.5% · guest 74.5%12:00 · Matt 25.5% · guest 74.5%15:00 · Matt 44.3% · guest 55.7%15:00 · Matt 44.3% · guest 55.7%18:00 · Matt 9.9% · guest 90.1%18:00 · Matt 9.9% · guest 90.1%21:00 · Matt 21.2% · guest 78.8%21:00 · Matt 21.2% · guest 78.8%24:00 · Matt 31.8% · guest 68.2%24:00 · Matt 31.8% · guest 68.2%27:00 · Matt 16.9% · guest 83.1%27:00 · Matt 16.9% · guest 83.1%30:00 · Matt 14.5% · guest 85.5%30:00 · Matt 14.5% · guest 85.5%33:00 · Matt 3.7% · guest 96.3%33:00 · Matt 3.7% · guest 96.3%36:00 · Matt 3.5% · guest 96.5%36:00 · Matt 3.5% · guest 96.5%39:00 · Matt 31% · guest 69%39:00 · Matt 31% · guest 69%42:00 · Matt 3.6% · guest 96.4%42:00 · Matt 3.6% · guest 96.4%45:00 · Matt 17.6% · guest 82.4%45:00 · Matt 17.6% · guest 82.4%48:00 · Matt 20.9% · guest 79.1%48:00 · Matt 20.9% · guest 79.1%51:00 · Matt 14.7% · guest 85.3%51:00 · Matt 14.7% · guest 85.3%54:00 · Matt 23.7% · guest 76.3%54:00 · Matt 23.7% · guest 76.3%57:00 · Matt 38% · guest 62%57:00 · Matt 38% · guest 62%1:00:00 · Matt 14.4% · guest 85.6%1:00:00 · Matt 14.4% · guest 85.6%1:03:00 · Matt 13.3% · guest 86.7%1:03:00 · Matt 13.3% · guest 86.7%1:06:00 · Matt 12% · guest 88%1:06:00 · Matt 12% · guest 88%1:09:00 · Matt 8.2% · guest 91.8%1:09:00 · Matt 8.2% · guest 91.8%1:12:00 · Matt 88.1% · guest 11.9%1:12:00 · Matt 88.1% · guest 11.9%
Sharpest disagreement ▶ 1:01:15 Direct rejection of CUDA moat narrative

Guest immediately rejects the host's premise that CUDA remains a formidable moat, citing that leading models like Gemini and Anthropic are trained without it and that switching cloud providers requires just eight keystrokes.

Hardest push from Matt ▶ 16:53 Devil's advocate on AI market demand and VC funding

Host directly challenges the sustainability of chip demand, questioning whether orders from AI labs are artificially inflated and financed by venture capital and private equity.

Biggest teaching moment ▶ 42:46 Detailed breakdown of memory bandwidth and decode sequentiality

Guest educates the host on memory bandwidth limitations by comparing the data movement for a single 70B parameter inference token to transferring 100 HD movies from memory to compute.

Matt holds his own ▶ 58:00 Host articulates disaggregated AWS Trainium and Cerebras architecture

Host demonstrates acute technical domain expertise by accurately describing how AWS Trainium handles the parallel pre-fill step while Cerebras handles the sequential decode step in a disaggregated system.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The AI Speed Revolution and Tokens Per Second 3411 Host sets the stage with clarifying questions on speed metrics and offers an apt UX analogy regarding broadband. Guest explains tokens per second per user and why latency matters for agentic workflows.
Navigating the AI Chip Landscape and Multi-Silicon Ecosystem 5522 Host demonstrates strong industry knowledge by citing specific silicon names like Google TPU, AWS Trainium, and OpenAI's Jalapeno project with Broadcom. Guest outlines the ASIC landscape and highlights NVIDIA's acquisition of Grok as validation of Cerebras's vision.
Geopolitics, China's AI Strategy, and Edge vs. Cloud Compute 4522 Host prompts discussion on China's AI stack, citing Huawei and DeepSeek. Guest educates on geopolitical constraints, comparing French nuclear power and Chinese grid infrastructure against US data center power bottlenecks.
AI Market Dynamics and Three Hidden Hardware Shortages 5635 Host presents a strong devil's advocate pushback on market valuation crashes and potential artificial demand driven by VC-funded labs. Guest counters the bubble hypothesis by walking through three specific hardware supply bottlenecks: HBM DRAM, TSMC CoWoS packaging, and 3nm foundry capacity.
The Cerebras Origin Story and Wafer-Scale Innovation 4612 Host guides Guest through Cerebras's origin story and its 10-year journey, probing into technical burn rates. Guest delivers a masterclass on why SRAM on wafer-scale chips eliminates the data movement bottleneck that plagues traditional GPUs.
Technical Deep Dive: Why Wafer-Scale Beats GPUs at Inference 5712 Host asks Guest to explain token generation mechanisms simply and probes into pre-fill versus decode phases. Guest explains the sequential nature of decode, illustrating that moving model weights is equivalent to shuffling 100 HD movies per generated word.
Reasoning Models, Verification, and Multimodal Performance 5612 Host asks informed questions covering RL training, model-data parallelism, verification guardrails, and multimodality. Guest details why distributed compute on GPUs requires complex tensor parallelism, whereas Cerebras's large memory footprint simplifies training flow.
Cloud Strategy, Data Center Expansion, and Eradicating the CUDA Moat 5644 Host brings up NVIDIA's CUDA moat as a potential barrier to entry. Guest directly rejects the premise, pointing out that state-of-the-art models like Gemini and Anthropic are trained without CUDA and that switching in the cloud takes eight keystrokes.

Statements from this episode (16)

Insight
Feldman: Tokens per second per user is the right AI speed metric
“The right metric is tokens per second per user. That that's how fast you get the first token all the way through the last token in, in your response.”
Andrew Feldman Jul 23, 2026 ▶ 2:40
Assertion Contradicted
Feldman: Cerebras sales were 10x higher than Groq's at acquisition
“And we were the fastest at it, and the largest, and, you know, our sales were more than 10 times the Grox, and they paid twenty billion dollars for the number two collector.”
Andrew Feldman Jul 23, 2026 ▶ 9:01
Disclosure
Feldman: Cerebras signed an OpenAI compute deal worth over $20 billion
“Remember, we did a huge deal. This is probably the largest deals in Silicon Valley history north of twenty billion dollars.”
Andrew Feldman Jul 23, 2026 ▶ 9:49
Opinion
Feldman: Chinese open-source AI models trail GPT, Anthropic, and Gemini
“They are behind in chips. But their approach was at the next level is open source models where they're producing some extraordinary models. Not as good as GPT or Anthropic or Google's Gemini, but very good.”
Andrew Feldman Jul 23, 2026 ▶ 12:59
Assertion Not checkable as stated
Feldman: Agentic AI workflows are driving CPU demand through the roof
“And so, as we do more and more AI work, and more and more agentic work, we're making more and more calls to CPUs, and therefore the demand for CPUs is through the roof.”
Andrew Feldman Jul 23, 2026 ▶ 25:16
Insight
Feldman: AI inference is bottlenecked by data movement, causing GPU slowness
“In inference in AI, it's the exact opposite. You move a huge amount of data, all the weights, from memory to compute, and you need one calculation to generate the next word. And then you have to do it again. So all the time is dominated by the movement of data…”
Andrew Feldman Jul 23, 2026 ▶ 32:16
Assertion Supported
Feldman: Cerebras built a 46,000 square millimeter wafer-scale chip
“The biggest chip that had ever been built before us was 800 square millimeters. 840 to be exact. And this is 46,000.”
Andrew Feldman Jul 23, 2026 ▶ 33:41
Disclosure
Feldman: Cerebras burned $8M monthly for 18 months before building successful chips
“And we had a, an 18 month period where we were spending eight million a month and we couldn't build them.”
Andrew Feldman Jul 23, 2026 ▶ 34:19
Assertion Partly supported
Feldman: Cerebras completed the largest semiconductor IPO in history on May 14th
“And so when we rang the bell and we went public on May 14th this year and the largest semiconductor IPO in history and we did something unusual.”
Andrew Feldman Jul 23, 2026 ▶ 37:56
Assertion Not checkable as stated
Feldman: GPUs suffer from high failure rates and infant mortality
“The JPs have a huge failure rate, so I'm sure you guys have spoken about this. Infant mortality is enormous, and they fail all the time.”
Andrew Feldman Jul 23, 2026 ▶ 40:08
Assertion Supported
Feldman: Cerebras moves weights to compute ~2,500x faster than standard GPUs
“And so the speed of moving waits to compute is about two and a half thousand times faster here than on a Wilben GP.”
Andrew Feldman Jul 23, 2026 ▶ 44:43
Disclosure
Feldman: Cerebras serves second-tier AI labs for model training
“We do RL and we do traditional training too. Not for the largest models, for the largest lab, but for the next tier.”
Andrew Feldman Jul 23, 2026 ▶ 45:43
Assertion Not checkable as stated
Feldman: Leading AI labs paused video generation development due to compute costs
“Obviously, what follows that Is video, because a video is just a collection of images. But that takes an enormous amount of compute right now. And that's one of the reasons it's been sort of set aside by the leading labs. So unbelievably computation intensive.”
Andrew Feldman Jul 23, 2026 ▶ 53:29
Disclosure
Feldman: Cerebras signed a 760-megawatt multi-year compute deal with OpenAI
“The deal is 760 megawatts, 250 megawatts in 26 on a multi-year lease. An additional 250 megawatts in 27, on a multi-year lease, and an additional in 28, a multi-year lease.”
Andrew Feldman Jul 23, 2026 ▶ 56:26
Assertion Not checkable as stated
Feldman: Nvidia CUDA lost 70% of frontier AI model training market share
“I think two years ago every state of the art model was trained in a Cuda flow. And right now, Gemini is trained without Cuda. Anthropical is trained without Cuda. Open AI as strange as could. So in a one or two year period, they lost 70% share. Of training mod…”
Andrew Feldman Jul 23, 2026 ▶ 1:01:20
Opinion
Feldman: AI is causing irreparable damage to the SaaS business model
“I think the business of dashboarding and the business, the AI doesn't the damages doing the SAS is, I think, unreparable. You could ask your AI, build me a tool like Salesforce. 30 seconds later, you have a working tool.”
Andrew Feldman Jul 23, 2026 ▶ 1:10:46
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.