Jul 23, 2026 · 1h 12m · mad
Cerebras CEO: Why GPUs Can't Do Fast Inference
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Cerebras Systems Co-Founder and CEO Andrew Feldman joins Matt Turck on The MAD Podcast to discuss the architectural limits of traditional GPUs, the engineering breakthroughs behind Cerebras' massive wafer-scale engine, and why ultra-fast AI inference speed is essential for the future of reasoning models and agentic workflows.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 20.3% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Guest immediately rejects the host's premise that CUDA remains a formidable moat, citing that leading models like Gemini and Anthropic are trained without it and that switching cloud providers requires just eight keystrokes.
Hardest push from Matt ▶ 16:53 Devil's advocate on AI market demand and VC fundingHost directly challenges the sustainability of chip demand, questioning whether orders from AI labs are artificially inflated and financed by venture capital and private equity.
Biggest teaching moment ▶ 42:46 Detailed breakdown of memory bandwidth and decode sequentialityGuest educates the host on memory bandwidth limitations by comparing the data movement for a single 70B parameter inference token to transferring 100 HD movies from memory to compute.
Matt holds his own ▶ 58:00 Host articulates disaggregated AWS Trainium and Cerebras architectureHost demonstrates acute technical domain expertise by accurately describing how AWS Trainium handles the parallel pre-fill step while Cerebras handles the sequential decode step in a disaggregated system.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| The AI Speed Revolution and Tokens Per Second | 3 | 4 | 1 | 1 | Host sets the stage with clarifying questions on speed metrics and offers an apt UX analogy regarding broadband. Guest explains tokens per second per user and why latency matters for agentic workflows. | |
| Navigating the AI Chip Landscape and Multi-Silicon Ecosystem | 5 | 5 | 2 | 2 | Host demonstrates strong industry knowledge by citing specific silicon names like Google TPU, AWS Trainium, and OpenAI's Jalapeno project with Broadcom. Guest outlines the ASIC landscape and highlights NVIDIA's acquisition of Grok as validation of Cerebras's vision. | |
| Geopolitics, China's AI Strategy, and Edge vs. Cloud Compute | 4 | 5 | 2 | 2 | Host prompts discussion on China's AI stack, citing Huawei and DeepSeek. Guest educates on geopolitical constraints, comparing French nuclear power and Chinese grid infrastructure against US data center power bottlenecks. | |
| AI Market Dynamics and Three Hidden Hardware Shortages | 5 | 6 | 3 | 5 | Host presents a strong devil's advocate pushback on market valuation crashes and potential artificial demand driven by VC-funded labs. Guest counters the bubble hypothesis by walking through three specific hardware supply bottlenecks: HBM DRAM, TSMC CoWoS packaging, and 3nm foundry capacity. | |
| The Cerebras Origin Story and Wafer-Scale Innovation | 4 | 6 | 1 | 2 | Host guides Guest through Cerebras's origin story and its 10-year journey, probing into technical burn rates. Guest delivers a masterclass on why SRAM on wafer-scale chips eliminates the data movement bottleneck that plagues traditional GPUs. | |
| Technical Deep Dive: Why Wafer-Scale Beats GPUs at Inference | 5 | 7 | 1 | 2 | Host asks Guest to explain token generation mechanisms simply and probes into pre-fill versus decode phases. Guest explains the sequential nature of decode, illustrating that moving model weights is equivalent to shuffling 100 HD movies per generated word. | |
| Reasoning Models, Verification, and Multimodal Performance | 5 | 6 | 1 | 2 | Host asks informed questions covering RL training, model-data parallelism, verification guardrails, and multimodality. Guest details why distributed compute on GPUs requires complex tensor parallelism, whereas Cerebras's large memory footprint simplifies training flow. | |
| Cloud Strategy, Data Center Expansion, and Eradicating the CUDA Moat | 5 | 6 | 4 | 4 | Host brings up NVIDIA's CUDA moat as a potential barrier to entry. Guest directly rejects the premise, pointing out that state-of-the-art models like Gemini and Anthropic are trained without CUDA and that switching in the cloud takes eight keystrokes. |