Sep 7, 2023 · 30m · no-priors
No Priors Ep. 31 | With Cerebras CEO Andrew Feldman
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of No Priors, hosts Elad Gil and Sarah Guo interview Cerebras Systems CEO Andrew Feldman to discuss wafer-scale semiconductor architecture, solutions to the global GPU compute crunch, and the evolving economic realities of training and deploying foundation AI models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 19.1% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Andrew aggressively frames Nvidia's market dominance as extortionate pricing and artificial dependency, drawing analogies to historic customer resentment against Intel.
Hardest push from the hosts ▶ 21:11 Sarah presses on internal bets between training and inferenceSarah directly questions Andrew on whether Cerebras is hedging its core architecture by making internal bets to separate training silicon from generative inference silicon.
Biggest teaching moment ▶ 15:46 Masterclass on chip development capital expenditure vs SaaSAndrew educates the audience and hosts on the brutal economic reality of chip manufacturing, contrasting fast SaaS prototyping with multi-year, multi-million dollar tapeout cycles.
The host holds their own ▶ 25:39 Elad connects fab forecasting errors to Nvidia's crypto inventory crashElad brings in sharp market historical context, highlighting how Nvidia suffered severe supply over-allocation during the prior crypto downturn to support Andrew's forecasting argument.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| The One Hundred Million Dollar Cerebras G42 Partnership | 4 | 3 | 1 | 0 | The conversation opens warmly with Sarah reminiscing about Cerebras's earliest 2015 pitch meetings in her office. Andrew details the initial 36-exaflop G42 partnership and early architectural bets on wafer-scale chips. The dynamic is collaborative and celebratory rather than adversarial. | |
| Real-World Training Efficiency versus Synthetic Hardware Benchmarks | 3 | 6 | 4 | 1 | When asked about performance benchmarks, Andrew rejects synthetic benchmarks entirely, citing his background at AMD and Sun where entire teams were dedicated to gaming them. He educates the hosts on real-world training times, data parallelism, and how wafer-scale SRAM avoids GPU distributed compute overhead. | |
| Disaggregating Compute and Memory to Outperform Standard GPUs | 3 | 6 | 2 | 0 | Sarah prompts Andrew to explain the architectural differences between traditional GPUs and Cerebras hardware for large-scale ML. Andrew explains how coupling memory to compute on GPU interposers forces inefficient hardware buying, whereas disaggregating memory allows running arbitrary parameter sizes. | |
| Multilingual Foundation Models, Cultural Nuance, and Tokenization | 5 | 4 | 1 | 1 | Andrew highlights serving costs and multilingual models, particularly Arabic and low-resource languages. Elad and Sarah engage deeply on cultural encoding, byte tokenization biases, and alignment mechanics in pre-training datasets. | |
| Engineering and Financial Realities of Starting Chip Companies | 3 | 7 | 2 | 0 | Sarah asks about the structural hurdles of founding semiconductor startups. Andrew delivers a comprehensive breakdown of the sixty-million-dollar tapeout cycle, zero-margin-for-error QA ratios, and why hardware cannot operate on rapid SaaS feedback loops. | |
| The AI Accelerator Landscape and Challenging NVIDIA Dominance | 5 | 5 | 4 | 2 | Andrew pulls no punches regarding Nvidia, accusing them of exploiting monopoly conditions and extorting customers. Sarah probes Andrew on whether general sparse linear algebra can outlast specialized transformer architectures or separate training and inference silicon. | |
| Supply Chain Inflexibility and the Ongoing GPU Shortage | 5 | 6 | 2 | 1 | Andrew outlines the inflexible economics of semiconductor fabs, noting that TSMC twenty-billion-dollar fabs cannot pivot quickly when forecasts are blown. Elad adds industry context by noting Nvidia's previous inventory write-downs during the crypto market crash. | |
| Economic Limits of Foundation Model Scaling and Deployment | 4 | 4 | 1 | 1 | The hosts and guest discuss economic scaling limits, highlighting that inference costs will keep production models in the 3B-13B parameter range. Sarah and Andrew discuss high-value enterprise proprietary datasets from companies like Bloomberg and Reuters as the next major differentiator. | |
| Podcast Conclusion, Social Channels, and Episode Transcripts | 0 | 0 | 0 | 0 | Sarah delivers the closing housekeeping notes, social media channels, and transcript website links. |