Oct 1, 2025 · 29m · latent-space

⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras

Andrew Feldman · 19m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Cerebras CEO Andrew Feldman discusses the company's $1.1 billion fundraise and details how their specialized wafer-scale silicon overcomes memory bandwidth bottlenecks to power the next generation of real-time AI inference.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.1 Guest teaching 3.9 Guest disagreement 1.4 The hosts pushing back 1.8
05100:0010:0020:000:03–3:36 · The hosts as informed peer 3/10 Announcing Cerebras' $1.1 Billion Fundraise The host opens the episode celebrating Cerebras' $1.1B round and establishes his connections across the AI ecosystem and Cognition. Feldman provides background context on early meetings with OpenAI founders and the progression of their hardware generations.3:37–8:41 · The hosts as informed peer 4/10 Selecting Early vs. Late-Stage Investors Feldman walks through memory architectures, using the cup and straw analogy to explain why on-chip SRAM memory bandwidth overcomes the DRAM and HBM bottlenecks seen in traditional GPUs. The host listens and engages lightly with background context on benchmarking.8:41–12:06 · The hosts as informed peer 6/10 Wafer-Scale Strategy and Continuous Benchmarking The host demonstrates domain familiarity by pointing out interconnect complexity and offering the metaphor of MLPerf as the Olympics versus Artificial Analysis as daily live traffic. Feldman details the hardware mess of wiring thousands of small SRAM chips versus wafer-scale silicon.12:11–16:44 · The hosts as informed peer 7/10 Balancing Training Workloads and Inference Demand The host showcases deep technical insight by quoting internal Google TPU training-to-inference ratios (2:3) to estimate Cerebras' 1:5 ratio, which Feldman confirms. They discuss why latency directly causes customer churn in LLM applications.16:44–18:52 · The hosts as informed peer 6/10 Foundational Chip Architecture Decisions and Linear Algebra The host brings up recent work in hybrid architectures and attention variants. Feldman details their foundational 2016 design choice to optimize sparse linear algebra rather than hard-coding 3x3 convolutions, which future-proofed the chip for transformers and diffusion.18:53–21:30 · The hosts as informed peer 5/10 Expanding Cerebras Cloud and Developer Infrastructure The host directly probes whether Cerebras Coder puts the company into direct competition with its own developer customers. Feldman clarifies their strategy, explicitly stating Cerebras has no ambition to build an IDE or compete with developer toolmakers.21:31–25:30 · The hosts as informed peer 6/10 Critical Engineering Challenges: Power, Routing, and I/O Bottlenecks Feldman discusses overlooked datacenter realities including huge power draw and upstream token routing. The host pushes back to clarify routing responsibilities and correctly diagnoses a customer benchmark failure as I/O-bound rather than compute-bound.25:32–28:40 · The hosts as informed peer 4/10 Future Vision: The Exponential Rise of Real-Time AI Inference When the host asks for a 10-year grandmaster plan, Feldman dismisses long-range roadmaps as impractical in fast-evolving AI markets. He breaks down the three mathematical drivers of inference compute demand and the non-negotiable requirement for sub-second latency.0:03–3:36 · Guest teaching 2/10 Announcing Cerebras' $1.1 Billion Fundraise The host opens the episode celebrating Cerebras' $1.1B round and establishes his connections across the AI ecosystem and Cognition. Feldman provides background context on early meetings with OpenAI founders and the progression of their hardware generations.3:37–8:41 · Guest teaching 5/10 Selecting Early vs. Late-Stage Investors Feldman walks through memory architectures, using the cup and straw analogy to explain why on-chip SRAM memory bandwidth overcomes the DRAM and HBM bottlenecks seen in traditional GPUs. The host listens and engages lightly with background context on benchmarking.8:41–12:06 · Guest teaching 4/10 Wafer-Scale Strategy and Continuous Benchmarking The host demonstrates domain familiarity by pointing out interconnect complexity and offering the metaphor of MLPerf as the Olympics versus Artificial Analysis as daily live traffic. Feldman details the hardware mess of wiring thousands of small SRAM chips versus wafer-scale silicon.12:11–16:44 · Guest teaching 3/10 Balancing Training Workloads and Inference Demand The host showcases deep technical insight by quoting internal Google TPU training-to-inference ratios (2:3) to estimate Cerebras' 1:5 ratio, which Feldman confirms. They discuss why latency directly causes customer churn in LLM applications.16:44–18:52 · Guest teaching 5/10 Foundational Chip Architecture Decisions and Linear Algebra The host brings up recent work in hybrid architectures and attention variants. Feldman details their foundational 2016 design choice to optimize sparse linear algebra rather than hard-coding 3x3 convolutions, which future-proofed the chip for transformers and diffusion.18:53–21:30 · Guest teaching 3/10 Expanding Cerebras Cloud and Developer Infrastructure The host directly probes whether Cerebras Coder puts the company into direct competition with its own developer customers. Feldman clarifies their strategy, explicitly stating Cerebras has no ambition to build an IDE or compete with developer toolmakers.21:31–25:30 · Guest teaching 5/10 Critical Engineering Challenges: Power, Routing, and I/O Bottlenecks Feldman discusses overlooked datacenter realities including huge power draw and upstream token routing. The host pushes back to clarify routing responsibilities and correctly diagnoses a customer benchmark failure as I/O-bound rather than compute-bound.25:32–28:40 · Guest teaching 4/10 Future Vision: The Exponential Rise of Real-Time AI Inference When the host asks for a 10-year grandmaster plan, Feldman dismisses long-range roadmaps as impractical in fast-evolving AI markets. He breaks down the three mathematical drivers of inference compute demand and the non-negotiable requirement for sub-second latency.0:03–3:36 · Guest disagreement 1/10 Announcing Cerebras' $1.1 Billion Fundraise The host opens the episode celebrating Cerebras' $1.1B round and establishes his connections across the AI ecosystem and Cognition. Feldman provides background context on early meetings with OpenAI founders and the progression of their hardware generations.3:37–8:41 · Guest disagreement 1/10 Selecting Early vs. Late-Stage Investors Feldman walks through memory architectures, using the cup and straw analogy to explain why on-chip SRAM memory bandwidth overcomes the DRAM and HBM bottlenecks seen in traditional GPUs. The host listens and engages lightly with background context on benchmarking.8:41–12:06 · Guest disagreement 1/10 Wafer-Scale Strategy and Continuous Benchmarking The host demonstrates domain familiarity by pointing out interconnect complexity and offering the metaphor of MLPerf as the Olympics versus Artificial Analysis as daily live traffic. Feldman details the hardware mess of wiring thousands of small SRAM chips versus wafer-scale silicon.12:11–16:44 · Guest disagreement 1/10 Balancing Training Workloads and Inference Demand The host showcases deep technical insight by quoting internal Google TPU training-to-inference ratios (2:3) to estimate Cerebras' 1:5 ratio, which Feldman confirms. They discuss why latency directly causes customer churn in LLM applications.16:44–18:52 · Guest disagreement 1/10 Foundational Chip Architecture Decisions and Linear Algebra The host brings up recent work in hybrid architectures and attention variants. Feldman details their foundational 2016 design choice to optimize sparse linear algebra rather than hard-coding 3x3 convolutions, which future-proofed the chip for transformers and diffusion.18:53–21:30 · Guest disagreement 2/10 Expanding Cerebras Cloud and Developer Infrastructure The host directly probes whether Cerebras Coder puts the company into direct competition with its own developer customers. Feldman clarifies their strategy, explicitly stating Cerebras has no ambition to build an IDE or compete with developer toolmakers.21:31–25:30 · Guest disagreement 2/10 Critical Engineering Challenges: Power, Routing, and I/O Bottlenecks Feldman discusses overlooked datacenter realities including huge power draw and upstream token routing. The host pushes back to clarify routing responsibilities and correctly diagnoses a customer benchmark failure as I/O-bound rather than compute-bound.25:32–28:40 · Guest disagreement 2/10 Future Vision: The Exponential Rise of Real-Time AI Inference When the host asks for a 10-year grandmaster plan, Feldman dismisses long-range roadmaps as impractical in fast-evolving AI markets. He breaks down the three mathematical drivers of inference compute demand and the non-negotiable requirement for sub-second latency.0:03–3:36 · The hosts pushing back 0/10 Announcing Cerebras' $1.1 Billion Fundraise The host opens the episode celebrating Cerebras' $1.1B round and establishes his connections across the AI ecosystem and Cognition. Feldman provides background context on early meetings with OpenAI founders and the progression of their hardware generations.3:37–8:41 · The hosts pushing back 1/10 Selecting Early vs. Late-Stage Investors Feldman walks through memory architectures, using the cup and straw analogy to explain why on-chip SRAM memory bandwidth overcomes the DRAM and HBM bottlenecks seen in traditional GPUs. The host listens and engages lightly with background context on benchmarking.8:41–12:06 · The hosts pushing back 1/10 Wafer-Scale Strategy and Continuous Benchmarking The host demonstrates domain familiarity by pointing out interconnect complexity and offering the metaphor of MLPerf as the Olympics versus Artificial Analysis as daily live traffic. Feldman details the hardware mess of wiring thousands of small SRAM chips versus wafer-scale silicon.12:11–16:44 · The hosts pushing back 1/10 Balancing Training Workloads and Inference Demand The host showcases deep technical insight by quoting internal Google TPU training-to-inference ratios (2:3) to estimate Cerebras' 1:5 ratio, which Feldman confirms. They discuss why latency directly causes customer churn in LLM applications.16:44–18:52 · The hosts pushing back 1/10 Foundational Chip Architecture Decisions and Linear Algebra The host brings up recent work in hybrid architectures and attention variants. Feldman details their foundational 2016 design choice to optimize sparse linear algebra rather than hard-coding 3x3 convolutions, which future-proofed the chip for transformers and diffusion.18:53–21:30 · The hosts pushing back 4/10 Expanding Cerebras Cloud and Developer Infrastructure The host directly probes whether Cerebras Coder puts the company into direct competition with its own developer customers. Feldman clarifies their strategy, explicitly stating Cerebras has no ambition to build an IDE or compete with developer toolmakers.21:31–25:30 · The hosts pushing back 4/10 Critical Engineering Challenges: Power, Routing, and I/O Bottlenecks Feldman discusses overlooked datacenter realities including huge power draw and upstream token routing. The host pushes back to clarify routing responsibilities and correctly diagnoses a customer benchmark failure as I/O-bound rather than compute-bound.25:32–28:40 · The hosts pushing back 2/10 Future Vision: The Exponential Rise of Real-Time AI Inference When the host asks for a 10-year grandmaster plan, Feldman dismisses long-range roadmaps as impractical in fast-evolving AI markets. He breaks down the three mathematical drivers of inference compute demand and the non-negotiable requirement for sub-second latency.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 26:05 Feldman rejects long-term 10-year roadmap framing

Feldman pushes back on the host's request for a 10-year grandmaster plan, arguing that extended roadmaps in fast-moving hardware markets are fundamentally flawed and likely to be wrong.

Hardest push from the hosts ▶ 18:52 Host presses on channel conflict with Cerebras Coder

The host confronts Feldman on whether launching Cerebras Coder creates conflict by directly competing with existing software customers like Cognition.

Biggest teaching moment ▶ 5:55 Feldman illustrates memory bandwidth via Coke and straw analogy

Feldman educates the host on why memory bandwidth is the primary operational constraint for LLM inference using the visual analogy of cup size versus straw diameter.

The host holds their own ▶ 13:25 Host cites internal Google TPU workload ratios

The host displays insider technical depth by citing internal Google TPU training-to-inference ratios (2:3) to accurately pinpoint Cerebras' operational workload split.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Announcing Cerebras' $1.1 Billion Fundraise 3210 The host opens the episode celebrating Cerebras' $1.1B round and establishes his connections across the AI ecosystem and Cognition. Feldman provides background context on early meetings with OpenAI founders and the progression of their hardware generations.
Selecting Early vs. Late-Stage Investors 4511 Feldman walks through memory architectures, using the cup and straw analogy to explain why on-chip SRAM memory bandwidth overcomes the DRAM and HBM bottlenecks seen in traditional GPUs. The host listens and engages lightly with background context on benchmarking.
Wafer-Scale Strategy and Continuous Benchmarking 6411 The host demonstrates domain familiarity by pointing out interconnect complexity and offering the metaphor of MLPerf as the Olympics versus Artificial Analysis as daily live traffic. Feldman details the hardware mess of wiring thousands of small SRAM chips versus wafer-scale silicon.
Balancing Training Workloads and Inference Demand 7311 The host showcases deep technical insight by quoting internal Google TPU training-to-inference ratios (2:3) to estimate Cerebras' 1:5 ratio, which Feldman confirms. They discuss why latency directly causes customer churn in LLM applications.
Foundational Chip Architecture Decisions and Linear Algebra 6511 The host brings up recent work in hybrid architectures and attention variants. Feldman details their foundational 2016 design choice to optimize sparse linear algebra rather than hard-coding 3x3 convolutions, which future-proofed the chip for transformers and diffusion.
Expanding Cerebras Cloud and Developer Infrastructure 5324 The host directly probes whether Cerebras Coder puts the company into direct competition with its own developer customers. Feldman clarifies their strategy, explicitly stating Cerebras has no ambition to build an IDE or compete with developer toolmakers.
Critical Engineering Challenges: Power, Routing, and I/O Bottlenecks 6524 Feldman discusses overlooked datacenter realities including huge power draw and upstream token routing. The host pushes back to clarify routing responsibilities and correctly diagnoses a customer benchmark failure as I/O-bound rather than compute-bound.
Future Vision: The Exponential Rise of Real-Time AI Inference 4422 When the host asks for a 10-year grandmaster plan, Feldman dismisses long-range roadmaps as impractical in fast-evolving AI markets. He breaks down the three mathematical drivers of inference compute demand and the non-negotiable requirement for sub-second latency.

Statements from this episode (17)

Assertion Supported
Feldman: Cerebras raised $1.1B at an $8.1B valuation
“So we announced a 1.1 billion dollar fundraise that we had completed. It was done at an 8.1 billion dollar post money valuation, and it was led by Fidelity and Atreides management.”
Andrew Feldman Oct 1, 2025 ▶ 0:35
Assertion Supported
Feldman: Sam Altman and Ilya Sutskever invested in Cerebras' early rounds
“In 2016, we met with Sam Altman and Ilya Suskovard at OpenAI and they were an idea and we were PowerPoint, right? That's amazing. And what AI was doing was identifying cats in pictures. And I think they ended up investing in us, both of them and many of their …”
Andrew Feldman Oct 1, 2025 ▶ 1:47
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Andrew Feldman Oct 1, 2025 ▶ 3:05
Insight
Feldman: Pre-IPO startups should target investors primarily focused on public markets
“I think in later stages, as you get close to IPO, you're looking for a very different type of investor. You're looking for an investor who Primarily does public markets.”
Andrew Feldman Oct 1, 2025 ▶ 4:23
Insight
Feldman: Memory bandwidth is the primary bottleneck in AI inference performance
“Inference performance comes from memory bandwidth and the memory bandwidth is the limiting factor. In inference performance. Remember, in order to generate a token, to generate a word, all the weights have to move from memory to compute. If you're constrained …”
Andrew Feldman Oct 1, 2025 ▶ 5:56
Assertion Supported
Feldman: Cerebras provides 2,625x more memory bandwidth than traditional GPUs
“And we have 2625 times more memory bandwidth than the GPU does.”
Andrew Feldman Oct 1, 2025 ▶ 6:22
Disclosure
Feldman: Cerebras serves Mistral AI's Le Chat assistant
“We serve Le Chat from Mistral.”
Andrew Feldman Oct 1, 2025 ▶ 8:17
Insight
Feldman: Multi-chip SRAM architectures make speculative decoding extremely difficult
“It limits the things you can do. It makes all sorts of cool AI techniques like speculative decode extremely difficult. Whereas if you have a giant chip, you might only need a handful.”
Andrew Feldman Oct 1, 2025 ▶ 10:49
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Andrew Feldman Oct 1, 2025 ▶ 11:03
Disclosure
Feldman: Cerebras trains models for Mayo Clinic, GSK, and US military
“We do a great deal of work with large enterprises in training, with Mayo Clinic, with GlaxoSmithKline, with the US military, with the Department of Energy, with our customers in the Middle East. We've trained leading models, language models in, in Arabic, in C…”
Andrew Feldman Oct 1, 2025 ▶ 12:40
Assertion Not checkable as stated
Feldman: AI startups are replacing closed-source models with fine-tuned open-source
“I think for sort of AI companies like Cognition, like all your competitors, like AlphaSense, like dozens of others, they are trying to replace closed source models with very, very fast open source models, and they're trying to drive the open source Accuracy dr…”
Andrew Feldman Oct 1, 2025 ▶ 14:20
Assertion Not checkable as stated
Enterprises retrain 10B-30B open models from scratch for legal data compliance
“We see at the large enterprise level, particularly those who have large data assets, a desire to train their own models and to go a little smaller, say in the 10 to thirty billion parameter category. I think especially the very large companies, enterprises hav…”
Andrew Feldman Oct 1, 2025 ▶ 15:27
Insight
Feldman: Running dozens of AI agents expands security attack surfaces geometrically
“When you spin off dozens of agents, the sort of attack surface of the of the solution expands geometrically.”
Andrew Feldman Oct 1, 2025 ▶ 16:06
Insight
Feldman: Chip architecture begins by deciding what not to be good at
“One of the hardest things in computer architecture, and one of the first things you do is you decide what you're not going to be good at. I'm going to build chip. What am I not going to be good at? We're not going to be good at general purpose compute. We're n…”
Andrew Feldman Oct 1, 2025 ▶ 17:24
Assertion Not checkable as stated
Feldman: Non-compute bottlenecks masked Cerebras' 20-25x speedup for a hyperscaler
“We were doing work with, we're doing work with one of the hyperscalers and for a product they have, and they said, look, we used your system and it didn't make us that much faster. And so we said, well, that's a surprise because we're 20, 25 times faster on th…”
Andrew Feldman Oct 1, 2025 ▶ 24:36
Insight
Feldman: AI taking 10-15 minutes for answers are proofs-of-concept, not products
“Number two is that for AI to deliver on its promise, to be embedded in our lives, it must be fast. There aren't things that are embedded in your life that make you wait 10 or 15 minutes to get a good answer. Those are proof of concepts, but those aren't produc…”
Andrew Feldman Oct 1, 2025 ▶ 27:36
Prediction Not checkable as stated
Feldman: Transformer architecture has several more years of viability
“And finally, I think the transformer, the current architecture has a way still to run. I think we will see that for several years more.”
Andrew Feldman Oct 1, 2025 ▶ 28:32
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.