Feb 20, 2020 · 20m · mad

Designing an AI Supercomputer // Michael James, Cerebras (FirstMark's Data Driven NYC)

Michael James · 16m spoken Matt Turck · 33s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In a presentation at FirstMark's Data Driven NYC, Cerebras co-founder Michael James explains how Cerebras designed and manufactured the Wafer Scale Engine—the world's largest chip—to solve fundamental memory and power bottlenecks in traditional AI computing. He details the hardware-software co-design, key manufacturing breakthroughs, and real-world scientific deployments powering the next generation of artificial intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 2.9% of the talking time here. How this is scored →

Matt as informed peer 0.0 Guest teaching 2.7 Guest disagreement 1.1 Matt pushing back 0.0
05100:0010:0020:000:51–3:01 · Matt as informed peer 0/10 The AI Wave & Early Industry Skepticism Michael James opens the presentation by describing the surge in AI interest and recounting how early advisors warned that creating dedicated computer hardware for AI was crazy. Because this segment is entirely a guest monologue, host expertise and host pushback are scored zero.3:01–5:45 · Matt as informed peer 0/10 Questioning Conventional Von Neumann Architecture James presents a technical breakdown of von Neumann architecture limitations, using a bird brain calculation to show that classical memory access requires excessive power. As a monologue segment, host metrics remain zero while the guest provides architectural context.5:45–8:09 · Matt as informed peer 0/10 A New Architectural Paradigm for Data Science James reveals the Cerebras Wafer-Scale Engine, walking the audience through its 462 square centimeter size, 1.2 trillion transistors, and 400,000 cores. The host does not speak in this presentation segment, keeping host-side scores at zero.8:09–10:36 · Matt as informed peer 0/10 Scale Comparison: Cerebras WSE vs. Standard GPU The guest explains how Cerebras solved silicon yield problems by routing around micro-defects across 400,000 independent cores instead of discarding whole wafers. Host participation is absent, maintaining zero host scores.10:36–13:10 · Matt as informed peer 0/10 Overcoming Engineering Challenges & Silicon Thermal Expansion James details a year-long physical engineering challenge involving thermal expansion and silicon cracking, explaining how electron microscopy uncovered unexpected laser vaporization residue. The host is inactive during this detailed monologue.13:10–15:16 · Matt as informed peer 0/10 Summary of Cerebras WSE Performance Advantages James demonstrates how software frameworks map neural network graphs like ResNet-50 onto WSE cores and highlights early adoption by US National Labs. Host metrics are zero as the presentation continues uninterrupted.15:16–15:48 · Matt as informed peer 0/10 Enabling New Categories of Learning Algorithms James concludes his talk by encouraging data scientists to break free from GPU-based constraints and invent new categories of learning algorithms. The host does not engage during this closing monologue span.0:51–3:01 · Guest teaching 1/10 The AI Wave & Early Industry Skepticism Michael James opens the presentation by describing the surge in AI interest and recounting how early advisors warned that creating dedicated computer hardware for AI was crazy. Because this segment is entirely a guest monologue, host expertise and host pushback are scored zero.3:01–5:45 · Guest teaching 3/10 Questioning Conventional Von Neumann Architecture James presents a technical breakdown of von Neumann architecture limitations, using a bird brain calculation to show that classical memory access requires excessive power. As a monologue segment, host metrics remain zero while the guest provides architectural context.5:45–8:09 · Guest teaching 3/10 A New Architectural Paradigm for Data Science James reveals the Cerebras Wafer-Scale Engine, walking the audience through its 462 square centimeter size, 1.2 trillion transistors, and 400,000 cores. The host does not speak in this presentation segment, keeping host-side scores at zero.8:09–10:36 · Guest teaching 3/10 Scale Comparison: Cerebras WSE vs. Standard GPU The guest explains how Cerebras solved silicon yield problems by routing around micro-defects across 400,000 independent cores instead of discarding whole wafers. Host participation is absent, maintaining zero host scores.10:36–13:10 · Guest teaching 4/10 Overcoming Engineering Challenges & Silicon Thermal Expansion James details a year-long physical engineering challenge involving thermal expansion and silicon cracking, explaining how electron microscopy uncovered unexpected laser vaporization residue. The host is inactive during this detailed monologue.13:10–15:16 · Guest teaching 3/10 Summary of Cerebras WSE Performance Advantages James demonstrates how software frameworks map neural network graphs like ResNet-50 onto WSE cores and highlights early adoption by US National Labs. Host metrics are zero as the presentation continues uninterrupted.15:16–15:48 · Guest teaching 2/10 Enabling New Categories of Learning Algorithms James concludes his talk by encouraging data scientists to break free from GPU-based constraints and invent new categories of learning algorithms. The host does not engage during this closing monologue span.0:51–3:01 · Guest disagreement 1/10 The AI Wave & Early Industry Skepticism Michael James opens the presentation by describing the surge in AI interest and recounting how early advisors warned that creating dedicated computer hardware for AI was crazy. Because this segment is entirely a guest monologue, host expertise and host pushback are scored zero.3:01–5:45 · Guest disagreement 2/10 Questioning Conventional Von Neumann Architecture James presents a technical breakdown of von Neumann architecture limitations, using a bird brain calculation to show that classical memory access requires excessive power. As a monologue segment, host metrics remain zero while the guest provides architectural context.5:45–8:09 · Guest disagreement 1/10 A New Architectural Paradigm for Data Science James reveals the Cerebras Wafer-Scale Engine, walking the audience through its 462 square centimeter size, 1.2 trillion transistors, and 400,000 cores. The host does not speak in this presentation segment, keeping host-side scores at zero.8:09–10:36 · Guest disagreement 1/10 Scale Comparison: Cerebras WSE vs. Standard GPU The guest explains how Cerebras solved silicon yield problems by routing around micro-defects across 400,000 independent cores instead of discarding whole wafers. Host participation is absent, maintaining zero host scores.10:36–13:10 · Guest disagreement 1/10 Overcoming Engineering Challenges & Silicon Thermal Expansion James details a year-long physical engineering challenge involving thermal expansion and silicon cracking, explaining how electron microscopy uncovered unexpected laser vaporization residue. The host is inactive during this detailed monologue.13:10–15:16 · Guest disagreement 1/10 Summary of Cerebras WSE Performance Advantages James demonstrates how software frameworks map neural network graphs like ResNet-50 onto WSE cores and highlights early adoption by US National Labs. Host metrics are zero as the presentation continues uninterrupted.15:16–15:48 · Guest disagreement 1/10 Enabling New Categories of Learning Algorithms James concludes his talk by encouraging data scientists to break free from GPU-based constraints and invent new categories of learning algorithms. The host does not engage during this closing monologue span.0:51–3:01 · Matt pushing back 0/10 The AI Wave & Early Industry Skepticism Michael James opens the presentation by describing the surge in AI interest and recounting how early advisors warned that creating dedicated computer hardware for AI was crazy. Because this segment is entirely a guest monologue, host expertise and host pushback are scored zero.3:01–5:45 · Matt pushing back 0/10 Questioning Conventional Von Neumann Architecture James presents a technical breakdown of von Neumann architecture limitations, using a bird brain calculation to show that classical memory access requires excessive power. As a monologue segment, host metrics remain zero while the guest provides architectural context.5:45–8:09 · Matt pushing back 0/10 A New Architectural Paradigm for Data Science James reveals the Cerebras Wafer-Scale Engine, walking the audience through its 462 square centimeter size, 1.2 trillion transistors, and 400,000 cores. The host does not speak in this presentation segment, keeping host-side scores at zero.8:09–10:36 · Matt pushing back 0/10 Scale Comparison: Cerebras WSE vs. Standard GPU The guest explains how Cerebras solved silicon yield problems by routing around micro-defects across 400,000 independent cores instead of discarding whole wafers. Host participation is absent, maintaining zero host scores.10:36–13:10 · Matt pushing back 0/10 Overcoming Engineering Challenges & Silicon Thermal Expansion James details a year-long physical engineering challenge involving thermal expansion and silicon cracking, explaining how electron microscopy uncovered unexpected laser vaporization residue. The host is inactive during this detailed monologue.13:10–15:16 · Matt pushing back 0/10 Summary of Cerebras WSE Performance Advantages James demonstrates how software frameworks map neural network graphs like ResNet-50 onto WSE cores and highlights early adoption by US National Labs. Host metrics are zero as the presentation continues uninterrupted.15:16–15:48 · Matt pushing back 0/10 Enabling New Categories of Learning Algorithms James concludes his talk by encouraging data scientists to break free from GPU-based constraints and invent new categories of learning algorithms. The host does not engage during this closing monologue span.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 18.3% · guest 81.7%15:00 · Matt 18.3% · guest 81.7%18:00 · Matt 2.8% · guest 97.2%18:00 · Matt 2.8% · guest 97.2%
Sharpest disagreement ▶ 2:15 Dismissing conventional advice to defer hardware development

Michael James explicitly rejects the standard industry advice to wait for researchers to settle on algorithms before building hardware, calling it 'the wrong advice' and insisting hardware co-design must happen during rapid change.

Hardest push from Matt ▶ 16:42 Host challenges guest on rival AI startup differentiation

Matt Turck directly challenges James by citing specific competitors like Graphcore and SambaNova, asking whether Cerebras is actually doing something distinct or merely pursuing the exact same market opportunity.

Biggest teaching moment ▶ 4:00 Educating audience on von Neumann memory power limits

Michael James walks through a detailed mathematical comparison showing that a traditional von Neumann architecture would require over half a megawatt just for memory transfers to emulate a bird's brain, demonstrating why conventional chips fail for AI workloads.

Matt holds his own ▶ 16:42 Host demonstrates market expertise by naming chip competitors

Matt Turck demonstrates strong industry context by explicitly naming rival AI hardware startups Graphcore and SambaNova to probe Cerebras' precise positioning in the semiconductor landscape.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The AI Wave & Early Industry Skepticism 0110 Michael James opens the presentation by describing the surge in AI interest and recounting how early advisors warned that creating dedicated computer hardware for AI was crazy. Because this segment is entirely a guest monologue, host expertise and host pushback are scored zero.
Questioning Conventional Von Neumann Architecture 0320 James presents a technical breakdown of von Neumann architecture limitations, using a bird brain calculation to show that classical memory access requires excessive power. As a monologue segment, host metrics remain zero while the guest provides architectural context.
A New Architectural Paradigm for Data Science 0310 James reveals the Cerebras Wafer-Scale Engine, walking the audience through its 462 square centimeter size, 1.2 trillion transistors, and 400,000 cores. The host does not speak in this presentation segment, keeping host-side scores at zero.
Scale Comparison: Cerebras WSE vs. Standard GPU 0310 The guest explains how Cerebras solved silicon yield problems by routing around micro-defects across 400,000 independent cores instead of discarding whole wafers. Host participation is absent, maintaining zero host scores.
Overcoming Engineering Challenges & Silicon Thermal Expansion 0410 James details a year-long physical engineering challenge involving thermal expansion and silicon cracking, explaining how electron microscopy uncovered unexpected laser vaporization residue. The host is inactive during this detailed monologue.
Summary of Cerebras WSE Performance Advantages 0310 James demonstrates how software frameworks map neural network graphs like ResNet-50 onto WSE cores and highlights early adoption by US National Labs. Host metrics are zero as the presentation continues uninterrupted.
Enabling New Categories of Learning Algorithms 0210 James concludes his talk by encouraging data scientists to break free from GPU-based constraints and invent new categories of learning algorithms. The host does not engage during this closing monologue span.

Statements from this episode (21)

Assertion Contradicted
Cerebras built the first supercomputer designed end-to-end for AI
“And it's also the first supercomputer that has been designed from beginning to end for artificial intelligence and deep learning.”
Michael James Feb 20, 2020 ▶ 0:33
Insight
New computing workloads always demand purpose-built hardware
“When we look throughout the history of computing, it has never happened that a new important workload has come along, and the correct answer to that workload was to use some pre-existing equipment that had been designed to do something else”
Michael James Feb 20, 2020 ▶ 4:03
Opinion
Von Neumann architecture might not be the right path for AI
“So, so maybe there's a sign that, that we're not on the right track with the architecture we have in, in von Neumann computers.”
Michael James Feb 20, 2020 ▶ 5:38
Insight
James: Deep learning models rely on data flow, not instruction-driven loops
“The models we're producing don't work based on instructions. It, it's not a central operator that says do this, do that. Instead it's the data that flows over our model and a simple feedback equation that tells that data how to manipulate and transform the mod…”
Michael James Feb 20, 2020 ▶ 6:25
Disclosure
Cerebras designed a chip as large as possible with interleaved memory
“We decided to build a machine that was as large as we could, that had memory interleaved directly with the computational substrate.”
Michael James Feb 20, 2020 ▶ 6:44
Assertion Supported
The Cerebras Wafer Scale Engine is the world's largest processor
“This is the world's largest processor.”
Michael James Feb 20, 2020 ▶ 7:13
Assertion Supported
Cerebras's WSE is the first chip with over 1 trillion transistors
“It is the first part that has over a trillion transistors, and we're here, and at 1.2, and it has 400,000 sparse linear algebra cores on it.”
Michael James Feb 20, 2020 ▶ 7:26
Assertion Supported
Cerebras's WSE delivers 9 petabytes per second of memory bandwidth
“There's 18 gigabytes of memory right, right on that, ah, on that silicon wafer, and it gives an unprecedented nine gigabytes per second, sorry, nine petabytes per second of memory bandwidth.”
Michael James Feb 20, 2020 ▶ 7:42
Assertion Not checkable as stated
Wafer-scale chips were previously presumed impossible to build
“And when we showed this the world got pretty excited, so there was a large press splash, and in part that's because this had been presumed to be impossible to achieve.”
Michael James Feb 20, 2020 ▶ 8:19
Assertion Supported
Cerebras routes around manufacturing defects across its 400,000 cores
“By having 400,000 independent cores, it's perfectly fine if we have a few hundred defects on there, and it's only our job to route the information flows around those defects.”
Michael James Feb 20, 2020 ▶ 8:53
Assertion Supported
No motherboard had ever housed a processor as large as Cerebras's
“No one before had ever put a processor that large in a motherboard, and so we had to also make a system to house this thing.”
Michael James Feb 20, 2020 ▶ 9:19
Assertion Supported
Cerebras CS-1 system enclosure contains a single wafer plus cooling hardware
“There's only one wafer mounted in there, and the entire rest of this is about, ah, cooling and power delivery.”
Michael James Feb 20, 2020 ▶ 9:39
What-if
Traditional power delivery would carry 40,000 amps and melt Cerebras's processor
“We would have 40,000 amps, and that would actually melt the part if we chose to do it that way, so we had to bring some of the geometric thinking beyond how we think geometry in the models into this three D structure.”
Michael James Feb 20, 2020 ▶ 10:02
Assertion Not checkable as stated
Michael James: Cerebras succeeded at reticle stitching on the first attempt
“Reticle stitching to get above the optical limit of how you manufacture these things that worked the first time we tried it.”
Michael James Feb 20, 2020 ▶ 10:50
Assertion Supported
Cerebras's chip provides 10,000x the memory bandwidth of standard GPUs
“So, having solved this, we were left with a machine that was pretty impressive on most of the metrics, being larger, more cores, 10,000 times the memory bandwidth and 33,000 times the fabric bandwidth.”
Michael James Feb 20, 2020 ▶ 13:10
Opinion
Programming Cerebras's engine is simpler than programming conventional computers
“So again, the simpler, ah, I think simpler than programming a conventional computer where you have to specify millions of instructions.”
Michael James Feb 20, 2020 ▶ 13:29
Assertion Supported
Cerebras maps neural network layers across 400,000 chip cores
“We have, ah, an optimization solver that places these over the 400,000 cores.”
Michael James Feb 20, 2020 ▶ 14:21
Assertion Supported
Cerebras's first customer was the U.S. National Labs
“It's pretty unusual in Silicon Valley to have your first customer be the U.S. National Labs”
Michael James Feb 20, 2020 ▶ 14:46
Assertion Supported
Lawrence Livermore integrated Cerebras into its Lassen supercomputer
“Lawrence Livermore has integrated into their Lassen supercomputer.”
Michael James Feb 20, 2020 ▶ 15:08
Assertion Supported
Cerebras' Wafer Scale Engine is fabricated at TSMC
“It is fabricated at Taiwan Semiconductor, which is one of the world's large major fab houses.”
Michael James Feb 20, 2020 ▶ 18:23
Prediction Open · timeframe Feb 2035
Quantum and optical computing remain 10 to 15 years away from production
“We think that because that's primary research, that, that's a good, ah, 10 to 15 years out, whereas what we looked to do was to take, ah, pre, ah, technology we could manufacture, maybe difficult, but all achievable, so we weren't trying to invent any new law …”
Michael James Feb 20, 2020 ▶ 19:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.