Feb 26, 2026 · 1h 13m · cheeky-pint

Reiner Pope of MatX on accelerating AI with transformer-optimized chips

Reiner Pope · 46m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Former Google TPU architect and MatX co-founder Reiner Pope breaks down the history, architectural trade-offs, and future of specialized AI silicon, highlighting how tailored hardware and software co-design are essential for scaling modern large language models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. John holds 29.2% of the talking time here. How this is scored →

John as informed peer 4.7 Guest teaching 5.8 Guest disagreement 1.4 John pushing back 1.8
05100:0015:0030:0045:001:00:001:35–5:30 · John as informed peer 5/10 Origins of TPU and Hardware Parallelization Collison explores the origins of TPU and asks whether the timing of TPUs and transformers was coincidental or rooted in parallelization insights. Pope educates Collison on the mechanical sympathy required for parallel hardware versus CPU instruction parallelism.5:31–7:55 · John as informed peer 4/10 Comparing CPUs and GPUs for AI Workloads Collison asks for an intuitive mental model comparing CPUs and GPUs beyond the tautology that GPUs do math. Pope provides a clear analogy comparing a truck carrying a large payload to a motorcycle navigating an obstacle course.7:55–11:42 · John as informed peer 5/10 Founding MatX and Core Value Proposition Pope outlines MatX's founding thesis around matrix sizes and lower precision arithmetic. Collison pushes on whether product risk exists given that LLMs are clearly established, prompting Pope to clarify that the bet was made prior to ChatGPT.11:43–21:47 · John as informed peer 5/10 Scaling MatX, Funding, and AI Supply Chain Bottlenecks Collison asks how MatX will compete with giants for supply chain allocations across wafers, HBM, and racks. Pope breaks down the exact trade-offs between SRAM and HBM memory latencies and explains how ironclad buyer contracts secure supplier capacity.21:47–25:29 · John as informed peer 4/10 MatX Architecture: Memory, Systolic Arrays, and Low Precision Pope explains MatX's architecture combining SRAM, HBM, splittable systolic arrays, and 4-bit precision. Collison asks basic clarifying questions on bit precision and systolic arrays while Pope explains the mathematical trade-offs.25:31–34:28 · John as informed peer 4/10 The Chip Design Process: From Verilog to Tape Out Collison explores the physical chip design cycle, comparing software iteration to chip tape-outs. When Collison asks why simulation doesn't eliminate all errors before spending thirty million dollars, Pope retorts that software companies also ship bugs to production.34:29–44:14 · John as informed peer 6/10 Software Moats, TSMC Dominance, and Space Data Centers Collison pushes Pope on why TSMC has maintained an uncontested monopoly at leading edge nodes despite geopolitical risks, and presses on the viability of space data centers. Pope counters that leading-edge nodes only matter for specific power-sensitive workloads and details chip redundancy models in space.44:18–56:04 · John as informed peer 4/10 Sponsor Segment: Stripe Billing for AI Products Segment starts with Collison reading a Stripe Billing ad read before discussing AI development bottlenecks and state management in context windows. Pope explains how context size struggles to grow while parameter counts expand.56:06–1:02:57 · John as informed peer 5/10 Team Culture, Numerics Research, and In-Head Iteration Pope outlines MatX's unique ML research team co-designing numerics and doing rapid in-head estimation before simulation. Collison connects Pope's gate-cost heuristic to Jeff Dean's famous 'numbers every engineer should know'.1:02:58–1:13:02 · John as informed peer 5/10 Why Rust, Systems Optimization, Cuckoo Hashing, and Architecture Frontiers Pope explains his preference for Rust over Go and Haskell due to low-level memory control and discusses optimizing SIMD vector instructions for cuckoo hashing. Collison probes into practical workloads and wraps up on architectural opportunities.1:35–5:30 · Guest teaching 5/10 Origins of TPU and Hardware Parallelization Collison explores the origins of TPU and asks whether the timing of TPUs and transformers was coincidental or rooted in parallelization insights. Pope educates Collison on the mechanical sympathy required for parallel hardware versus CPU instruction parallelism.5:31–7:55 · Guest teaching 6/10 Comparing CPUs and GPUs for AI Workloads Collison asks for an intuitive mental model comparing CPUs and GPUs beyond the tautology that GPUs do math. Pope provides a clear analogy comparing a truck carrying a large payload to a motorcycle navigating an obstacle course.7:55–11:42 · Guest teaching 5/10 Founding MatX and Core Value Proposition Pope outlines MatX's founding thesis around matrix sizes and lower precision arithmetic. Collison pushes on whether product risk exists given that LLMs are clearly established, prompting Pope to clarify that the bet was made prior to ChatGPT.11:43–21:47 · Guest teaching 6/10 Scaling MatX, Funding, and AI Supply Chain Bottlenecks Collison asks how MatX will compete with giants for supply chain allocations across wafers, HBM, and racks. Pope breaks down the exact trade-offs between SRAM and HBM memory latencies and explains how ironclad buyer contracts secure supplier capacity.21:47–25:29 · Guest teaching 7/10 MatX Architecture: Memory, Systolic Arrays, and Low Precision Pope explains MatX's architecture combining SRAM, HBM, splittable systolic arrays, and 4-bit precision. Collison asks basic clarifying questions on bit precision and systolic arrays while Pope explains the mathematical trade-offs.25:31–34:28 · Guest teaching 7/10 The Chip Design Process: From Verilog to Tape Out Collison explores the physical chip design cycle, comparing software iteration to chip tape-outs. When Collison asks why simulation doesn't eliminate all errors before spending thirty million dollars, Pope retorts that software companies also ship bugs to production.34:29–44:14 · Guest teaching 5/10 Software Moats, TSMC Dominance, and Space Data Centers Collison pushes Pope on why TSMC has maintained an uncontested monopoly at leading edge nodes despite geopolitical risks, and presses on the viability of space data centers. Pope counters that leading-edge nodes only matter for specific power-sensitive workloads and details chip redundancy models in space.44:18–56:04 · Guest teaching 5/10 Sponsor Segment: Stripe Billing for AI Products Segment starts with Collison reading a Stripe Billing ad read before discussing AI development bottlenecks and state management in context windows. Pope explains how context size struggles to grow while parameter counts expand.56:06–1:02:57 · Guest teaching 6/10 Team Culture, Numerics Research, and In-Head Iteration Pope outlines MatX's unique ML research team co-designing numerics and doing rapid in-head estimation before simulation. Collison connects Pope's gate-cost heuristic to Jeff Dean's famous 'numbers every engineer should know'.1:02:58–1:13:02 · Guest teaching 6/10 Why Rust, Systems Optimization, Cuckoo Hashing, and Architecture Frontiers Pope explains his preference for Rust over Go and Haskell due to low-level memory control and discusses optimizing SIMD vector instructions for cuckoo hashing. Collison probes into practical workloads and wraps up on architectural opportunities.1:35–5:30 · Guest disagreement 2/10 Origins of TPU and Hardware Parallelization Collison explores the origins of TPU and asks whether the timing of TPUs and transformers was coincidental or rooted in parallelization insights. Pope educates Collison on the mechanical sympathy required for parallel hardware versus CPU instruction parallelism.5:31–7:55 · Guest disagreement 1/10 Comparing CPUs and GPUs for AI Workloads Collison asks for an intuitive mental model comparing CPUs and GPUs beyond the tautology that GPUs do math. Pope provides a clear analogy comparing a truck carrying a large payload to a motorcycle navigating an obstacle course.7:55–11:42 · Guest disagreement 2/10 Founding MatX and Core Value Proposition Pope outlines MatX's founding thesis around matrix sizes and lower precision arithmetic. Collison pushes on whether product risk exists given that LLMs are clearly established, prompting Pope to clarify that the bet was made prior to ChatGPT.11:43–21:47 · Guest disagreement 1/10 Scaling MatX, Funding, and AI Supply Chain Bottlenecks Collison asks how MatX will compete with giants for supply chain allocations across wafers, HBM, and racks. Pope breaks down the exact trade-offs between SRAM and HBM memory latencies and explains how ironclad buyer contracts secure supplier capacity.21:47–25:29 · Guest disagreement 1/10 MatX Architecture: Memory, Systolic Arrays, and Low Precision Pope explains MatX's architecture combining SRAM, HBM, splittable systolic arrays, and 4-bit precision. Collison asks basic clarifying questions on bit precision and systolic arrays while Pope explains the mathematical trade-offs.25:31–34:28 · Guest disagreement 2/10 The Chip Design Process: From Verilog to Tape Out Collison explores the physical chip design cycle, comparing software iteration to chip tape-outs. When Collison asks why simulation doesn't eliminate all errors before spending thirty million dollars, Pope retorts that software companies also ship bugs to production.34:29–44:14 · Guest disagreement 2/10 Software Moats, TSMC Dominance, and Space Data Centers Collison pushes Pope on why TSMC has maintained an uncontested monopoly at leading edge nodes despite geopolitical risks, and presses on the viability of space data centers. Pope counters that leading-edge nodes only matter for specific power-sensitive workloads and details chip redundancy models in space.44:18–56:04 · Guest disagreement 1/10 Sponsor Segment: Stripe Billing for AI Products Segment starts with Collison reading a Stripe Billing ad read before discussing AI development bottlenecks and state management in context windows. Pope explains how context size struggles to grow while parameter counts expand.56:06–1:02:57 · Guest disagreement 1/10 Team Culture, Numerics Research, and In-Head Iteration Pope outlines MatX's unique ML research team co-designing numerics and doing rapid in-head estimation before simulation. Collison connects Pope's gate-cost heuristic to Jeff Dean's famous 'numbers every engineer should know'.1:02:58–1:13:02 · Guest disagreement 1/10 Why Rust, Systems Optimization, Cuckoo Hashing, and Architecture Frontiers Pope explains his preference for Rust over Go and Haskell due to low-level memory control and discusses optimizing SIMD vector instructions for cuckoo hashing. Collison probes into practical workloads and wraps up on architectural opportunities.1:35–5:30 · John pushing back 2/10 Origins of TPU and Hardware Parallelization Collison explores the origins of TPU and asks whether the timing of TPUs and transformers was coincidental or rooted in parallelization insights. Pope educates Collison on the mechanical sympathy required for parallel hardware versus CPU instruction parallelism.5:31–7:55 · John pushing back 1/10 Comparing CPUs and GPUs for AI Workloads Collison asks for an intuitive mental model comparing CPUs and GPUs beyond the tautology that GPUs do math. Pope provides a clear analogy comparing a truck carrying a large payload to a motorcycle navigating an obstacle course.7:55–11:42 · John pushing back 2/10 Founding MatX and Core Value Proposition Pope outlines MatX's founding thesis around matrix sizes and lower precision arithmetic. Collison pushes on whether product risk exists given that LLMs are clearly established, prompting Pope to clarify that the bet was made prior to ChatGPT.11:43–21:47 · John pushing back 2/10 Scaling MatX, Funding, and AI Supply Chain Bottlenecks Collison asks how MatX will compete with giants for supply chain allocations across wafers, HBM, and racks. Pope breaks down the exact trade-offs between SRAM and HBM memory latencies and explains how ironclad buyer contracts secure supplier capacity.21:47–25:29 · John pushing back 1/10 MatX Architecture: Memory, Systolic Arrays, and Low Precision Pope explains MatX's architecture combining SRAM, HBM, splittable systolic arrays, and 4-bit precision. Collison asks basic clarifying questions on bit precision and systolic arrays while Pope explains the mathematical trade-offs.25:31–34:28 · John pushing back 2/10 The Chip Design Process: From Verilog to Tape Out Collison explores the physical chip design cycle, comparing software iteration to chip tape-outs. When Collison asks why simulation doesn't eliminate all errors before spending thirty million dollars, Pope retorts that software companies also ship bugs to production.34:29–44:14 · John pushing back 4/10 Software Moats, TSMC Dominance, and Space Data Centers Collison pushes Pope on why TSMC has maintained an uncontested monopoly at leading edge nodes despite geopolitical risks, and presses on the viability of space data centers. Pope counters that leading-edge nodes only matter for specific power-sensitive workloads and details chip redundancy models in space.44:18–56:04 · John pushing back 2/10 Sponsor Segment: Stripe Billing for AI Products Segment starts with Collison reading a Stripe Billing ad read before discussing AI development bottlenecks and state management in context windows. Pope explains how context size struggles to grow while parameter counts expand.56:06–1:02:57 · John pushing back 1/10 Team Culture, Numerics Research, and In-Head Iteration Pope outlines MatX's unique ML research team co-designing numerics and doing rapid in-head estimation before simulation. Collison connects Pope's gate-cost heuristic to Jeff Dean's famous 'numbers every engineer should know'.1:02:58–1:13:02 · John pushing back 1/10 Why Rust, Systems Optimization, Cuckoo Hashing, and Architecture Frontiers Pope explains his preference for Rust over Go and Haskell due to low-level memory control and discusses optimizing SIMD vector instructions for cuckoo hashing. Collison probes into practical workloads and wraps up on architectural opportunities.

speaking balance: gold is John, purple is the guest (3 minute bins)

0:00 · John 34.2% · guest 65.8%0:00 · John 34.2% · guest 65.8%3:00 · John 26.3% · guest 73.7%3:00 · John 26.3% · guest 73.7%6:00 · John 29.6% · guest 70.4%6:00 · John 29.6% · guest 70.4%9:00 · John 18.8% · guest 81.2%9:00 · John 18.8% · guest 81.2%12:00 · John 49.3% · guest 50.7%12:00 · John 49.3% · guest 50.7%15:00 · John 44.2% · guest 55.8%15:00 · John 44.2% · guest 55.8%18:00 · John 27.4% · guest 72.6%18:00 · John 27.4% · guest 72.6%21:00 · John 19.7% · guest 80.3%21:00 · John 19.7% · guest 80.3%24:00 · John 23.5% · guest 76.5%24:00 · John 23.5% · guest 76.5%27:00 · John 31.5% · guest 68.5%27:00 · John 31.5% · guest 68.5%30:00 · John 22.9% · guest 77.1%30:00 · John 22.9% · guest 77.1%33:00 · John 36.4% · guest 63.6%33:00 · John 36.4% · guest 63.6%36:00 · John 39.4% · guest 60.6%36:00 · John 39.4% · guest 60.6%39:00 · John 28.6% · guest 71.4%39:00 · John 28.6% · guest 71.4%42:00 · John 45.8% · guest 54.2%42:00 · John 45.8% · guest 54.2%45:00 · John 29.1% · guest 70.9%45:00 · John 29.1% · guest 70.9%48:00 · John 17.2% · guest 82.8%48:00 · John 17.2% · guest 82.8%51:00 · John 44.5% · guest 55.5%51:00 · John 44.5% · guest 55.5%54:00 · John 35.3% · guest 64.7%54:00 · John 35.3% · guest 64.7%57:00 · John 24.2% · guest 75.8%57:00 · John 24.2% · guest 75.8%1:00:00 · John 23.7% · guest 76.3%1:00:00 · John 23.7% · guest 76.3%1:03:00 · John 14.3% · guest 85.7%1:03:00 · John 14.3% · guest 85.7%1:06:00 · John 7.9% · guest 92.1%1:06:00 · John 7.9% · guest 92.1%1:09:00 · John 36.3% · guest 63.7%1:09:00 · John 36.3% · guest 63.7%1:12:00 · John 3.1% · guest 96.9%1:12:00 · John 3.1% · guest 96.9%
Sharpest disagreement ▶ 32:43 Software companies ship bugs too

Pope dryly dismisses Collison's suggestion that thirty million dollar tape-outs shouldn't have logical bugs by comparing it directly to software companies shipping bugs to production.

Hardest push from John ▶ 38:03 Pushing on why TSMC faces no competitors

Collison repeatedly challenges the assumption that TSMC's position is inevitable, pointing out that ship-building and airplane manufacturing have competitive markets.

Biggest teaching moment ▶ 15:22 Explaining the SRAM vs HBM memory tradeoff

Pope gives a detailed technical breakdown of why Grok and Cerebras struggle with throughput despite low latency, explaining how combining SRAM and HBM solves both.

John holds their own ▶ 1:01:00 Connecting gate heuristics to Jeff Dean

Collison demonstrates deep familiarity with high-scale systems folklore by connecting Pope's internal 'go/gates' rules to Jeff Dean's classic systems performance numbers.

the scores for every segment, with the reasoning behind each
ChapterTopicJohn as informed peerGuest teachingGuest disagreementJohn pushing backWhy
Origins of TPU and Hardware Parallelization 5522 Collison explores the origins of TPU and asks whether the timing of TPUs and transformers was coincidental or rooted in parallelization insights. Pope educates Collison on the mechanical sympathy required for parallel hardware versus CPU instruction parallelism.
Comparing CPUs and GPUs for AI Workloads 4611 Collison asks for an intuitive mental model comparing CPUs and GPUs beyond the tautology that GPUs do math. Pope provides a clear analogy comparing a truck carrying a large payload to a motorcycle navigating an obstacle course.
Founding MatX and Core Value Proposition 5522 Pope outlines MatX's founding thesis around matrix sizes and lower precision arithmetic. Collison pushes on whether product risk exists given that LLMs are clearly established, prompting Pope to clarify that the bet was made prior to ChatGPT.
Scaling MatX, Funding, and AI Supply Chain Bottlenecks 5612 Collison asks how MatX will compete with giants for supply chain allocations across wafers, HBM, and racks. Pope breaks down the exact trade-offs between SRAM and HBM memory latencies and explains how ironclad buyer contracts secure supplier capacity.
MatX Architecture: Memory, Systolic Arrays, and Low Precision 4711 Pope explains MatX's architecture combining SRAM, HBM, splittable systolic arrays, and 4-bit precision. Collison asks basic clarifying questions on bit precision and systolic arrays while Pope explains the mathematical trade-offs.
The Chip Design Process: From Verilog to Tape Out 4722 Collison explores the physical chip design cycle, comparing software iteration to chip tape-outs. When Collison asks why simulation doesn't eliminate all errors before spending thirty million dollars, Pope retorts that software companies also ship bugs to production.
Software Moats, TSMC Dominance, and Space Data Centers 6524 Collison pushes Pope on why TSMC has maintained an uncontested monopoly at leading edge nodes despite geopolitical risks, and presses on the viability of space data centers. Pope counters that leading-edge nodes only matter for specific power-sensitive workloads and details chip redundancy models in space.
Sponsor Segment: Stripe Billing for AI Products 4512 Segment starts with Collison reading a Stripe Billing ad read before discussing AI development bottlenecks and state management in context windows. Pope explains how context size struggles to grow while parameter counts expand.
Team Culture, Numerics Research, and In-Head Iteration 5611 Pope outlines MatX's unique ML research team co-designing numerics and doing rapid in-head estimation before simulation. Collison connects Pope's gate-cost heuristic to Jeff Dean's famous 'numbers every engineer should know'.
Why Rust, Systems Optimization, Cuckoo Hashing, and Architecture Frontiers 5611 Pope explains his preference for Rust over Go and Haskell due to low-level memory control and discusses optimizing SIMD vector instructions for cuckoo hashing. Collison probes into practical workloads and wraps up on architectural opportunities.

Statements from this episode (26)

Assertion Not checkable as stated
Pope: Almost every senior AI researcher passed through Google Brain
“Pretty much anyone who's maybe, I don't know, over 30 and at a large lab has been at Google Brain at some point.”
Reiner Pope Feb 26, 2026 ▶ 0:56
Opinion
Pope: Google's decision to build TPUs specifically for neural nets paid off
“They at least had the option, like the opportunity to design the TPUs for neural nets at least rather than graphics applications like NVIDIA. And so the overall architecture starting with single core doing what was at the time reasonably large systolic arrays …”
Reiner Pope Feb 26, 2026 ▶ 1:10
Opinion
Pope: TPUv1 announcement catalyzed the 2016-2017 AI chip startup wave
“TPUv one was announced in 2016, I think. That was what actually kind of led to the creation of all of those 2016, 20 17 startups. So Cerebus, Gronk, Graphcore, SambaNova, all of those.”
Reiner Pope Feb 26, 2026 ▶ 1:37
Assertion Supported
Pope: Google built TPUv1 in 12-18 months with 20-30 engineers
“TPUv one actually was, I think, is a really impressive project. It was done on a very short timeline, maybe, I don't know the full details, but maybe about a year or so, maybe a year and a half with a skeleton team of 20, 30 people.”
Reiner Pope Feb 26, 2026 ▶ 1:50
Assertion Contradicted
Pope: Google completely stopped publishing its AI research around 2022
“In twenty-twenty-two was about the time when just Google completely stopped publishing its research. And so all the good papers are from before that as a result.”
Reiner Pope Feb 26, 2026 ▶ 3:13
Insight
Pope: Chip physics enforces parallelism because signals take 100 cycles to cross
“So, I mean, it is just true hardware is massively parallel. Like, you've got tens of billions, hundreds of billions of transistors on your chip, and it takes, like, maybe a hundred clock cycles to get from one side of the chip to the other, and so you can't, l…”
Reiner Pope Feb 26, 2026 ▶ 3:41
Insight
Pope: Instruction control dominates CPU cost, unlike GPUs with large payloads
“Reading what do I have to do next? Okay, how do I do that? That is most of the cost on a CPU, whereas if you just keep the same instructions but make the payload a hundred times bigger, then you can shift most of the cost to be in the actual work that you want…”
Reiner Pope Feb 26, 2026 ▶ 7:14
Insight
Pope: The best AI inference chip is also a great training chip
“I think the best inference chip today will be a train, a really good training chip as well.”
Reiner Pope Feb 26, 2026 ▶ 11:16
Prediction Open · timeframe Feb 2029
Pope: MatX will sell AI inference chips first due to lower risk
“Our product is both training and inference, but I think the first sales will be an inference. That's mostly just a market effect where It's easier to buy, like, it's not as big of a risk to go to buy an inference cluster than as a training cluster.”
Reiner Pope Feb 26, 2026 ▶ 11:23
Disclosure
Pope: MatX raised a $500M Series B co-led by Jane Street
“So we this is a we've raised a series B round. It's led by Jane Street and Situational Awareness Situational awareness, that is Leopold Ashenbrenner's fund. He wrote the definitive book on, on, on where, on AGI and where it's going. And then Jane Street, they'…”
Reiner Pope Feb 26, 2026 ▶ 11:47
Opinion
Pope: Groq and Cerebras are uncompetitive on dollars per token
“And then there's the Grok and Cerebris that are much better at latency because they've got this the SRAM, weights are in SRAM very low latency. The problem is, and the challenge when you go to a Grok or a Cerebris system is that the throughput you get there, i…”
Reiner Pope Feb 26, 2026 ▶ 15:49
Disclosure
Pope: MatX combines SRAM and HBM for cheap, low-latency AI chips
“It is actually possible to do both in the same chip. You, it's kind of an obvious thing. You say you take the HPM, you take the SRAM, put them together in the same chip, you put the weights in SRAM, and you put the all of the inference data in HPM. That is wha…”
Reiner Pope Feb 26, 2026 ▶ 16:12
Assertion Supported
Pope: SRAM is an order of magnitude faster per token than HBM
“There's, yeah, there's just some simple math of, like, how long does it take you to read through all of HBM? It takes about 20 milliseconds, and so that's the amount of time per token it runs. Yes. Whereas the amount of time to read through all of SRAM is much…”
Reiner Pope Feb 26, 2026 ▶ 16:57
Prediction Not checkable as stated
Pope predicts supply crunches across the entire AI hardware supply chain
“The supply chain we're gonna have crunches on, on all of the supply chain, really. So if you look at the, sort of, the big components of what any company, but like us, for example, build-out there is dependency on Logic Dies from typically TSMC, maybe Samsung,…”
Reiner Pope Feb 26, 2026 ▶ 18:54
Insight
Pope: Large systolic arrays are unbeatable in area and power efficiency
“Make a really large systolic array. You can't beat that in area or power efficiency.”
Reiner Pope Feb 26, 2026 ▶ 22:34
Insight
Pope: Mixture of experts maps well to systolic arrays, attention does not
“The mixture of expert layer maps really well, but the attention does not.”
Reiner Pope Feb 26, 2026 ▶ 23:16
Disclosure
Pope: MatX splits large systolic arrays without sacrificing efficiency
“Take a really large systolic array, but have a way to split it up into pieces without losing efficiency. So sort of that is the core of the design for us.”
Reiner Pope Feb 26, 2026 ▶ 23:28
Prediction Not checkable as stated
Pope: 4-bit precision will probably be the primary AI format, like Nvidia
“We think probably the main thing will be similar to where NVIDIA is at, which is four bit precision.”
Reiner Pope Feb 26, 2026 ▶ 25:01
Assertion Partly supported
Pope: Chip tape-outs cost $30M and fail 50% of the time
“The ideal, which companies tend to hit about 50% of the time is that your first tape out costs like thirty million dollars your first... The ideal is that your first tape out is actually, is your production thing. So you do a tape out, you make maybe a thousan…”
Reiner Pope Feb 26, 2026 ▶ 31:21
Assertion Supported
Pope: Metal layer chip respins cost $100K versus $30M full tape-outs
“In good cases and in many cases, you can redo just the metal layers, which costs you only like a 100,000 dollars. As opposed to the- Pay the thirty million dollars again. But in bad cases, like, if you've made something serious and you can't fix that at the me…”
Reiner Pope Feb 26, 2026 ▶ 31:57
Assertion Supported
Pope: OpenAI is starting to design its own custom chips
“Google does. OpenAI is starting.”
Reiner Pope Feb 26, 2026 ▶ 40:30
Assertion Partly supported
Pope: NVIDIA includes eight spare chips in a 64-chip rack
“Nvidia has eight spare chips and a rack of 64.”
Reiner Pope Feb 26, 2026 ▶ 42:08
Assertion Not checkable as stated
Pope: AI labs apply RL to software code, not chip architecture
“Models are extremely good at Rust and Python. They've done a lot of RL on them. They have not done as much RL on Verilog. They've done almost none on, okay, write, write me a markdown file that describes a chip architecture.”
Reiner Pope Feb 26, 2026 ▶ 45:39
Assertion Supported
Pope: Foundry turnaround from tape-out to first chips is 4-5 months
“What we see is time from, like, tape out to first chip, to chips back. Again, depends on node, but it's ballpark four or five months.”
Reiner Pope Feb 26, 2026 ▶ 51:28
Prediction Open · timeframe Feb 2029
Pope: Model parameter counts will grow much faster than context lengths
“Really tied into this context thing, I think the context size will stay ballpark the same way it is, maybe a few times larger. But the parameter count will go up. Like, parameter count should grow much, much faster than context length, actually, just because o…”
Reiner Pope Feb 26, 2026 ▶ 54:27
Disclosure
Pope: MatX trains small LLMs from scratch daily for hardware co-design
“So our ML team is actual, real ML research. What they do every day is they train small LLMs from scratch, focusing on numerics and attention.”
Reiner Pope Feb 26, 2026 ▶ 56:57
Made with StarZero

Turn any episode into a week of clips.

This entire site, about 28 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.