Disclosure certainty 4/5 debate potential 2/5

Pope: MatX combines SRAM and HBM for cheap, low-latency AI chips

Reiner Pope · Reiner Pope of MatX on accelerating AI with transformer-optimized chips · Feb 26, 2026 · at 16:12

Reiner Pope is the co-founder and CEO of MatX, developing specialized chip architectures for large language models.

0:00 / 0:19exact quote · 19.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“It is actually possible to do both in the same chip. You, it's kind of an obvious thing. You say you take the HPM, you take the SRAM, put them together in the same chip, you put the weights in SRAM, and you put the all of the inference data in HPM. That is what we are doing, in fact and I think that actually hits a really nice sweet spot where you can get a little latency and also be very cheap.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Reiner Pope

Opinion
Pope: Groq and Cerebras are uncompetitive on dollars per token
“And then there's the Grok and Cerebris that are much better at latency because they've got this the SRAM, weights are in SRAM very low latency. The problem is, and the challenge when you go to a Grok or a Cerebris system is that the throughput you get there, i…”
Reiner Pope Feb 26, 2026 ▶ 15:49 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
Assertion Contradicted
Pope: Google completely stopped publishing its AI research around 2022
“In twenty-twenty-two was about the time when just Google completely stopped publishing its research. And so all the good papers are from before that as a result.”
Reiner Pope Feb 26, 2026 ▶ 3:13 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
Prediction Open · timeframe Feb 2029
Pope: Model parameter counts will grow much faster than context lengths
“Really tied into this context thing, I think the context size will stay ballpark the same way it is, maybe a few times larger. But the parameter count will go up. Like, parameter count should grow much, much faster than context length, actually, just because o…”
Reiner Pope Feb 26, 2026 ▶ 54:27 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
Insight
Pope: The best AI inference chip is also a great training chip
“I think the best inference chip today will be a train, a really good training chip as well.”
Reiner Pope Feb 26, 2026 ▶ 11:16 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
Insight
Pope: Mixture of experts maps well to systolic arrays, attention does not
“The mixture of expert layer maps really well, but the attention does not.”
Reiner Pope Feb 26, 2026 ▶ 23:16 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
Disclosure
Pope: MatX splits large systolic arrays without sacrificing efficiency
“Take a really large systolic array, but have a way to split it up into pieces without losing efficiency. So sort of that is the core of the design for us.”
Reiner Pope Feb 26, 2026 ▶ 23:28 Reiner Pope of MatX on accelerating AI with transformer-optimized chips
Made with StarZero

Turn any episode into a week of clips.

This entire site, about 28 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.