Mamba

product on 5 shows · 7 statements across 5 episodes · said 40 times in 16 episodes since 2024

Latent Space 30 No Priors 4 the MAD Podcast 4 the a16z Podcast 1 20VC 1

Mentions by year, every show

tap a year for its mentions
00134258202420252026episodesmentions
048202420252026episodes it came up in
001.5438202420252026episodesmentions per episode

Latent Space 30the MAD Podcast 4No Priors 420VC 1the a16z Podcast 1

2026 3 mentions in 3 episodes 1 per episode
2025 15 mentions in 5 episodes 3 per episode
2024 22 mentions in 8 episodes 3 per episode

every mention on every show, scene by scene, with the transcript →

7 statements about Mamba, every show

a16z Assertion Supported
Transformers perform all Bayesian tasks, Mamba does most, and MLPs fail
“Transformer does everything. Mamba does most of it. LSTMs do only partially, and MLPs fail completely.”
Vishal Misra Mar 17, 2026 ▶ 21:13 Why Scale Will Not Solve AGI | Vishal Misra - The a16z Show
Groq hardware is too inflexible to run alternative architectures like Mamba SSMs
“Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like Designed it at the hardware level, right?”
Quentin Anthony Nov 3, 2025 ▶ 22:42 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
LATENT SPACE Prediction Not checkable as stated
Lenz: Hybrid Transformer models are here to stay for long context
“If I had to guess hybrid models are here to stay just because the efficiency without sacrificing the performance is, is too much to give up. You know, the attention is so expensive. The quadratic cost and the linear memory cost is so much that I think for long…”
Barak Lenz Oct 11, 2025 ▶ 9:13 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Lenz: Middle attention placement at 1:8 ratio optimizes hybrid models
“Putting it the first or the last performed worse than in the middle. And one to eight was good enough. You know, you might get very slight improvements with one to six, but it was marginal, maybe within the standard deviation.”
Barak Lenz Oct 11, 2025 ▶ 10:16 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
NO PRIORS Assertion Supported
Gu: Mamba successfully applied state-space models to language modeling
“Recently proposed a model called Mamba which was kind of brought these to language modeling and showed really good results there.”
Albert Gu Jun 27, 2024 ▶ 2:24 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu
NO PRIORS Assertion Supported
Gu: Researchers are applying Mamba-based foundation models to DNA sequences
“So I actually just heard from some collaborators today that they applied a mama based model on DNA modeling. They're basically bringing this idea of foundation models To DNA, which is kind of this new idea.”
Albert Gu Jun 27, 2024 ▶ 12:02 No Priors Ep. 70 | With Cartesia Co-Founders Karan Goel & Albert Gu

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.