Mamba

also referred to as: mamba ssm

4 statements across 3 episodes · 3 bullish · 1 bearish · 3 people on the record · first statement Dec 24, 2024 by Eugene Cheah · said 30 times in 10 episodes since 2024 · across every show →

Mentions by year

brought up most by Barak Lenz (8), Shawn Wang (6), Quentin Anthony (3), Eugene Cheah (2), Vipul Ved Prakash (1), Mikhail Parakhin (1), Jack Morris (1), Dan Fu (1)

tap a year for its mentions
0083155202420252026episodesmentions
035202420252026episodes it came up in
0022.545202420252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Mamba, oldest first

Dec 24, 2024 positive
Insight
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”
Eugene Cheah Dec 24, 2024 ▶ 28:15 2024 in Post-Transformer Architectures: State Space Models, RWKV [Latent Space LIVE! @ NeurIPS 2024]
Oct 11, 2025 positive
Insight
Lenz: Middle attention placement at 1:8 ratio optimizes hybrid models
“Putting it the first or the last performed worse than in the middle. And one to eight was good enough. You know, you might get very slight improvements with one to six, but it was marginal, maybe within the standard deviation.”
Barak Lenz Oct 11, 2025 ▶ 10:16 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Oct 11, 2025 bullish
Prediction Not checkable as stated
Lenz: Hybrid Transformer models are here to stay for long context
“If I had to guess hybrid models are here to stay just because the efficiency without sacrificing the performance is, is too much to give up. You know, the attention is so expensive. The quadratic cost and the linear memory cost is so much that I think for long…”
Barak Lenz Oct 11, 2025 ▶ 9:13 Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Nov 3, 2025 negative
Opinion
Groq hardware is too inflexible to run alternative architectures like Mamba SSMs
“Grok is very inflexible hardware so that you kind of, if you want to do like a Mamba SSM on it, you're going to have a really hard time because instead of having like a low level CUDA compiler, like everything is like Designed it at the hardware level, right?”
Quentin Anthony Nov 3, 2025 ▶ 22:42 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.