“Since then, we've released several models, recent model lines in called Jamba, which I think the fascinating part about it is, is the first hybrid model. It's not just attention.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Barak Lenz
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Barak LenzOct 11, 2025▶ 40:33Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
AssertionNot checkable as stated
Lenz: Local smartphone AI requires hybrid models due to KV cache limits
“So if you wanted to do something local on your phone to search your images, as an example, you can't do that without a hybrid architecture or without doing drastically changes because the model plus KVCache won't fit.”
Barak LenzOct 11, 2025▶ 13:09Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Opinion
Lenz: Most enterprises avoid reasoning models due to high latency
“Most enterprises don't really want to use reasoning models. The latencies is too high”
Barak LenzOct 11, 2025▶ 33:15Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
PredictionNot checkable as stated
Lenz: Hybrid Transformer models are here to stay for long context
“If I had to guess hybrid models are here to stay just because the efficiency without sacrificing the performance is, is too much to give up. You know, the attention is so expensive. The quadratic cost and the linear memory cost is so much that I think for long…”
Barak LenzOct 11, 2025▶ 9:13Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Insight
Lenz: Middle attention placement at 1:8 ratio optimizes hybrid models
“Putting it the first or the last performed worse than in the middle. And one to eight was good enough. You know, you might get very slight improvements with one to six, but it was marginal, maybe within the standard deviation.”
Barak LenzOct 11, 2025▶ 10:16Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
PredictionNot checkable as stated
Lenz: Full attention models will decline as sequence lengths rise
“I can definitely see sequence length rising, and I can't see full attention models being as prominent as they are today. So, so, so at least they'll have less full attention layers and I hope they'll have more innovations like Mumbai.”
Barak LenzOct 11, 2025▶ 11:03Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.