Oct 11, 2025 · 45m · latent-space

Building Jamba 3B: the tiny Hybrid Transformer State Space Reasoning Model - Barak Lenz, CTO of AI21

Barak Lenz · 34m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Barak Lenz, CTO of AI21 Labs, explores the architectural design of hybrid Transformer-Mamba models like Jamba 3B and the Maestro enterprise orchestration platform. He outlines how AI21 solves memory bottlenecks for long-context edge computing, applies quantitative experimentation rigor, and builds closed-loop AI systems for enterprises.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.0 Guest teaching 5.2 Guest disagreement 1.5 The hosts pushing back 2.0
05100:0015:0030:0045:004:23–12:01 · The hosts as informed peer 5/10 Origins and Advantages of Hybrid Transformer-Mamba Models The host demonstrates domain familiarity by citing Mixtral, the hardware lottery, and Noam Shazeer's 1:8 attention ratio paper. Barak provides detailed architectural insight on hybrid Mamba-attention ablations and optimal layer placement.12:01–15:46 · The hosts as informed peer 4/10 Unveiling Jamba 3B Dense for Edge and Long Context The host prompts on edge deployment features while Barak delivers an in-depth breakdown of memory constraints. Barak explains how KV cache memory footprint matches total model parameter size at modest context lengths on 3B edge models.15:46–21:29 · The hosts as informed peer 5/10 Experimentation Methodology and Frontier Model Training Scale The host probes for exact cluster size and cites DeepSpeed and Imbue open-source stacks. Barak dodges the exact hardware figures and counters that public recipes cannot replace custom frontier engineering.21:29–28:30 · The hosts as informed peer 6/10 Technical Hiring Philosophy and Transition from Algorithmic Trading The host explicitly pushes back against Barak's unstructured interview method, arguing founders need structured first-principle questions to attract talent. Both then bond over quantitative trading analogies, regime change, and alpha decay.28:30–41:44 · The hosts as informed peer 6/10 Building Model-Agnostic AI Systems with Maestro The host references AI21's blog post and challenges RL search spaces regarding calibration and the fog of war. Barak reframes brute-force RL rollouts into a principled multi-armed bandit and action-planning paradigm.41:44–44:51 · The hosts as informed peer 4/10 Enterprise Vision for Continuous Learning and Conclusion The host clarifies model agnosticism in enterprise setups as Barak outlines Maestro's continuous learning loop vision. The segment concludes on a friendly, collaborative note.4:23–12:01 · Guest teaching 5/10 Origins and Advantages of Hybrid Transformer-Mamba Models The host demonstrates domain familiarity by citing Mixtral, the hardware lottery, and Noam Shazeer's 1:8 attention ratio paper. Barak provides detailed architectural insight on hybrid Mamba-attention ablations and optimal layer placement.12:01–15:46 · Guest teaching 6/10 Unveiling Jamba 3B Dense for Edge and Long Context The host prompts on edge deployment features while Barak delivers an in-depth breakdown of memory constraints. Barak explains how KV cache memory footprint matches total model parameter size at modest context lengths on 3B edge models.15:46–21:29 · Guest teaching 6/10 Experimentation Methodology and Frontier Model Training Scale The host probes for exact cluster size and cites DeepSpeed and Imbue open-source stacks. Barak dodges the exact hardware figures and counters that public recipes cannot replace custom frontier engineering.21:29–28:30 · Guest teaching 4/10 Technical Hiring Philosophy and Transition from Algorithmic Trading The host explicitly pushes back against Barak's unstructured interview method, arguing founders need structured first-principle questions to attract talent. Both then bond over quantitative trading analogies, regime change, and alpha decay.28:30–41:44 · Guest teaching 6/10 Building Model-Agnostic AI Systems with Maestro The host references AI21's blog post and challenges RL search spaces regarding calibration and the fog of war. Barak reframes brute-force RL rollouts into a principled multi-armed bandit and action-planning paradigm.41:44–44:51 · Guest teaching 4/10 Enterprise Vision for Continuous Learning and Conclusion The host clarifies model agnosticism in enterprise setups as Barak outlines Maestro's continuous learning loop vision. The segment concludes on a friendly, collaborative note.4:23–12:01 · Guest disagreement 1/10 Origins and Advantages of Hybrid Transformer-Mamba Models The host demonstrates domain familiarity by citing Mixtral, the hardware lottery, and Noam Shazeer's 1:8 attention ratio paper. Barak provides detailed architectural insight on hybrid Mamba-attention ablations and optimal layer placement.12:01–15:46 · Guest disagreement 1/10 Unveiling Jamba 3B Dense for Edge and Long Context The host prompts on edge deployment features while Barak delivers an in-depth breakdown of memory constraints. Barak explains how KV cache memory footprint matches total model parameter size at modest context lengths on 3B edge models.15:46–21:29 · Guest disagreement 3/10 Experimentation Methodology and Frontier Model Training Scale The host probes for exact cluster size and cites DeepSpeed and Imbue open-source stacks. Barak dodges the exact hardware figures and counters that public recipes cannot replace custom frontier engineering.21:29–28:30 · Guest disagreement 2/10 Technical Hiring Philosophy and Transition from Algorithmic Trading The host explicitly pushes back against Barak's unstructured interview method, arguing founders need structured first-principle questions to attract talent. Both then bond over quantitative trading analogies, regime change, and alpha decay.28:30–41:44 · Guest disagreement 2/10 Building Model-Agnostic AI Systems with Maestro The host references AI21's blog post and challenges RL search spaces regarding calibration and the fog of war. Barak reframes brute-force RL rollouts into a principled multi-armed bandit and action-planning paradigm.41:44–44:51 · Guest disagreement 0/10 Enterprise Vision for Continuous Learning and Conclusion The host clarifies model agnosticism in enterprise setups as Barak outlines Maestro's continuous learning loop vision. The segment concludes on a friendly, collaborative note.4:23–12:01 · The hosts pushing back 1/10 Origins and Advantages of Hybrid Transformer-Mamba Models The host demonstrates domain familiarity by citing Mixtral, the hardware lottery, and Noam Shazeer's 1:8 attention ratio paper. Barak provides detailed architectural insight on hybrid Mamba-attention ablations and optimal layer placement.12:01–15:46 · The hosts pushing back 0/10 Unveiling Jamba 3B Dense for Edge and Long Context The host prompts on edge deployment features while Barak delivers an in-depth breakdown of memory constraints. Barak explains how KV cache memory footprint matches total model parameter size at modest context lengths on 3B edge models.15:46–21:29 · The hosts pushing back 3/10 Experimentation Methodology and Frontier Model Training Scale The host probes for exact cluster size and cites DeepSpeed and Imbue open-source stacks. Barak dodges the exact hardware figures and counters that public recipes cannot replace custom frontier engineering.21:29–28:30 · The hosts pushing back 5/10 Technical Hiring Philosophy and Transition from Algorithmic Trading The host explicitly pushes back against Barak's unstructured interview method, arguing founders need structured first-principle questions to attract talent. Both then bond over quantitative trading analogies, regime change, and alpha decay.28:30–41:44 · The hosts pushing back 3/10 Building Model-Agnostic AI Systems with Maestro The host references AI21's blog post and challenges RL search spaces regarding calibration and the fog of war. Barak reframes brute-force RL rollouts into a principled multi-armed bandit and action-planning paradigm.41:44–44:51 · The hosts pushing back 0/10 Enterprise Vision for Continuous Learning and Conclusion The host clarifies model agnosticism in enterprise setups as Barak outlines Maestro's continuous learning loop vision. The segment concludes on a friendly, collaborative note.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 0%45:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 20:43 Barak rejects open-source training recipe parity

Barak pushes back against the host's suggestion that public frameworks like DeepSpeed or Verl replicate the custom engineering required to train large models robustly.

Hardest push from the hosts ▶ 22:41 Host challenges unstructured interviewing approach

The host explicitly states his pushback against Barak's conversational interview style, arguing founders must provide clear recruiting anchor questions to attract frontier talent.

Biggest teaching moment ▶ 12:12 Barak explains KV cache explosion on edge devices

Barak educates the host on edge constraints, detailing how image tokenization and 16k context KV caches quickly equal the entire memory footprint of a 3B model.

The host holds their own ▶ 27:36 Host displays quantitative trading domain knowledge

The host draws on his fund background to analyze market regime shifts, static datasets, and hidden beta risk in algorithmic trading.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins and Advantages of Hybrid Transformer-Mamba Models 5511 The host demonstrates domain familiarity by citing Mixtral, the hardware lottery, and Noam Shazeer's 1:8 attention ratio paper. Barak provides detailed architectural insight on hybrid Mamba-attention ablations and optimal layer placement.
Unveiling Jamba 3B Dense for Edge and Long Context 4610 The host prompts on edge deployment features while Barak delivers an in-depth breakdown of memory constraints. Barak explains how KV cache memory footprint matches total model parameter size at modest context lengths on 3B edge models.
Experimentation Methodology and Frontier Model Training Scale 5633 The host probes for exact cluster size and cites DeepSpeed and Imbue open-source stacks. Barak dodges the exact hardware figures and counters that public recipes cannot replace custom frontier engineering.
Technical Hiring Philosophy and Transition from Algorithmic Trading 6425 The host explicitly pushes back against Barak's unstructured interview method, arguing founders need structured first-principle questions to attract talent. Both then bond over quantitative trading analogies, regime change, and alpha decay.
Building Model-Agnostic AI Systems with Maestro 6623 The host references AI21's blog post and challenges RL search spaces regarding calibration and the fog of war. Barak reframes brute-force RL rollouts into a principled multi-armed bandit and action-planning paradigm.
Enterprise Vision for Continuous Learning and Conclusion 4400 The host clarifies model agnosticism in enterprise setups as Barak outlines Maestro's continuous learning loop vision. The segment concludes on a friendly, collaborative note.

Statements from this episode (16)

Assertion Contradicted
Lenz: AI21's Jamba is the first hybrid model architecture
“Since then, we've released several models, recent model lines in called Jamba, which I think the fascinating part about it is, is the first hybrid model. It's not just attention.”
Barak Lenz Oct 11, 2025 ▶ 2:20
Disclosure
Lenz: AI21 will release a 3B dense Jamba model for edge devices
“And we're actually releasing new versions of our Jamba models soon, including a three B dense model. It's aimed for long context on edge devices.”
Barak Lenz Oct 11, 2025 ▶ 2:33
Prediction Not checkable as stated
Lenz: Hybrid Transformer models are here to stay for long context
“If I had to guess hybrid models are here to stay just because the efficiency without sacrificing the performance is, is too much to give up. You know, the attention is so expensive. The quadratic cost and the linear memory cost is so much that I think for long…”
Barak Lenz Oct 11, 2025 ▶ 9:13
Insight
Lenz: Middle attention placement at 1:8 ratio optimizes hybrid models
“Putting it the first or the last performed worse than in the middle. And one to eight was good enough. You know, you might get very slight improvements with one to six, but it was marginal, maybe within the standard deviation.”
Barak Lenz Oct 11, 2025 ▶ 10:16
Prediction Not checkable as stated
Lenz: Full attention models will decline as sequence lengths rise
“I can definitely see sequence length rising, and I can't see full attention models being as prominent as they are today. So, so, so at least they'll have less full attention layers and I hope they'll have more innovations like Mumbai.”
Barak Lenz Oct 11, 2025 ▶ 11:03
Assertion Not checkable as stated
Lenz: Local smartphone AI requires hybrid models due to KV cache limits
“So if you wanted to do something local on your phone to search your images, as an example, you can't do that without a hybrid architecture or without doing drastically changes because the model plus KVCache won't fit.”
Barak Lenz Oct 11, 2025 ▶ 13:09
Assertion Partly supported
AI21 CTO: Jamba 3B uses 1:12 attention ratio to cut memory footprint
“So this is a three B dense with only two attentional layers. It's one to 12 and not one to eight, because we wanted to maximize the efficiency. It has very few attention heads, so everything is geared To have, you know, long context with very little memory.”
Barak Lenz Oct 11, 2025 ▶ 15:08
Insight
Lenz: 3B models cannot be generalized and must be task-tailored
“Because again, three B models, they need to be tailored for tasks. It's not like you could squeeze whatever you wanted into them.”
Barak Lenz Oct 11, 2025 ▶ 15:36
Assertion Not checkable as stated
Lenz: No open-source infrastructure can train very large models
“And we're using our own infrastructure to train our models. There still isn't an open source infrastructure that I could say, use this to train your very, very large model.”
Barak Lenz Oct 11, 2025 ▶ 20:10
Insight
Lenz: Engineers seeking online answers are not at the frontier
“And I'm looking for people that try to solve problems on their own. Because if you think you're gonna found the answers online, I think you're not in the frontier.”
Barak Lenz Oct 11, 2025 ▶ 21:55
Insight
Lenz: Training foundation models mirrors developing algorithmic trading strategies
“So, so it's very similar in terms of how you interpret results. You want to treat everything as a black box. You want to establish your bounds, you know, what are you, and I've had tons of experience in algotrading, both from making money and not making money …”
Barak Lenz Oct 11, 2025 ▶ 25:52
Assertion Not checkable as stated
Lenz: GPT-4o is an AI orchestration system, not a raw model
“GPT-IV-O is already an AI system. It's not calling a model directly. It can do certain, you know, it can do tool calls. It can orchestrate this entire thing.”
Barak Lenz Oct 11, 2025 ▶ 30:09
Insight
AI21 CTO: AI systems must be model-agnostic and action-oriented
“AI systems need to be model agnostic. They shouldn't care about which model that they use, and they should look at what I call actions, which is a combination of a model with a prompt and maybe a set of tools that it can use and say, what can an action do for …”
Barak Lenz Oct 11, 2025 ▶ 31:46
Opinion
Lenz: Most enterprises avoid reasoning models due to high latency
“Most enterprises don't really want to use reasoning models. The latencies is too high”
Barak Lenz Oct 11, 2025 ▶ 33:15
Insight
Lenz: RL training wastes compute on saturated or impossible examples
“Once you've trained a few hundred steps of let's say GOP, Most of your training is just wasted on example that are either too hard for you and you didn't get any success on them or too easy and everything was a success.”
Barak Lenz Oct 11, 2025 ▶ 37:49
Opinion
Lenz: Model providers should not dictate enterprise AI policies
“Right now, if you're using a model, you're taking in their own policy. Even if I want to use GPT-OSS, I've taken in a lot of different policies about what to abstain from, what's considered dangerous and not dangerous, how I should behave, etc. And I don't thi…”
Barak Lenz Oct 11, 2025 ▶ 40:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.