Jamba
other on 1 show · 0 statements across 0 episodes
Mentions by year, every show
tap a year for its mentions
Latent Space 18
2026 2 mentions in 1 episode
2025 13 mentions in 2 episodes 7 per episode
every mention on every show, scene by scene, with the transcript →
4 statements about Jamba, every show
Lenz: Middle attention placement at 1:8 ratio optimizes hybrid models
“Putting it the first or the last performed worse than in the middle. And one to eight was good enough. You know, you might get very slight improvements with one to six, but it was marginal, maybe within the standard deviation.”
Lenz: AI21's Jamba is the first hybrid model architecture
“Since then, we've released several models, recent model lines in called Jamba, which I think the fascinating part about it is, is the first hybrid model. It's not just attention.”
Cheah: Hybrid SSM-transformer models outperform pure baselines of both
“None of us understand why a hybrid with a state-based model, the RWA state space, and transformer performs better than the baseline of both. It's like when you train one, you expect, and then you replace, you expect the same results. That's our pitch. That's o…”