why aren't all 8 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
AI21 CTO: AI systems must be model-agnostic and action-oriented
“AI systems need to be model agnostic. They shouldn't care about which model that they use, and they should look at what I call actions, which is a combination of a model with a prompt and maybe a set of tools that it can use and say, what can an action do for …”
Assertion Not checkable as stated
Lenz: No open-source infrastructure can train very large models
“And we're using our own infrastructure to train our models. There still isn't an open source infrastructure that I could say, use this to train your very, very large model.”
Insight
Lenz: Engineers seeking online answers are not at the frontier
“And I'm looking for people that try to solve problems on their own. Because if you think you're gonna found the answers online, I think you're not in the frontier.”
Insight
Lenz: Middle attention placement at 1:8 ratio optimizes hybrid models
“Putting it the first or the last performed worse than in the middle. And one to eight was good enough. You know, you might get very slight improvements with one to six, but it was marginal, maybe within the standard deviation.”
Prediction Not checkable as stated
Lenz: Full attention models will decline as sequence lengths rise
“I can definitely see sequence length rising, and I can't see full attention models being as prominent as they are today. So, so, so at least they'll have less full attention layers and I hope they'll have more innovations like Mumbai.”
Prediction Not checkable as stated
Lenz: Hybrid Transformer models are here to stay for long context
“If I had to guess hybrid models are here to stay just because the efficiency without sacrificing the performance is, is too much to give up. You know, the attention is so expensive. The quadratic cost and the linear memory cost is so much that I think for long…”
Assertion Contradicted
Lenz: AI21's Jamba is the first hybrid model architecture
“Since then, we've released several models, recent model lines in called Jamba, which I think the fascinating part about it is, is the first hybrid model. It's not just attention.”
Disclosure
Lenz: AI21 will release a 3B dense Jamba model for edge devices
“And we're actually releasing new versions of our Jamba models soon, including a three B dense model. It's aimed for long context on edge devices.”