Everything Albert Gu said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Gu: Transformers struggle significantly on raw pixel or audio waveform data
“People think that like you can throw a transformer at like anything and it just works. Actually it doesn't really like if you try to throw it at like the raw pixel level or the raw sample level and in audio waveforms I think it doesn't work nearly as well.”
Gu: High-quality speech synthesis requires multimodal foundation models
“And so actually to really get, like, perfect even just TTS or, like, speech-to-speech you actually really need to have, like, a model that has, More understanding, like at least of the language, but kind of like, it's not really an isolated component anymore. …”
Gu: Optimal hybrid models use a 10:1 ratio of SSM to attention
“People have found that the optimal ratio tends to be mostly SSM layers with a little bit of attention. So maybe a ratio of like 10 to one, I know of at least Probably like five groups that have independently verified that this is kind of the optimal ratio of t…”
Gu: Mamba successfully applied state-space models to language modeling
“Recently proposed a model called Mamba which was kind of brought these to language modeling and showed really good results there.”
Gu: State Space Models can be applied to almost all data types
“So it really can be applied to pretty much everything. So just like kind of Transformers, these are applied to everything. So can these sort of models over the course of research over a few years, we kind of realized that there are different advantages for dif…”
Gu: Aesthetic elegance was the primary driver behind inventing State Space Models
“People ask me, like, how do I treat my research problems, and my, I can't explain. My answer is just aesthetic. It's just like, there's something that I find elegant, and we're aesthetically pleasing about things, and to me, that's almost the most important th…”
Gu: Early SSMs excelled at raw signals but lagged Transformers on text
“The first types of models we were looking at were really good actually at modeling kind of these raw waveforms raw pixels, things like that, but not as good at modeling text, and transformers are way better there.”
Gu: Researchers are applying Mamba-based foundation models to DNA sequences
“So I actually just heard from some collaborators today that they applied a mama based model on DNA modeling. They're basically bringing this idea of foundation models To DNA, which is kind of this new idea.”