LSTM
topic on 6 shows · 10 statements across 9 episodes
Acquired
the Y Combinator Startup Podcast
Latent Space
No Priors
the MAD Podcast
the a16z Podcast
10 statements about LSTM, every show
AI capabilities would have been achieved even without inventing transformers
“I think if we hadn't invented the transformer, we would have gotten there with whatever LSTM you know, state space model, whatever, anything else people were developing, we would have gotten there.”
Transformers perform all Bayesian tasks, Mamba does most, and MLPs fail
“Transformer does everything. Mamba does most of it. LSTMs do only partially, and MLPs fail completely.”
Dean: Transformers delivered 10x to 100x compute efficiency over LSTMs
“Transformers similarly gave you a 10 X to a hundred X improvement in, you know compute cost to a given quality level versus say LSTMs at the time.”
Incorporating LSTMs Cut Google Translate Error Rate by 60%
“And indeed, in 2016, they incorporated into Google Translate these LSTMs. It reduces the error rate by 60%.”
Socher: Transformer results would take 10x compute and engineering with LSTMs
“Probably if it wasn't for transformers, it would have just been like 10 X more engineering and data needed to get to similar results, even with like past models like LSTMs and so on.”
Karpathy: Clean AI scaling laws are a property of transformers, not LSTMs
“When people talk about the scaling loss in neural networks, the scaling laws are actually a to a large extent of a property of the transformer. Before the transformer, people were playing with LSTMs and stacking them, etc. You don't actually get like clean sca…”
Vinyals: RNNs and LSTMs Never Remembered Beyond a Few Hundred Words
“We come from a world where we had recurrent neural networks and LSTMs that actually had infinite memory, although it was not very capable, right? You, the models in, in practice, they never remember more than a few hundred words or so.”
Polosukhin: LSTMs Were Too Slow for Production as Documents Scaled
“The state of the art at this time was LSTMs, Recurring Neural Networks, which you could not launch in production at all because they're too slow and take a fair bit of time to process as documents scale.”
Eck: Alex Graves advanced LSTMs more than anyone, including its creator
“Among the three of us, by far, Alex Graves has done the most with LSTM. So he continued, after he finished his PhD, and he continued doggedly to try to understand how recurrent neural networks worked, how to train them, and how to make them useful for sequence…”