Yi Tay: Gradient descent learning paradigm is AI's bottleneck, not architecture
Yi Tay · Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay · Jan 23, 2026 · at 49:07
Yi Tay, Senior Research Scientist at Google DeepMind, discusses whether scaling context windows to hundreds of millions of tokens requires replacing Transformer architectures or changing learning algorithms.
“It's not architecture itself. That's, that there's a problem that we, that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the tokens. I think it's more about the learning algorithm itself.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →