Huang: Context scaling requires positional interpolation rather than extrapolation
Mark Huang · How to train a Million Context LLM — with Mark Huang of Gradient.ai · May 31, 2024 · at 21:04
Mark Huang, co-founder of Gradient.ai, discusses technical mechanisms of RoPE scaling and positional embeddings for expanding LLM context windows.
“There's it's super confusing, but it's like, there's extrapolate positional extrapolation, and then there's interpolation. You want interpolation. It's been shown that just pure extrapolation makes the model a lot worse, and it's harder to attend to stuff, whereas the interpolation is like, you're squeezing everything back in to what the original context length was to a certain extent, and then allowing for it to overlap Different sequences that it's already seen as if it actually occurred when you see a billion contexts of sequence tokens.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →