Liu: Long-context recall depends directly on needle-in-haystack training loss
Beyang Liu · The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph · Dec 17, 2023 · at 19:44
Beyang Liu (Sourcegraph co-founder and CTO) discusses context ranking and positional loss variability across LLMs.
“The skill with which models are able to take advantage of context is always going to be dependent on how that factors into the impact on the training loss, right? So like, If you want long context window models to work well, then you have to have a ton of data where it's like, here's, like, a billion lines of text, and I'm gonna ask a question about, like, something that's, like, you know, embedded deeply into it, and, like, give me the right answer. And unless you have that training set, then, of course, you're gonna have variability in terms of, like, where it attends to”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →