Insight certainty 4/5 debate potential 2/5

Liu: Long-context recall depends directly on needle-in-haystack training loss

Beyang Liu · The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph · Dec 17, 2023 · at 19:44

Beyang Liu (Sourcegraph co-founder and CTO) discusses context ranking and positional loss variability across LLMs.

0:00 / 0:28exact quote · 29.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The skill with which models are able to take advantage of context is always going to be dependent on how that factors into the impact on the training loss, right? So like, If you want long context window models to work well, then you have to have a ton of data where it's like, here's, like, a billion lines of text, and I'm gonna ask a question about, like, something that's, like, you know, embedded deeply into it, and, like, give me the right answer. And unless you have that training set, then, of course, you're gonna have variability in terms of, like, where it attends to”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Beyang Liu

Assertion Not checkable as stated
Liu: Cody matches GitHub Copilot completion acceptance rates using open-source StarCoder
“Like today, Cody uses StarCoder for inline completions, and with the benefit of the context that we provide, we actually show, like, comparable completion acceptance rate metrics. It's kind of like the standard metric that folks use to evaluate inline completi…”
Beyang Liu Dec 17, 2023 ▶ 15:22 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Opinion
Liu: Pure transformer models are insufficient to support autonomous AI agents
“We're actually a little bit, I think, more bearish than the average, you know, AI hypefluencer out there on the feasibility of agents with purely kind of like transformer-based models.”
Beyang Liu Dec 17, 2023 ▶ 24:08 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Prediction Not checkable as stated
Liu: Reliable AI coding workflows require search-based algorithmic backbones
“The way that we get to this, like, more reliable, multi-step workflows that can do things beyond, you know, generate unit test is, is, it's really gonna be, like, a search-based approach, where, where you use an LLM as, kind of, like, an advisor or a proposal …”
Beyang Liu Dec 17, 2023 ▶ 29:29 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Opinion
Liu: Open-source AI models are currently state-of-the-art for code completion
“Yeah, I mean, for completions, open source is, is state of the art right now.”
Beyang Liu Dec 17, 2023 ▶ 1:09:58 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Insight
Liu: Reliable single-step generation is a strict prerequisite for true AI agents
“If you want to get to the point where you can actually be truly agentic or like multi-step automated a necessary part of that is like the single step has to be robust and reliable.”
Beyang Liu Dec 17, 2023 ▶ 1:30:29 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Prediction Not checkable as stated
Liu: AI coding assistants must pull context beyond Git repositories to succeed
“And I don't think the AI developer will be any different. It will need to pull context from all these different sources.”
Beyang Liu Dec 17, 2023 ▶ 37:22 The "Normsky" architecture for AI coding agents — with Beyang Liu + Steve Yegge of SourceGraph
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.