Prediction certainty 3/5 debate potential 3/5

Biderman: Model accuracy will still degrade at 10M context window scale

Dan Biderman · The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO · Jul 13, 2026 · at 17:31

Dan Biderman, co-founder and CEO of Engram, discusses the architectural limits of long-context LLMs and why scaling context windows does not solve context rot.

0:00 / 0:26exact quote · 26.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“But two is like, for the agentic tasks of 18 months from now, inside those major repositories of knowledge, and asking the models more and more things in underspecified ways, I suspect that the accuracy of the models would go down. The phenomenon of context fraud, right? The model has to read more, it will be less accurate, and we know this, and it will remain the same thing at, even at the ten million context window scale.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dan Biderman

Prediction Not checkable as stated
Biderman: AI-native companies will amass trillions of internal tokens within 18 months
“In 18 months, many companies would have maybe trillions of tokens, which of internal company data, proprietary data. I'm talking about like maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.”
Dan Biderman Jul 13, 2026 ▶ 14:42 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Not checkable as stated
Biderman: Hard engineering tasks will require test-time gradient updates
“We think that eventually part of the solution for very hard tasks in, in science and engineering and defense and all that stuff will involve some form of gradient based updates during during doing these long horizon tasks.”
Dan Biderman Jul 13, 2026 ▶ 21:57 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Not checkable as stated
Biderman: In 18 months, data scale will require weight-based learning
“Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.”
Dan Biderman Jul 13, 2026 ▶ 26:55 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Open · timeframe Jul 2031
Biderman: PC hardware will soon run near-trillion-parameter models locally
“And in the long, long term, I do think these things will actually run on people's devices, and we're seeing right now the new hardware on personal computers is already, ah, you know, soon approaching the ability to run inference on close to trillion parameters…”
Dan Biderman Jul 13, 2026 ▶ 28:12 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Disclosure
Biderman: Engram trains models to decide what to memorize vs keep in notes
“The way to work on it is to train models both, to train models to manage it themselves, and that's an active area for us. Have the model know, like, without any explicit supervision signal to determine this kind of stuff I can pull from my brain, and that kind…”
Dan Biderman Jul 13, 2026 ▶ 30:56 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Insight
Biderman: Manually partitioning LLM memory vs retrieval becomes unmanageable whack-a-mole
“And now the thing is, if you start manually, heuristically saying this is in, this is out, then it becomes a whack-a-mole. Every, every person in every enterprise has different data, and you can really very easily pick and choose what goes in and what goes out…”
Dan Biderman Jul 13, 2026 ▶ 31:47 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.