Assertion certainty 4/5 debate potential 2/5

Biderman: Harmless enterprise queries on frontier models cost thousands of dollars

Dan Biderman · The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO · Jul 13, 2026 · at 25:50

Dan Biderman, CEO of Engram, argues that using brute-force long-context LLMs for holistic enterprise queries is prohibitively expensive.

0:00 / 0:11exact quote · 11.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“And now you can solve these tasks with frontier models and compaction. And when you ask them to do so, they will consume thousands of dollars for queries that we think are harmless. That every employee in the company would be able to answer.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dan Biderman

Prediction Not checkable as stated
Biderman: AI-native companies will amass trillions of internal tokens within 18 months
“In 18 months, many companies would have maybe trillions of tokens, which of internal company data, proprietary data. I'm talking about like maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.”
Dan Biderman Jul 13, 2026 ▶ 14:42 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Not checkable as stated
Biderman: Hard engineering tasks will require test-time gradient updates
“We think that eventually part of the solution for very hard tasks in, in science and engineering and defense and all that stuff will involve some form of gradient based updates during during doing these long horizon tasks.”
Dan Biderman Jul 13, 2026 ▶ 21:57 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Not checkable as stated
Biderman: In 18 months, data scale will require weight-based learning
“Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.”
Dan Biderman Jul 13, 2026 ▶ 26:55 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Prediction Open · timeframe Jul 2031
Biderman: PC hardware will soon run near-trillion-parameter models locally
“And in the long, long term, I do think these things will actually run on people's devices, and we're seeing right now the new hardware on personal computers is already, ah, you know, soon approaching the ability to run inference on close to trillion parameters…”
Dan Biderman Jul 13, 2026 ▶ 28:12 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Disclosure
Biderman: Engram trains models to decide what to memorize vs keep in notes
“The way to work on it is to train models both, to train models to manage it themselves, and that's an active area for us. Have the model know, like, without any explicit supervision signal to determine this kind of stuff I can pull from my brain, and that kind…”
Dan Biderman Jul 13, 2026 ▶ 30:56 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Insight
Biderman: Manually partitioning LLM memory vs retrieval becomes unmanageable whack-a-mole
“And now the thing is, if you start manually, heuristically saying this is in, this is out, then it becomes a whack-a-mole. Every, every person in every enterprise has different data, and you can really very easily pick and choose what goes in and what goes out…”
Dan Biderman Jul 13, 2026 ▶ 31:47 The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.