Jul 13, 2026 · 49m · latent-space

The AI Memory Problem: Why Long Context Isn’t Enough — Dan Biderman, Engram Co-founder & CEO

Dan Biderman · 37m spoken Allen Park · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In a unique cooking show format, Engram co-founder and CEO Dan Biderman prepares Mediterranean meatballs while discussing the fundamental limits of long-context LLMs and traditional RAG. He outlines Engram's vision for modular parametric memory, test-time training, and personalized continual learning to unlock true AI efficiency and enterprise intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 3.3 Guest teaching 5.0 Guest disagreement 0.9 The hosts pushing back 0.8
05100:0015:0030:0045:000:13–7:10 · The hosts as informed peer 2/10 Welcome to Latent Space Cooking Show: Introducing Mediterranean Meatballs Allen introduces the cooking show format and explores Dan's transition from Israeli naval special operations and computational neuroscience into AI founding. The tone is casual and conversational while prepping ingredients.7:10–11:49 · The hosts as informed peer 3/10 The Genesis of Engram: Efficiency, Cartridges, and AI Intuition Dan explains Engram's origin in semi-supervised learning and introduces the concept of parametric cartridges. He illustrates parametric intuition versus in-context notes using the analogy of reading a cookbook versus having a chef's trained palate.11:49–15:54 · The hosts as informed peer 4/10 Parametric Intuition vs. Textual Notes and the Data Explosion Allen asks how parametric intuition differs fundamentally from retrieving relevant cookbook sections via RAG. Dan clarifies that weight-level learning is needed alongside notes, warning that enterprise knowledge bases will soon hit internet scale.15:56–18:00 · The hosts as informed peer 3/10 Seasoning Meatballs and Context Rot in Multi-Million Context Windows Allen questions whether the main issue with frontier models is the token expense of starting from scratch. Dan clarifies that context rot degrades model reasoning on underspecified agentic tasks regardless of window size.18:01–24:28 · The hosts as informed peer 4/10 Systems Bottlenecks: KV Cache Inefficiencies and Test-Time Training Allen steelmans the counterargument that enterprises do not need weight updates when RAG and context compaction exist. Dan demonstrates the systems failure of KV caching, noting an 80GB HBM state generated by a small Wikipedia article on Llama 70B.24:31–27:02 · The hosts as informed peer 3/10 Enterprise Use Cases: Holistic Reasoning Beyond Traditional RAG Dan provides concrete enterprise legal examples, showing why queries like incomplete M&A deals fail in standard RAG because the answer requires holistic synthesis across all client matters.27:03–29:58 · The hosts as informed peer 4/10 Continual Learning: Personalized Adapters and Enterprise Deployments Allen asks whether Engram uses parameter-efficient fine-tuning like LoRA for enterprise corpora. Dan outlines their vision of personal and enterprise weights acting like digital pets that continuously improve with user interaction.29:59–34:23 · The hosts as informed peer 4/10 Autonomous Memory Management: What Models Internalize vs. Externalize Allen asks how Engram determines what information lives in weights versus external context. Dan explains the goal of training models to autonomously decide what to internalize versus externalize, avoiding heuristic rules.34:23–38:04 · The hosts as informed peer 4/10 Token Efficiency, Model Routing, and Frontiers of Intelligence Allen asks about multi-agent routing trade-offs between cheap and expensive models. Dan reframes intelligence as achieving higher reasoning while expending less energy, acknowledging routing as an active open problem.38:05–42:22 · The hosts as informed peer 2/10 Inside Engram: Team Culture, Co-Founder Roles, and Startup Dynamics Allen asks about managing a heavily academic research team. Dan jokes about their lack of non-research diversity and explains their culture of pairing seasoned PhDs with rising talent to ship real products.42:22–45:16 · The hosts as informed peer 3/10 Scaling Infrastructure and Engram's Engineering Hiring Call Dan highlights the enormous systems challenge of swapping millions of adapter endpoints between disk and GPU HBM during live inference, issuing a hiring call for performance engineers.45:24–47:42 · The hosts as informed peer 3/10 Finishing the Dish: Coupling Efficiency with Frontier Intelligence Dan strongly rejects the notion that efficiency implies a budget commodity product, arguing efficiency is foundational for frontier intelligence. They taste and review the finished dish with Sean.0:13–7:10 · Guest teaching 1/10 Welcome to Latent Space Cooking Show: Introducing Mediterranean Meatballs Allen introduces the cooking show format and explores Dan's transition from Israeli naval special operations and computational neuroscience into AI founding. The tone is casual and conversational while prepping ingredients.7:10–11:49 · Guest teaching 5/10 The Genesis of Engram: Efficiency, Cartridges, and AI Intuition Dan explains Engram's origin in semi-supervised learning and introduces the concept of parametric cartridges. He illustrates parametric intuition versus in-context notes using the analogy of reading a cookbook versus having a chef's trained palate.11:49–15:54 · Guest teaching 6/10 Parametric Intuition vs. Textual Notes and the Data Explosion Allen asks how parametric intuition differs fundamentally from retrieving relevant cookbook sections via RAG. Dan clarifies that weight-level learning is needed alongside notes, warning that enterprise knowledge bases will soon hit internet scale.15:56–18:00 · Guest teaching 5/10 Seasoning Meatballs and Context Rot in Multi-Million Context Windows Allen questions whether the main issue with frontier models is the token expense of starting from scratch. Dan clarifies that context rot degrades model reasoning on underspecified agentic tasks regardless of window size.18:01–24:28 · Guest teaching 7/10 Systems Bottlenecks: KV Cache Inefficiencies and Test-Time Training Allen steelmans the counterargument that enterprises do not need weight updates when RAG and context compaction exist. Dan demonstrates the systems failure of KV caching, noting an 80GB HBM state generated by a small Wikipedia article on Llama 70B.24:31–27:02 · Guest teaching 6/10 Enterprise Use Cases: Holistic Reasoning Beyond Traditional RAG Dan provides concrete enterprise legal examples, showing why queries like incomplete M&A deals fail in standard RAG because the answer requires holistic synthesis across all client matters.27:03–29:58 · Guest teaching 5/10 Continual Learning: Personalized Adapters and Enterprise Deployments Allen asks whether Engram uses parameter-efficient fine-tuning like LoRA for enterprise corpora. Dan outlines their vision of personal and enterprise weights acting like digital pets that continuously improve with user interaction.29:59–34:23 · Guest teaching 6/10 Autonomous Memory Management: What Models Internalize vs. Externalize Allen asks how Engram determines what information lives in weights versus external context. Dan explains the goal of training models to autonomously decide what to internalize versus externalize, avoiding heuristic rules.34:23–38:04 · Guest teaching 6/10 Token Efficiency, Model Routing, and Frontiers of Intelligence Allen asks about multi-agent routing trade-offs between cheap and expensive models. Dan reframes intelligence as achieving higher reasoning while expending less energy, acknowledging routing as an active open problem.38:05–42:22 · Guest teaching 2/10 Inside Engram: Team Culture, Co-Founder Roles, and Startup Dynamics Allen asks about managing a heavily academic research team. Dan jokes about their lack of non-research diversity and explains their culture of pairing seasoned PhDs with rising talent to ship real products.42:22–45:16 · Guest teaching 6/10 Scaling Infrastructure and Engram's Engineering Hiring Call Dan highlights the enormous systems challenge of swapping millions of adapter endpoints between disk and GPU HBM during live inference, issuing a hiring call for performance engineers.45:24–47:42 · Guest teaching 5/10 Finishing the Dish: Coupling Efficiency with Frontier Intelligence Dan strongly rejects the notion that efficiency implies a budget commodity product, arguing efficiency is foundational for frontier intelligence. They taste and review the finished dish with Sean.0:13–7:10 · Guest disagreement 0/10 Welcome to Latent Space Cooking Show: Introducing Mediterranean Meatballs Allen introduces the cooking show format and explores Dan's transition from Israeli naval special operations and computational neuroscience into AI founding. The tone is casual and conversational while prepping ingredients.7:10–11:49 · Guest disagreement 1/10 The Genesis of Engram: Efficiency, Cartridges, and AI Intuition Dan explains Engram's origin in semi-supervised learning and introduces the concept of parametric cartridges. He illustrates parametric intuition versus in-context notes using the analogy of reading a cookbook versus having a chef's trained palate.11:49–15:54 · Guest disagreement 1/10 Parametric Intuition vs. Textual Notes and the Data Explosion Allen asks how parametric intuition differs fundamentally from retrieving relevant cookbook sections via RAG. Dan clarifies that weight-level learning is needed alongside notes, warning that enterprise knowledge bases will soon hit internet scale.15:56–18:00 · Guest disagreement 1/10 Seasoning Meatballs and Context Rot in Multi-Million Context Windows Allen questions whether the main issue with frontier models is the token expense of starting from scratch. Dan clarifies that context rot degrades model reasoning on underspecified agentic tasks regardless of window size.18:01–24:28 · Guest disagreement 2/10 Systems Bottlenecks: KV Cache Inefficiencies and Test-Time Training Allen steelmans the counterargument that enterprises do not need weight updates when RAG and context compaction exist. Dan demonstrates the systems failure of KV caching, noting an 80GB HBM state generated by a small Wikipedia article on Llama 70B.24:31–27:02 · Guest disagreement 1/10 Enterprise Use Cases: Holistic Reasoning Beyond Traditional RAG Dan provides concrete enterprise legal examples, showing why queries like incomplete M&A deals fail in standard RAG because the answer requires holistic synthesis across all client matters.27:03–29:58 · Guest disagreement 0/10 Continual Learning: Personalized Adapters and Enterprise Deployments Allen asks whether Engram uses parameter-efficient fine-tuning like LoRA for enterprise corpora. Dan outlines their vision of personal and enterprise weights acting like digital pets that continuously improve with user interaction.29:59–34:23 · Guest disagreement 1/10 Autonomous Memory Management: What Models Internalize vs. Externalize Allen asks how Engram determines what information lives in weights versus external context. Dan explains the goal of training models to autonomously decide what to internalize versus externalize, avoiding heuristic rules.34:23–38:04 · Guest disagreement 1/10 Token Efficiency, Model Routing, and Frontiers of Intelligence Allen asks about multi-agent routing trade-offs between cheap and expensive models. Dan reframes intelligence as achieving higher reasoning while expending less energy, acknowledging routing as an active open problem.38:05–42:22 · Guest disagreement 1/10 Inside Engram: Team Culture, Co-Founder Roles, and Startup Dynamics Allen asks about managing a heavily academic research team. Dan jokes about their lack of non-research diversity and explains their culture of pairing seasoned PhDs with rising talent to ship real products.42:22–45:16 · Guest disagreement 0/10 Scaling Infrastructure and Engram's Engineering Hiring Call Dan highlights the enormous systems challenge of swapping millions of adapter endpoints between disk and GPU HBM during live inference, issuing a hiring call for performance engineers.45:24–47:42 · Guest disagreement 2/10 Finishing the Dish: Coupling Efficiency with Frontier Intelligence Dan strongly rejects the notion that efficiency implies a budget commodity product, arguing efficiency is foundational for frontier intelligence. They taste and review the finished dish with Sean.0:13–7:10 · The hosts pushing back 0/10 Welcome to Latent Space Cooking Show: Introducing Mediterranean Meatballs Allen introduces the cooking show format and explores Dan's transition from Israeli naval special operations and computational neuroscience into AI founding. The tone is casual and conversational while prepping ingredients.7:10–11:49 · The hosts pushing back 1/10 The Genesis of Engram: Efficiency, Cartridges, and AI Intuition Dan explains Engram's origin in semi-supervised learning and introduces the concept of parametric cartridges. He illustrates parametric intuition versus in-context notes using the analogy of reading a cookbook versus having a chef's trained palate.11:49–15:54 · The hosts pushing back 1/10 Parametric Intuition vs. Textual Notes and the Data Explosion Allen asks how parametric intuition differs fundamentally from retrieving relevant cookbook sections via RAG. Dan clarifies that weight-level learning is needed alongside notes, warning that enterprise knowledge bases will soon hit internet scale.15:56–18:00 · The hosts pushing back 1/10 Seasoning Meatballs and Context Rot in Multi-Million Context Windows Allen questions whether the main issue with frontier models is the token expense of starting from scratch. Dan clarifies that context rot degrades model reasoning on underspecified agentic tasks regardless of window size.18:01–24:28 · The hosts pushing back 3/10 Systems Bottlenecks: KV Cache Inefficiencies and Test-Time Training Allen steelmans the counterargument that enterprises do not need weight updates when RAG and context compaction exist. Dan demonstrates the systems failure of KV caching, noting an 80GB HBM state generated by a small Wikipedia article on Llama 70B.24:31–27:02 · The hosts pushing back 1/10 Enterprise Use Cases: Holistic Reasoning Beyond Traditional RAG Dan provides concrete enterprise legal examples, showing why queries like incomplete M&A deals fail in standard RAG because the answer requires holistic synthesis across all client matters.27:03–29:58 · The hosts pushing back 0/10 Continual Learning: Personalized Adapters and Enterprise Deployments Allen asks whether Engram uses parameter-efficient fine-tuning like LoRA for enterprise corpora. Dan outlines their vision of personal and enterprise weights acting like digital pets that continuously improve with user interaction.29:59–34:23 · The hosts pushing back 1/10 Autonomous Memory Management: What Models Internalize vs. Externalize Allen asks how Engram determines what information lives in weights versus external context. Dan explains the goal of training models to autonomously decide what to internalize versus externalize, avoiding heuristic rules.34:23–38:04 · The hosts pushing back 1/10 Token Efficiency, Model Routing, and Frontiers of Intelligence Allen asks about multi-agent routing trade-offs between cheap and expensive models. Dan reframes intelligence as achieving higher reasoning while expending less energy, acknowledging routing as an active open problem.38:05–42:22 · The hosts pushing back 0/10 Inside Engram: Team Culture, Co-Founder Roles, and Startup Dynamics Allen asks about managing a heavily academic research team. Dan jokes about their lack of non-research diversity and explains their culture of pairing seasoned PhDs with rising talent to ship real products.42:22–45:16 · The hosts pushing back 0/10 Scaling Infrastructure and Engram's Engineering Hiring Call Dan highlights the enormous systems challenge of swapping millions of adapter endpoints between disk and GPU HBM during live inference, issuing a hiring call for performance engineers.45:24–47:42 · The hosts pushing back 0/10 Finishing the Dish: Coupling Efficiency with Frontier Intelligence Dan strongly rejects the notion that efficiency implies a budget commodity product, arguing efficiency is foundational for frontier intelligence. They taste and review the finished dish with Sean.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 45:43 Dan rejects the efficiency versus quality dichotomy

Dan forcefully rejects the conventional industry framing that building efficiency tools relegates a company to a budget tier rather than frontier intelligence.

Hardest push from the hosts ▶ 18:22 Allen steelmans the counterargument against in-weight learning

Allen challenges Dan's core value proposition by asking why enterprises cannot simply rely on context compaction, cheaper open-source models, and traditional RAG.

Biggest teaching moment ▶ 22:10 Dan reveals the KV cache memory explosion

Dan demonstrates the extreme systems inefficiency of in-context prefill by calculating that a small Wikipedia article creates an 80GB GPU memory footprint on Llama 70B.

The host holds their own ▶ 11:49 Allen presses on how parametric memory beats RAG chunking

Allen probes Dan's chef metaphor by asking why extracting relevant notes through RAG would not provide identical understanding in-context.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Welcome to Latent Space Cooking Show: Introducing Mediterranean Meatballs 2100 Allen introduces the cooking show format and explores Dan's transition from Israeli naval special operations and computational neuroscience into AI founding. The tone is casual and conversational while prepping ingredients.
The Genesis of Engram: Efficiency, Cartridges, and AI Intuition 3511 Dan explains Engram's origin in semi-supervised learning and introduces the concept of parametric cartridges. He illustrates parametric intuition versus in-context notes using the analogy of reading a cookbook versus having a chef's trained palate.
Parametric Intuition vs. Textual Notes and the Data Explosion 4611 Allen asks how parametric intuition differs fundamentally from retrieving relevant cookbook sections via RAG. Dan clarifies that weight-level learning is needed alongside notes, warning that enterprise knowledge bases will soon hit internet scale.
Seasoning Meatballs and Context Rot in Multi-Million Context Windows 3511 Allen questions whether the main issue with frontier models is the token expense of starting from scratch. Dan clarifies that context rot degrades model reasoning on underspecified agentic tasks regardless of window size.
Systems Bottlenecks: KV Cache Inefficiencies and Test-Time Training 4723 Allen steelmans the counterargument that enterprises do not need weight updates when RAG and context compaction exist. Dan demonstrates the systems failure of KV caching, noting an 80GB HBM state generated by a small Wikipedia article on Llama 70B.
Enterprise Use Cases: Holistic Reasoning Beyond Traditional RAG 3611 Dan provides concrete enterprise legal examples, showing why queries like incomplete M&A deals fail in standard RAG because the answer requires holistic synthesis across all client matters.
Continual Learning: Personalized Adapters and Enterprise Deployments 4500 Allen asks whether Engram uses parameter-efficient fine-tuning like LoRA for enterprise corpora. Dan outlines their vision of personal and enterprise weights acting like digital pets that continuously improve with user interaction.
Autonomous Memory Management: What Models Internalize vs. Externalize 4611 Allen asks how Engram determines what information lives in weights versus external context. Dan explains the goal of training models to autonomously decide what to internalize versus externalize, avoiding heuristic rules.
Token Efficiency, Model Routing, and Frontiers of Intelligence 4611 Allen asks about multi-agent routing trade-offs between cheap and expensive models. Dan reframes intelligence as achieving higher reasoning while expending less energy, acknowledging routing as an active open problem.
Inside Engram: Team Culture, Co-Founder Roles, and Startup Dynamics 2210 Allen asks about managing a heavily academic research team. Dan jokes about their lack of non-research diversity and explains their culture of pairing seasoned PhDs with rising talent to ship real products.
Scaling Infrastructure and Engram's Engineering Hiring Call 3600 Dan highlights the enormous systems challenge of swapping millions of adapter endpoints between disk and GPU HBM during live inference, issuing a hiring call for performance engineers.
Finishing the Dish: Coupling Efficiency with Frontier Intelligence 3520 Dan strongly rejects the notion that efficiency implies a budget commodity product, arguing efficiency is foundational for frontier intelligence. They taste and review the finished dish with Sean.

Statements from this episode (18)

Insight
Biderman: Israeli culture gives people multiple sequential shots on goal
“Israel as a culture is a place where you can basically, you get multiple shots at goal. If you're not the best in high school you still might have a good position in the military. And if you're really good, more doors open up for you for university. And even i…”
Dan Biderman Jul 13, 2026 ▶ 5:58
Prediction Not checkable as stated
Biderman: Semi-supervised learning will become super crucial again
“My PhD was focusing on On, on a field that's not super in vogue today, but I think will become super crucial again, which is semi supervised learning”
Dan Biderman Jul 13, 2026 ▶ 7:42
Insight
Biderman: Current LLMs lack chef-like intuition and rely on robotic reading
“Current LLMs are like coming into the kitchen first time, every time, reading the textbook, cooking the dish, measuring everything, but they don't have the intuition of a chef that's pinching salt and kneading dough and things like this.”
Dan Biderman Jul 13, 2026 ▶ 11:12
Prediction Not checkable as stated
Biderman: AI-native companies will amass trillions of internal tokens within 18 months
“In 18 months, many companies would have maybe trillions of tokens, which of internal company data, proprietary data. I'm talking about like maybe trillions. It sounds exaggerated, but I don't think it's an impossibility if they're really AI native.”
Dan Biderman Jul 13, 2026 ▶ 14:42
Prediction Not checkable as stated
Biderman: Model accuracy will still degrade at 10M context window scale
“But two is like, for the agentic tasks of 18 months from now, inside those major repositories of knowledge, and asking the models more and more things in underspecified ways, I suspect that the accuracy of the models would go down. The phenomenon of context fr…”
Dan Biderman Jul 13, 2026 ▶ 17:31
Prediction Not checkable as stated
Biderman: Hard engineering tasks will require test-time gradient updates
“We think that eventually part of the solution for very hard tasks in, in science and engineering and defense and all that stuff will involve some form of gradient based updates during during doing these long horizon tasks.”
Dan Biderman Jul 13, 2026 ▶ 21:57
Assertion Partly supported
Biderman: Processing a Wikipedia article in Llama 70B consumes 80GB HBM
“If you take a Lama, a 70 B model, and you load one article from Wikipedia, which is a few tens of kilobytes, and you have the model read this The brain state of the model when reading this few tens of kilobytes is like, 80 gigabytes. 80 gigabytes on, on the HB…”
Dan Biderman Jul 13, 2026 ▶ 22:39
Disclosure
Biderman: Engram works with Harvey on large file systems
“And so, for example, these are the kinds of things we work with Harvey.”
Dan Biderman Jul 13, 2026 ▶ 25:11
Assertion Not checkable as stated
Biderman: Harmless enterprise queries on frontier models cost thousands of dollars
“And now you can solve these tasks with frontier models and compaction. And when you ask them to do so, they will consume thousands of dollars for queries that we think are harmless. That every employee in the company would be able to answer.”
Dan Biderman Jul 13, 2026 ▶ 25:50
Prediction Not checkable as stated
Biderman: In 18 months, data scale will require weight-based learning
“Other parts of it are bets that in 18 months from now, the scale of the data will require the methods that we know from pre-training work.”
Dan Biderman Jul 13, 2026 ▶ 26:55
Disclosure
Biderman: Engram aims to give every user personalized, continually-learning weights
“Our ambition in the long term, ah, is that. Every person has a model, or a part of the model, or a set of weights that, that represents their knowledge, their expertise, learns from them, that the more time they spend with the model, the better it gets for the…”
Dan Biderman Jul 13, 2026 ▶ 27:13
Prediction Open · timeframe Jul 2031
Biderman: PC hardware will soon run near-trillion-parameter models locally
“And in the long, long term, I do think these things will actually run on people's devices, and we're seeing right now the new hardware on personal computers is already, ah, you know, soon approaching the ability to run inference on close to trillion parameters…”
Dan Biderman Jul 13, 2026 ▶ 28:12
Disclosure
Biderman: Engram trains models to decide what to memorize vs keep in notes
“The way to work on it is to train models both, to train models to manage it themselves, and that's an active area for us. Have the model know, like, without any explicit supervision signal to determine this kind of stuff I can pull from my brain, and that kind…”
Dan Biderman Jul 13, 2026 ▶ 30:56
Insight
Biderman: Manually partitioning LLM memory vs retrieval becomes unmanageable whack-a-mole
“And now the thing is, if you start manually, heuristically saying this is in, this is out, then it becomes a whack-a-mole. Every, every person in every enterprise has different data, and you can really very easily pick and choose what goes in and what goes out…”
Dan Biderman Jul 13, 2026 ▶ 31:47
Prediction Not checkable as stated
Biderman: AI models must learn to autonomously filter out erroneous user feedback
“Increasingly the models will get better, and increasingly they'll know more things than we do, so the model in some way has to learn and understand and kind of, like, discern what, which feedback is valuable and which feedback should be ignored.”
Dan Biderman Jul 13, 2026 ▶ 33:54
Prediction Not checkable as stated
Biderman: AI solutions will rely on model routing, not single monolithic models
“So I think routing will be part of the solution there for sure, and I think Many people, not just myself, say this solution is multi-modal. It's not Engram taking over. There's one model, and you teach it things, and you can close Stargate. That's not our appr…”
Dan Biderman Jul 13, 2026 ▶ 36:56
Insight
Biderman: Personalized AI adapters require hot-swapping millions of endpoints at inference
“And if you truly believe that we can get to the level where we have those kinds of parameter efficient adapters for every person and team, you suddenly think about deployments that involve millions of different endpoints stored in different places that need to…”
Dan Biderman Jul 13, 2026 ▶ 43:43
Insight
Biderman: AI efficiency and frontier intelligence cannot be decoupled
“The point for me is, the principle is, any kind of, like, efficiency and intelligence, they cannot really be decoupled. Sometimes people think if you're building something that's more efficient, that can save you dollars, therefore you're not in the premium ca…”
Dan Biderman Jul 13, 2026 ▶ 45:46
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.