Insight certainty 4/5 debate potential 2/5

Shah: True LLM User Understanding Requires Four Memory Capabilities

Dhravya Shah · ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory · Mar 9, 2026 · at 3:58

Dhravya Shah, founder of Supermemory, explains why naive vector storage is insufficient for LLM memory and outlines the four core architectural pillars required.

0:00 / 0:39exact quote · 39.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So, to make an LLM truly good at understanding a user, you have to handle four things. One is knowledge updates, so you have to invalidate, steal knowledge, and build on top of it, which is different from the storing vectors. You have to have some sort of temporal reasoning, so you have to have, like, give the agent an understanding of time and how things have passed for personalization, and you need some sort of forgetfulness to forget things that are not relevant anymore at all. But you also need this concept we call user profiles to always, like, memory is not just a retrieval call. The LLM has this very small profile of the user that it will utilize on every single turn.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Dhravya Shah

Opinion
Shah: Claude Code avoids indexing because vendors profit from high token usage
“They could index the code and cursor does, but plot code does not. And I believe this is also because they are not incentivized to utilize less tokens by the end user, because they're just offsetting the cost of indexing, et cetera, to tokens and time that the…”
Dhravya Shah Mar 9, 2026 ▶ 13:36 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Insight
Shah: OpenClaw memory fails because LLMs frequently skip search tool calls
“The way OpenClaw has with QMD or with whatever memory plugin you use, it inherently relies on tools to search through these memory.md files that it prepares. So, you know, like, what did I decide about the API? Then agent will decide to search, and sometimes i…”
Dhravya Shah Mar 9, 2026 ▶ 8:37 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Assertion Supported
Shah: OpenClaw's 15-message replay fails prompt caching and costs 10x more
“The way OpenClaw does it is it essentially sends back the last 15 messages in the conversation and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like, you are not doing any, like, you're not utilizing a…”
Dhravya Shah Mar 9, 2026 ▶ 21:03 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Opinion
Shah: Most vector databases become too slow or expensive at scale
“So, so, better databases, like, at that time, and even now, except a few, like, TurboPuffer and Chroma, Were not really scalable for the scale I was getting to where they were either getting too slow or too expensive to run.”
Dhravya Shah Mar 9, 2026 ▶ 2:25 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Opinion
Shah: Triplet-Based Knowledge Graphs Degrade AI Memory Performance
“And we think that triplets actually Lead to worse performance because you have to traverse them a lot to get to any information.”
Dhravya Shah Mar 9, 2026 ▶ 5:15 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Assertion Supported
Shah: Supermemory outperformed Claude Code and OpenClaw benchmarks by almost 50%
“So the Claude code one performed the worst, and OpenClaw slightly more than that, and SuperMemory is the highest, and you can see that, you know, it's like a pretty significant difference, like almost 50%.”
Dhravya Shah Mar 9, 2026 ▶ 14:33 ⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.