Mar 9, 2026 · 27m · latent-space

⚡️ OpenClaw's Memory Sucks and the fix is simple — Dhravya Shah, Supermemory

Dhravya Shah · 19m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Supermemory founder Dhravya Shah explains why naive RAG and file-based agent memory fail in modern AI systems, presenting dynamic knowledge graphs, deterministic context hooks, and open benchmarks as the foundation for scalable context infrastructure.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.1 Guest teaching 4.1 Guest disagreement 2.4 The hosts pushing back 2.8
05100:0010:0020:000:03–3:16 · The hosts as informed peer 4/10 From Buildspace Side Project to Viral Open Source Architecture Host kicks off the interview with familiar rapport, joking about Buildspace's founder and noting he starred the Supermemory repository when it first went viral.3:17–6:03 · The hosts as informed peer 6/10 Re-Architecting AI Memory Beyond Naive RAG and Triplets Guest outlines the evolution from naive vector RAG to knowledge graphs, while the host demonstrates domain familiarity by instantly completing definitions around triplet entities and graph traversal trade-offs.6:04–8:08 · The hosts as informed peer 3/10 Supermemory Today: Context Infrastructure and Company Profile Standard founder profile segment where Dhravya explains Supermemory's VC funding, transition from consumer tool to context infrastructure API, and plugin traction.8:08–12:19 · The hosts as informed peer 6/10 Diagnosing OpenClaw's Memory Flaws and the Hooks Solution Guest diagnoses why tool-based search fails in OpenClaw compared to automated hooks. Host contributes sharp technical context by contrasting Toby Lutke's local QMD philosophy with Cloudflare cloud architectures and tool-calling failure modes in models like Grok.12:19–16:21 · The hosts as informed peer 6/10 File System Memory, Token Incentives, and MemoryBench Guest presents a hot take that AI coding assistants lack token reduction incentives and presents a benchmark favoring Supermemory. The host pushes back directly, forcing the guest to admit the benchmark tested their own reproduction rather than native tools.16:22–20:38 · The hosts as informed peer 5/10 Critiquing AI Memory Benchmarks and the User Profile Dilemma Guest deconstructs popular AI memory benchmarks, strongly criticizing Locomo and explaining why extracting everything on LongMemEval creates unrealistic production costs. When host doubts this flaw, guest educates him on extraction pricing and episodic user profiling.20:39–24:55 · The hosts as informed peer 6/10 OpenClaw Edge Cases, Token Economics, and Hybrid Retrieval Host presses on whether real developers actually care about memory token costs. Guest pushes back with real-world customer economic realities and introduces Supermemory's hybrid RAG fallback architecture.24:56–26:56 · The hosts as informed peer 5/10 Supermemory Future Roadmap, Voice Agents, and Industry Outlook Wrap-up segment covering voice agents and future roadmaps. Host lists several competing memory architectures and invites the guest to deliver an upcoming technical workshop.0:03–3:16 · Guest teaching 2/10 From Buildspace Side Project to Viral Open Source Architecture Host kicks off the interview with familiar rapport, joking about Buildspace's founder and noting he starred the Supermemory repository when it first went viral.3:17–6:03 · Guest teaching 5/10 Re-Architecting AI Memory Beyond Naive RAG and Triplets Guest outlines the evolution from naive vector RAG to knowledge graphs, while the host demonstrates domain familiarity by instantly completing definitions around triplet entities and graph traversal trade-offs.6:04–8:08 · Guest teaching 2/10 Supermemory Today: Context Infrastructure and Company Profile Standard founder profile segment where Dhravya explains Supermemory's VC funding, transition from consumer tool to context infrastructure API, and plugin traction.8:08–12:19 · Guest teaching 5/10 Diagnosing OpenClaw's Memory Flaws and the Hooks Solution Guest diagnoses why tool-based search fails in OpenClaw compared to automated hooks. Host contributes sharp technical context by contrasting Toby Lutke's local QMD philosophy with Cloudflare cloud architectures and tool-calling failure modes in models like Grok.12:19–16:21 · Guest teaching 5/10 File System Memory, Token Incentives, and MemoryBench Guest presents a hot take that AI coding assistants lack token reduction incentives and presents a benchmark favoring Supermemory. The host pushes back directly, forcing the guest to admit the benchmark tested their own reproduction rather than native tools.16:22–20:38 · Guest teaching 7/10 Critiquing AI Memory Benchmarks and the User Profile Dilemma Guest deconstructs popular AI memory benchmarks, strongly criticizing Locomo and explaining why extracting everything on LongMemEval creates unrealistic production costs. When host doubts this flaw, guest educates him on extraction pricing and episodic user profiling.20:39–24:55 · Guest teaching 5/10 OpenClaw Edge Cases, Token Economics, and Hybrid Retrieval Host presses on whether real developers actually care about memory token costs. Guest pushes back with real-world customer economic realities and introduces Supermemory's hybrid RAG fallback architecture.24:56–26:56 · Guest teaching 2/10 Supermemory Future Roadmap, Voice Agents, and Industry Outlook Wrap-up segment covering voice agents and future roadmaps. Host lists several competing memory architectures and invites the guest to deliver an upcoming technical workshop.0:03–3:16 · Guest disagreement 1/10 From Buildspace Side Project to Viral Open Source Architecture Host kicks off the interview with familiar rapport, joking about Buildspace's founder and noting he starred the Supermemory repository when it first went viral.3:17–6:03 · Guest disagreement 2/10 Re-Architecting AI Memory Beyond Naive RAG and Triplets Guest outlines the evolution from naive vector RAG to knowledge graphs, while the host demonstrates domain familiarity by instantly completing definitions around triplet entities and graph traversal trade-offs.6:04–8:08 · Guest disagreement 1/10 Supermemory Today: Context Infrastructure and Company Profile Standard founder profile segment where Dhravya explains Supermemory's VC funding, transition from consumer tool to context infrastructure API, and plugin traction.8:08–12:19 · Guest disagreement 3/10 Diagnosing OpenClaw's Memory Flaws and the Hooks Solution Guest diagnoses why tool-based search fails in OpenClaw compared to automated hooks. Host contributes sharp technical context by contrasting Toby Lutke's local QMD philosophy with Cloudflare cloud architectures and tool-calling failure modes in models like Grok.12:19–16:21 · Guest disagreement 4/10 File System Memory, Token Incentives, and MemoryBench Guest presents a hot take that AI coding assistants lack token reduction incentives and presents a benchmark favoring Supermemory. The host pushes back directly, forcing the guest to admit the benchmark tested their own reproduction rather than native tools.16:22–20:38 · Guest disagreement 4/10 Critiquing AI Memory Benchmarks and the User Profile Dilemma Guest deconstructs popular AI memory benchmarks, strongly criticizing Locomo and explaining why extracting everything on LongMemEval creates unrealistic production costs. When host doubts this flaw, guest educates him on extraction pricing and episodic user profiling.20:39–24:55 · Guest disagreement 3/10 OpenClaw Edge Cases, Token Economics, and Hybrid Retrieval Host presses on whether real developers actually care about memory token costs. Guest pushes back with real-world customer economic realities and introduces Supermemory's hybrid RAG fallback architecture.24:56–26:56 · Guest disagreement 1/10 Supermemory Future Roadmap, Voice Agents, and Industry Outlook Wrap-up segment covering voice agents and future roadmaps. Host lists several competing memory architectures and invites the guest to deliver an upcoming technical workshop.0:03–3:16 · The hosts pushing back 1/10 From Buildspace Side Project to Viral Open Source Architecture Host kicks off the interview with familiar rapport, joking about Buildspace's founder and noting he starred the Supermemory repository when it first went viral.3:17–6:03 · The hosts pushing back 1/10 Re-Architecting AI Memory Beyond Naive RAG and Triplets Guest outlines the evolution from naive vector RAG to knowledge graphs, while the host demonstrates domain familiarity by instantly completing definitions around triplet entities and graph traversal trade-offs.6:04–8:08 · The hosts pushing back 1/10 Supermemory Today: Context Infrastructure and Company Profile Standard founder profile segment where Dhravya explains Supermemory's VC funding, transition from consumer tool to context infrastructure API, and plugin traction.8:08–12:19 · The hosts pushing back 3/10 Diagnosing OpenClaw's Memory Flaws and the Hooks Solution Guest diagnoses why tool-based search fails in OpenClaw compared to automated hooks. Host contributes sharp technical context by contrasting Toby Lutke's local QMD philosophy with Cloudflare cloud architectures and tool-calling failure modes in models like Grok.12:19–16:21 · The hosts pushing back 6/10 File System Memory, Token Incentives, and MemoryBench Guest presents a hot take that AI coding assistants lack token reduction incentives and presents a benchmark favoring Supermemory. The host pushes back directly, forcing the guest to admit the benchmark tested their own reproduction rather than native tools.16:22–20:38 · The hosts pushing back 4/10 Critiquing AI Memory Benchmarks and the User Profile Dilemma Guest deconstructs popular AI memory benchmarks, strongly criticizing Locomo and explaining why extracting everything on LongMemEval creates unrealistic production costs. When host doubts this flaw, guest educates him on extraction pricing and episodic user profiling.20:39–24:55 · The hosts pushing back 5/10 OpenClaw Edge Cases, Token Economics, and Hybrid Retrieval Host presses on whether real developers actually care about memory token costs. Guest pushes back with real-world customer economic realities and introduces Supermemory's hybrid RAG fallback architecture.24:56–26:56 · The hosts pushing back 1/10 Supermemory Future Roadmap, Voice Agents, and Industry Outlook Wrap-up segment covering voice agents and future roadmaps. Host lists several competing memory architectures and invites the guest to deliver an upcoming technical workshop.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 0%27:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 17:02 Locomo benchmark takedown

Dhravya aggressively dismisses the Locomo benchmark as outdated and flawed, arguing you can easily game 100% scores simply by dumping raw context into newer models.

Hardest push from the hosts ▶ 14:47 Benchmark reproduction challenge

Alessio interrupts the guest's celebratory benchmark results to insist on methodological clarity, forcing Dhravya to concede that the benchmark evaluated Supermemory's custom re-implementation rather than native Claude code.

Biggest teaching moment ▶ 16:43 Explaining extraction cost fallacies

When Alessio suggests that maximizing memory extraction to win benchmarks does not seem problematic, Dhravya breaks down the prohibitive real-world token extraction costs and lack of forgetfulness testing.

The host holds their own ▶ 11:11 Local vs cloud memory trade-off critique

Alessio demonstrates technical depth by contrasting Toby Lutke's local-first QMD approach with Cloudflare architectures, illustrating the real-world flaw of relying on models to voluntarily trigger search tools.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
From Buildspace Side Project to Viral Open Source Architecture 4211 Host kicks off the interview with familiar rapport, joking about Buildspace's founder and noting he starred the Supermemory repository when it first went viral.
Re-Architecting AI Memory Beyond Naive RAG and Triplets 6521 Guest outlines the evolution from naive vector RAG to knowledge graphs, while the host demonstrates domain familiarity by instantly completing definitions around triplet entities and graph traversal trade-offs.
Supermemory Today: Context Infrastructure and Company Profile 3211 Standard founder profile segment where Dhravya explains Supermemory's VC funding, transition from consumer tool to context infrastructure API, and plugin traction.
Diagnosing OpenClaw's Memory Flaws and the Hooks Solution 6533 Guest diagnoses why tool-based search fails in OpenClaw compared to automated hooks. Host contributes sharp technical context by contrasting Toby Lutke's local QMD philosophy with Cloudflare cloud architectures and tool-calling failure modes in models like Grok.
File System Memory, Token Incentives, and MemoryBench 6546 Guest presents a hot take that AI coding assistants lack token reduction incentives and presents a benchmark favoring Supermemory. The host pushes back directly, forcing the guest to admit the benchmark tested their own reproduction rather than native tools.
Critiquing AI Memory Benchmarks and the User Profile Dilemma 5744 Guest deconstructs popular AI memory benchmarks, strongly criticizing Locomo and explaining why extracting everything on LongMemEval creates unrealistic production costs. When host doubts this flaw, guest educates him on extraction pricing and episodic user profiling.
OpenClaw Edge Cases, Token Economics, and Hybrid Retrieval 6535 Host presses on whether real developers actually care about memory token costs. Guest pushes back with real-world customer economic realities and introduces Supermemory's hybrid RAG fallback architecture.
Supermemory Future Roadmap, Voice Agents, and Industry Outlook 5211 Wrap-up segment covering voice agents and future roadmaps. Host lists several competing memory architectures and invites the guest to deliver an upcoming technical workshop.

Statements from this episode (15)

Assertion Not checkable as stated
Shah: Supermemory was among 2024's fastest-growing open-source projects
“I open sourced SuperMemory during BuildSpace and it Like the project absolutely blew up. It got like, you know, it was one of the biggest projects, like fastest growing projects of 20, 24.”
Dhravya Shah Mar 9, 2026 ▶ 1:50
Opinion
Shah: Most vector databases become too slow or expensive at scale
“So, so, better databases, like, at that time, and even now, except a few, like, TurboPuffer and Chroma, Were not really scalable for the scale I was getting to where they were either getting too slow or too expensive to run.”
Dhravya Shah Mar 9, 2026 ▶ 2:25
Insight
Shah: True LLM User Understanding Requires Four Memory Capabilities
“So, to make an LLM truly good at understanding a user, you have to handle four things. One is knowledge updates, so you have to invalidate, steal knowledge, and build on top of it, which is different from the storing vectors. You have to have some sort of temp…”
Dhravya Shah Mar 9, 2026 ▶ 3:58
Opinion
Shah: Triplet-Based Knowledge Graphs Degrade AI Memory Performance
“And we think that triplets actually Lead to worse performance because you have to traverse them a lot to get to any information.”
Dhravya Shah Mar 9, 2026 ▶ 5:15
Disclosure
Shah: Supermemory Pivoted from Consumer App to Memory Infrastructure
“And I was like, this consumer app is probably not useful anymore. This core infrastructure is more useful now. And that is our core business now. And that's what we offer.”
Dhravya Shah Mar 9, 2026 ▶ 5:53
Insight
Shah: OpenClaw memory fails because LLMs frequently skip search tool calls
“The way OpenClaw has with QMD or with whatever memory plugin you use, it inherently relies on tools to search through these memory.md files that it prepares. So, you know, like, what did I decide about the API? Then agent will decide to search, and sometimes i…”
Dhravya Shah Mar 9, 2026 ▶ 8:37
Assertion Supported
Shah: Supermemory dynamically injects under 2,000 tokens via deterministic hooks
“So we basically turned the tools based approach to hooks based approach that actually is dealing from the super memory graph, which keeps the content fresh. So it handles updates, et cetera. So now, you know, for every tool call or for every user message, a ho…”
Dhravya Shah Mar 9, 2026 ▶ 10:13
Opinion
Shah: Claude Code avoids indexing because vendors profit from high token usage
“They could index the code and cursor does, but plot code does not. And I believe this is also because they are not incentivized to utilize less tokens by the end user, because they're just offsetting the cost of indexing, et cetera, to tokens and time that the…”
Dhravya Shah Mar 9, 2026 ▶ 13:36
Assertion Supported
Shah: Supermemory outperformed Claude Code and OpenClaw benchmarks by almost 50%
“So the Claude code one performed the worst, and OpenClaw slightly more than that, and SuperMemory is the highest, and you can see that, you know, it's like a pretty significant difference, like almost 50%.”
Dhravya Shah Mar 9, 2026 ▶ 14:33
Opinion
Shah: LongMemEval is good but reward-hacks memory extraction volume
“LongMemival is a very good benchmark because it calculates, it tests for the right things. It tests for the right things, the data set is great, everything is really good in LongMemival, except for the fact that you can win LongMemival by remembering as many t…”
Dhravya Shah Mar 9, 2026 ▶ 16:36
Opinion
Shah: LoCoMo benchmark is flawed because models pass via context dumping
“Locomo is, like, really bad. And people actually, it's not just me, like, pretty much everyone knows that Locomo is not the right thing to judge. So the reason is, like, Locomo is essentially testing for retrieval capacity. And it's not testing for the real me…”
Dhravya Shah Mar 9, 2026 ▶ 17:38
Insight
Shah: Pure retrieval AI memory fails on non-literal questions
“Traditionally you have this retrieval thing that happens. So you get a question, you retrieve something, and an answer is generated based on that. But we think that it will not work for most non-literal questions, like find the best monitor for me. If I've nev…”
Dhravya Shah Mar 9, 2026 ▶ 19:24
Assertion Supported
Shah: OpenClaw's 15-message replay fails prompt caching and costs 10x more
“The way OpenClaw does it is it essentially sends back the last 15 messages in the conversation and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like, you are not doing any, like, you're not utilizing a…”
Dhravya Shah Mar 9, 2026 ▶ 21:03
Assertion Supported
Shah: No other memory provider offers hybrid memory and raw RAG fallback
“Essentially, you will return the memories first, and then if there's any raw chunks that match up, we also return those to make sure the agent knows just enough information to answer the question. And no other provider does this right now”
Dhravya Shah Mar 9, 2026 ▶ 23:28
Disclosure
Shah: Supermemory infrastructure costs only two cents per million tokens
“For us, we have optimized infrastructure down to, it costs us two cents for a million tokens, because we have our own model, our own database, our own everything.”
Dhravya Shah Mar 9, 2026 ▶ 24:43
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.