Apr 21, 2025 · 34m · latent-space

Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)

Charles Packer · 21m spoken Shawn Wang · 4m spoken Charlie Snell · 2m spoken Alessio Fanelli · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Latent Space podcast episode, Letta AI researchers Charles Packer, Kevin Lin, and Charlie Snell introduce 'Sleep-Time Compute,' a post-training paradigm that utilizes idle GPU time to pre-compute reasoning and optimize stateful agent performance. They present empirical benchmarks across frontier models and outline practical architectures like MemGPT v2 to transform reactive LLM interfaces into proactive, persistent systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.8% of the talking time here. How this is scored →

The hosts as informed peer 5.1 Guest teaching 4.7 Guest disagreement 1.3 The hosts pushing back 2.1
05100:0010:0020:0030:000:00–4:34 · The hosts as informed peer 5/10 Introducing Sleep-Time Compute and Guest Research Backgrounds Swyx introduces the guests, noting their prior research background on MemGPT and test-time compute. Swyx inquires about the strategic naming shift from LLM OS to stateful agents, prompting Charles to clarify how the systems infrastructure enables the end product.4:35–8:38 · The hosts as informed peer 4/10 Core Motivation: Stateful Context and Pre-Computing Intermediate Inferences Charles explains the core concept of sleep-time compute through pre-computing intermediate mathematical inferences and codebase exploration. Alessio asks about the origins of the biological brain analogy, which Charles clarifies by connecting cognitive memory consolidation to database background indexing.8:39–10:59 · The hosts as informed peer 4/10 Defining Sleep-Time Compute: Cognitive versus Systems-Level Analogies Alessio presses on how sleep time is technically demarcated in an LLM lifecycle. Charles educates the hosts by defining sleep time as any post-training time that is not test time, arguing that systems-level memory hierarchies offer a much sharper formalization than biological sleep.11:00–14:05 · The hosts as informed peer 6/10 Empirical Methodology and Pareto Frontier Scaling on Benchmarks Alessio walks through an on-screen diagram of context pre-computation, reframing it as an agent recommending answers to itself. Charles and Charlie Snell validate the intuition while explaining the empirical goal of establishing Pareto frontier improvements on standard reasoning benchmarks.14:06–17:29 · The hosts as informed peer 6/10 Agentic Sleep Policies and Query Predictability Metrics Swyx challenges the team on whether 'sleep time' is merely marketing for standard data pre-processing. Charles and Kevin counter by explaining that determining when and how much to think ahead based on query predictability is what makes the mechanism fundamentally agentic.17:30–22:46 · The hosts as informed peer 6/10 GSM8K Benchmark Results and Test-Time Latency Cost Asymmetry Kevin explains the GSM8K benchmark curve where sleep time yields major gains when immediate answers are forced. Swyx observes that the curves converge at high test-time compute, prompting Charles to highlight the massive real-world latency cost asymmetry of user-facing tokens.22:46–27:16 · The hosts as informed peer 5/10 Frontier Model Scaling Dynamics and Reasoning Latency Analysis Alessio and Swyx explore latency variations across reasoning models like o1-mini and Claude 3.7 Sonnet. Charlie and Charles unpack their findings on query predictability and explain why sleep-time Pareto improvements hold across diverse thinking token parameterizations.27:17–30:53 · The hosts as informed peer 4/10 Letta Architectures: MemGPT v2 and Background Document Processing Charles details Letta's implementation architectures, including MemGPT v2 background agents managing context windows under strict token budgets for low-latency voice and document parsing pipelines. The hosts give space for an extensive technical architectural overview.30:54–34:01 · The hosts as informed peer 6/10 Stateful Agent Prerequisites, Industry Comparisons, and Concluding Thoughts Swyx demonstrates domain expertise by advising the team to formulate an intuitive messaging ratio analogous to Noam Brown's inference compute framing, and brings up concurrent tools like LangMem. Charles emphasizes that memory is an absolute prerequisite for sleep-time compute scaling.0:00–4:34 · Guest teaching 3/10 Introducing Sleep-Time Compute and Guest Research Backgrounds Swyx introduces the guests, noting their prior research background on MemGPT and test-time compute. Swyx inquires about the strategic naming shift from LLM OS to stateful agents, prompting Charles to clarify how the systems infrastructure enables the end product.4:35–8:38 · Guest teaching 5/10 Core Motivation: Stateful Context and Pre-Computing Intermediate Inferences Charles explains the core concept of sleep-time compute through pre-computing intermediate mathematical inferences and codebase exploration. Alessio asks about the origins of the biological brain analogy, which Charles clarifies by connecting cognitive memory consolidation to database background indexing.8:39–10:59 · Guest teaching 6/10 Defining Sleep-Time Compute: Cognitive versus Systems-Level Analogies Alessio presses on how sleep time is technically demarcated in an LLM lifecycle. Charles educates the hosts by defining sleep time as any post-training time that is not test time, arguing that systems-level memory hierarchies offer a much sharper formalization than biological sleep.11:00–14:05 · Guest teaching 4/10 Empirical Methodology and Pareto Frontier Scaling on Benchmarks Alessio walks through an on-screen diagram of context pre-computation, reframing it as an agent recommending answers to itself. Charles and Charlie Snell validate the intuition while explaining the empirical goal of establishing Pareto frontier improvements on standard reasoning benchmarks.14:06–17:29 · Guest teaching 5/10 Agentic Sleep Policies and Query Predictability Metrics Swyx challenges the team on whether 'sleep time' is merely marketing for standard data pre-processing. Charles and Kevin counter by explaining that determining when and how much to think ahead based on query predictability is what makes the mechanism fundamentally agentic.17:30–22:46 · Guest teaching 4/10 GSM8K Benchmark Results and Test-Time Latency Cost Asymmetry Kevin explains the GSM8K benchmark curve where sleep time yields major gains when immediate answers are forced. Swyx observes that the curves converge at high test-time compute, prompting Charles to highlight the massive real-world latency cost asymmetry of user-facing tokens.22:46–27:16 · Guest teaching 5/10 Frontier Model Scaling Dynamics and Reasoning Latency Analysis Alessio and Swyx explore latency variations across reasoning models like o1-mini and Claude 3.7 Sonnet. Charlie and Charles unpack their findings on query predictability and explain why sleep-time Pareto improvements hold across diverse thinking token parameterizations.27:17–30:53 · Guest teaching 6/10 Letta Architectures: MemGPT v2 and Background Document Processing Charles details Letta's implementation architectures, including MemGPT v2 background agents managing context windows under strict token budgets for low-latency voice and document parsing pipelines. The hosts give space for an extensive technical architectural overview.30:54–34:01 · Guest teaching 4/10 Stateful Agent Prerequisites, Industry Comparisons, and Concluding Thoughts Swyx demonstrates domain expertise by advising the team to formulate an intuitive messaging ratio analogous to Noam Brown's inference compute framing, and brings up concurrent tools like LangMem. Charles emphasizes that memory is an absolute prerequisite for sleep-time compute scaling.0:00–4:34 · Guest disagreement 1/10 Introducing Sleep-Time Compute and Guest Research Backgrounds Swyx introduces the guests, noting their prior research background on MemGPT and test-time compute. Swyx inquires about the strategic naming shift from LLM OS to stateful agents, prompting Charles to clarify how the systems infrastructure enables the end product.4:35–8:38 · Guest disagreement 1/10 Core Motivation: Stateful Context and Pre-Computing Intermediate Inferences Charles explains the core concept of sleep-time compute through pre-computing intermediate mathematical inferences and codebase exploration. Alessio asks about the origins of the biological brain analogy, which Charles clarifies by connecting cognitive memory consolidation to database background indexing.8:39–10:59 · Guest disagreement 2/10 Defining Sleep-Time Compute: Cognitive versus Systems-Level Analogies Alessio presses on how sleep time is technically demarcated in an LLM lifecycle. Charles educates the hosts by defining sleep time as any post-training time that is not test time, arguing that systems-level memory hierarchies offer a much sharper formalization than biological sleep.11:00–14:05 · Guest disagreement 1/10 Empirical Methodology and Pareto Frontier Scaling on Benchmarks Alessio walks through an on-screen diagram of context pre-computation, reframing it as an agent recommending answers to itself. Charles and Charlie Snell validate the intuition while explaining the empirical goal of establishing Pareto frontier improvements on standard reasoning benchmarks.14:06–17:29 · Guest disagreement 2/10 Agentic Sleep Policies and Query Predictability Metrics Swyx challenges the team on whether 'sleep time' is merely marketing for standard data pre-processing. Charles and Kevin counter by explaining that determining when and how much to think ahead based on query predictability is what makes the mechanism fundamentally agentic.17:30–22:46 · Guest disagreement 2/10 GSM8K Benchmark Results and Test-Time Latency Cost Asymmetry Kevin explains the GSM8K benchmark curve where sleep time yields major gains when immediate answers are forced. Swyx observes that the curves converge at high test-time compute, prompting Charles to highlight the massive real-world latency cost asymmetry of user-facing tokens.22:46–27:16 · Guest disagreement 1/10 Frontier Model Scaling Dynamics and Reasoning Latency Analysis Alessio and Swyx explore latency variations across reasoning models like o1-mini and Claude 3.7 Sonnet. Charlie and Charles unpack their findings on query predictability and explain why sleep-time Pareto improvements hold across diverse thinking token parameterizations.27:17–30:53 · Guest disagreement 1/10 Letta Architectures: MemGPT v2 and Background Document Processing Charles details Letta's implementation architectures, including MemGPT v2 background agents managing context windows under strict token budgets for low-latency voice and document parsing pipelines. The hosts give space for an extensive technical architectural overview.30:54–34:01 · Guest disagreement 1/10 Stateful Agent Prerequisites, Industry Comparisons, and Concluding Thoughts Swyx demonstrates domain expertise by advising the team to formulate an intuitive messaging ratio analogous to Noam Brown's inference compute framing, and brings up concurrent tools like LangMem. Charles emphasizes that memory is an absolute prerequisite for sleep-time compute scaling.0:00–4:34 · The hosts pushing back 2/10 Introducing Sleep-Time Compute and Guest Research Backgrounds Swyx introduces the guests, noting their prior research background on MemGPT and test-time compute. Swyx inquires about the strategic naming shift from LLM OS to stateful agents, prompting Charles to clarify how the systems infrastructure enables the end product.4:35–8:38 · The hosts pushing back 1/10 Core Motivation: Stateful Context and Pre-Computing Intermediate Inferences Charles explains the core concept of sleep-time compute through pre-computing intermediate mathematical inferences and codebase exploration. Alessio asks about the origins of the biological brain analogy, which Charles clarifies by connecting cognitive memory consolidation to database background indexing.8:39–10:59 · The hosts pushing back 2/10 Defining Sleep-Time Compute: Cognitive versus Systems-Level Analogies Alessio presses on how sleep time is technically demarcated in an LLM lifecycle. Charles educates the hosts by defining sleep time as any post-training time that is not test time, arguing that systems-level memory hierarchies offer a much sharper formalization than biological sleep.11:00–14:05 · The hosts pushing back 2/10 Empirical Methodology and Pareto Frontier Scaling on Benchmarks Alessio walks through an on-screen diagram of context pre-computation, reframing it as an agent recommending answers to itself. Charles and Charlie Snell validate the intuition while explaining the empirical goal of establishing Pareto frontier improvements on standard reasoning benchmarks.14:06–17:29 · The hosts pushing back 4/10 Agentic Sleep Policies and Query Predictability Metrics Swyx challenges the team on whether 'sleep time' is merely marketing for standard data pre-processing. Charles and Kevin counter by explaining that determining when and how much to think ahead based on query predictability is what makes the mechanism fundamentally agentic.17:30–22:46 · The hosts pushing back 3/10 GSM8K Benchmark Results and Test-Time Latency Cost Asymmetry Kevin explains the GSM8K benchmark curve where sleep time yields major gains when immediate answers are forced. Swyx observes that the curves converge at high test-time compute, prompting Charles to highlight the massive real-world latency cost asymmetry of user-facing tokens.22:46–27:16 · The hosts pushing back 1/10 Frontier Model Scaling Dynamics and Reasoning Latency Analysis Alessio and Swyx explore latency variations across reasoning models like o1-mini and Claude 3.7 Sonnet. Charlie and Charles unpack their findings on query predictability and explain why sleep-time Pareto improvements hold across diverse thinking token parameterizations.27:17–30:53 · The hosts pushing back 1/10 Letta Architectures: MemGPT v2 and Background Document Processing Charles details Letta's implementation architectures, including MemGPT v2 background agents managing context windows under strict token budgets for low-latency voice and document parsing pipelines. The hosts give space for an extensive technical architectural overview.30:54–34:01 · The hosts pushing back 3/10 Stateful Agent Prerequisites, Industry Comparisons, and Concluding Thoughts Swyx demonstrates domain expertise by advising the team to formulate an intuitive messaging ratio analogous to Noam Brown's inference compute framing, and brings up concurrent tools like LangMem. Charles emphasizes that memory is an absolute prerequisite for sleep-time compute scaling.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 23.1% · guest 76.9%0:00 · the hosts 23.1% · guest 76.9%3:00 · the hosts 20.7% · guest 79.3%3:00 · the hosts 20.7% · guest 79.3%6:00 · the hosts 22.4% · guest 77.6%6:00 · the hosts 22.4% · guest 77.6%9:00 · the hosts 34.7% · guest 65.3%9:00 · the hosts 34.7% · guest 65.3%12:00 · the hosts 24.6% · guest 75.4%12:00 · the hosts 24.6% · guest 75.4%15:00 · the hosts 14.2% · guest 85.8%15:00 · the hosts 14.2% · guest 85.8%18:00 · the hosts 10.6% · guest 89.4%18:00 · the hosts 10.6% · guest 89.4%21:00 · the hosts 18.5% · guest 81.5%21:00 · the hosts 18.5% · guest 81.5%24:00 · the hosts 8.3% · guest 91.7%24:00 · the hosts 8.3% · guest 91.7%27:00 · the hosts 24.3% · guest 75.7%27:00 · the hosts 24.3% · guest 75.7%30:00 · the hosts 39% · guest 61%30:00 · the hosts 39% · guest 61%33:00 · the hosts 19.8% · guest 80.2%33:00 · the hosts 19.8% · guest 80.2%
Sharpest disagreement ▶ 14:28 Charles rejecting brute force data pre-processing framing

Charles firmly pushes back against Swyx's skeptical characterization of sleep time compute as mere data pre-processing, explaining the mathematical limits and why dynamic context expansion is inherently agentic.

Hardest push from the hosts ▶ 14:05 Swyx questioning if sleep time is just rebranded pre-processing

Swyx directly challenges the core premise of the paper by asking whether the authors are simply rebranding classic data pre-processing under biological terminology.

Biggest teaching moment ▶ 9:03 Charles defining formal sleep time boundaries and systems analogies

Charles clarifies Alessio's confusion about when an LLM is 'sleeping' by establishing that sleep time encompasses all idle post-training GPU cycles and explaining why computer systems memory hierarchies are more precise than cognitive metaphors.

The host holds their own ▶ 31:45 Swyx offering Noam Brown ratio framework for messaging

Swyx showcases strong industry insight by advising the Letta team to distill their complex multi-axis charts into a clear, punchy ratio comparing seconds of sleep time to test time, citing Noam Brown's inference scaling strategy.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introducing Sleep-Time Compute and Guest Research Backgrounds 5312 Swyx introduces the guests, noting their prior research background on MemGPT and test-time compute. Swyx inquires about the strategic naming shift from LLM OS to stateful agents, prompting Charles to clarify how the systems infrastructure enables the end product.
Core Motivation: Stateful Context and Pre-Computing Intermediate Inferences 4511 Charles explains the core concept of sleep-time compute through pre-computing intermediate mathematical inferences and codebase exploration. Alessio asks about the origins of the biological brain analogy, which Charles clarifies by connecting cognitive memory consolidation to database background indexing.
Defining Sleep-Time Compute: Cognitive versus Systems-Level Analogies 4622 Alessio presses on how sleep time is technically demarcated in an LLM lifecycle. Charles educates the hosts by defining sleep time as any post-training time that is not test time, arguing that systems-level memory hierarchies offer a much sharper formalization than biological sleep.
Empirical Methodology and Pareto Frontier Scaling on Benchmarks 6412 Alessio walks through an on-screen diagram of context pre-computation, reframing it as an agent recommending answers to itself. Charles and Charlie Snell validate the intuition while explaining the empirical goal of establishing Pareto frontier improvements on standard reasoning benchmarks.
Agentic Sleep Policies and Query Predictability Metrics 6524 Swyx challenges the team on whether 'sleep time' is merely marketing for standard data pre-processing. Charles and Kevin counter by explaining that determining when and how much to think ahead based on query predictability is what makes the mechanism fundamentally agentic.
GSM8K Benchmark Results and Test-Time Latency Cost Asymmetry 6423 Kevin explains the GSM8K benchmark curve where sleep time yields major gains when immediate answers are forced. Swyx observes that the curves converge at high test-time compute, prompting Charles to highlight the massive real-world latency cost asymmetry of user-facing tokens.
Frontier Model Scaling Dynamics and Reasoning Latency Analysis 5511 Alessio and Swyx explore latency variations across reasoning models like o1-mini and Claude 3.7 Sonnet. Charlie and Charles unpack their findings on query predictability and explain why sleep-time Pareto improvements hold across diverse thinking token parameterizations.
Letta Architectures: MemGPT v2 and Background Document Processing 4611 Charles details Letta's implementation architectures, including MemGPT v2 background agents managing context windows under strict token budgets for low-latency voice and document parsing pipelines. The hosts give space for an extensive technical architectural overview.
Stateful Agent Prerequisites, Industry Comparisons, and Concluding Thoughts 6413 Swyx demonstrates domain expertise by advising the team to formulate an intuitive messaging ratio analogous to Noam Brown's inference compute framing, and brings up concurrent tools like LangMem. Charles emphasizes that memory is an absolute prerequisite for sleep-time compute scaling.

Statements from this episode (15)

Insight
Packer: Sleep-time compute during idle downtime is a major missed opportunity
“And practically speaking, you know, machines, they're not like humans, they can be run all the time. And there's a ton of downtime, both in advance of like questions being asked also like After questions have been asked too. So I think beyond just scaling at t…”
Charles Packer Apr 21, 2025 ▶ 1:02
Insight
Packer: Stateful AI agents require an LLM OS to maintain state
“To have a stateful agent, you need like an LMOS because you need something other than the LM to kind of maintain state.”
Charles Packer Apr 21, 2025 ▶ 3:56
Assertion Supported
Packer: Most test-time compute benchmarks assume stateless context delivery
“Most of the evaluations and work in this area, they kind of assume you get all of your context at test time. You get a math problem, you get like the entire setup and the question at test time.”
Charles Packer Apr 21, 2025 ▶ 5:23
Insight
Packer: Sleep-time compute re-represents token state into easily queryable formats
“In the test time compute setting, you know, here, the state is tokens and the kind of like sleep time, like indexing process is a re-representation of those tokens into something that is like more easily queryable and more flexible.”
Charles Packer Apr 21, 2025 ▶ 7:40
Insight
Packer: System-level memory analogies outperform cognitive metaphors for AI architecture
“I think the cognitive analogies are almost like a subset, kind of like Kevin was saying, of the systems level analogies. And I think the system level analogies are just a lot sharper because, you know, at the end of the day, like with tokens, it's like a memor…”
Charles Packer Apr 21, 2025 ▶ 10:04
Prediction Held up
Packer: ChatGPT will likely use sleep-time compute to learn offline
“Like if you activate sleep time compute on a chatbot like ChatGPT, it can like learn about you as you're not on ChatGPT.com. I think that's, you know, kind of what they're probably going to try to do. That's the direction they're going in.”
Charles Packer Apr 21, 2025 ▶ 10:35
Assertion Not checkable as stated
Packer: Nobody in the AI industry is actively scaling sleep-time compute
“How much we really can kind of take advantage of sleep time compute, which is effectively completely on mine today. Like nobody's really Scaling in the sleep time compute direction.”
Charles Packer Apr 21, 2025 ▶ 12:38
Insight
Packer: Sleep-time compute cannot be brute-forced due to diminishing returns
“There's definitely the aspect of, there's diminishing returns. So depending on what you're trying to do, you know, you will kind of reach a limit of how much you can re-represent the context. I think in this case, you know, with these like GSM, AK style questi…”
Charles Packer Apr 21, 2025 ▶ 14:44
Insight
Packer: Sleep-time compute avoids the user latency costs of test-time compute
“Assigning higher costs to tokens that come at test time, because once you're at test time, it kind of implies that a user is waiting. Something is waiting. It's either another process, a user, an event, and there is real cost to like every single token or like…”
Charles Packer Apr 21, 2025 ▶ 21:33
Insight
Packer: True AI agents run continuously rather than waiting for triggers
“And I think that's another aspect of like what makes something agentic, like not having to have a user send an event to trigger the machine to turn on, just allowing these machines to run all the time.”
Charles Packer Apr 21, 2025 ▶ 22:37
Assertion Supported
Snell: Sleep-time compute gains widen when upcoming questions are predictable
“We measured how predictable is the question from the sort of context. And we find that like, as the questions become more predictable, the benefit you see from applying sleep compute kind of widens.”
Charlie Snell Apr 21, 2025 ▶ 25:05
Assertion Supported
Packer: Sleep-time compute offers Pareto improvements across Claude 3.7 and DeepSeek
“It's like pretty consistent across like both 3.7 deep seek, three mini, which all like the way you actually scale the x-axis here is fundamentally quite different in each case with 3.7 extended thinking mode. The parameter you provide to scale it is different …”
Charles Packer Apr 21, 2025 ▶ 26:27
Disclosure
Packer: MemGPT v2 uses background agents for aggressive memory management
“We have two different ones we're releasing. One is like more chat focused. So it's basically like memgptv two. So chatting with an agent, but the main agent isn't actually managing the memory and you have like a very aggressive like background process that can…”
Charles Packer Apr 21, 2025 ▶ 29:26
Prediction Not checkable as stated
Packer: Background sleep-time agent architectures will be standard within two years
“I think those two, yeah, I think similar to memgpt, I think they're definitely like very good reference designs for just what's coming next. I think this sort of thing is, is just gonna be like the norm in like a year or two years.”
Charles Packer Apr 21, 2025 ▶ 30:44
Insight
Packer: Persistent memory is fundamentally required to scale sleep-time compute
“I think with the stateful agents thing in particular, I think one interesting thing about scaling sleep time compute is you need memory for it. You just fundamentally cannot do this sort of scaling without memory. Memory is like, you know, kind of table stakes…”
Charles Packer Apr 21, 2025 ▶ 32:44
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.