Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Packer: Sleep-time compute offers Pareto improvements across Claude 3.7 and DeepSeek

Charles Packer · Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin) · Apr 21, 2025 · at 26:27

Charles Packer, Letta co-founder and researcher, discusses benchmark results showing sleep-time compute consistently boosts efficiency across multiple major frontier LLMs.

0:00 / 0:46exact quote · 46.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“It's like pretty consistent across like both 3.7 deep seek, three mini, which all like the way you actually scale the x-axis here is fundamentally quite different in each case with 3.7 extended thinking mode. The parameter you provide to scale it is different from like O three reasoning effort. And then to, I guess, enable this sort of like to send, to tell like R one on the API, whether or not to like go to 10 K tokens versus like two K tokens also is like a, you know, different mechanism. Kind of like, irregardless of how you like push the agent to go further and also like kind of cross across frontier model categories, this sort of observation kind of holds that Yeah. It's kind of like a Pareto, Pareto benefit to apply sleep time compute.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Charles Packer

Insight
Packer: Sleep-time compute during idle downtime is a major missed opportunity
“And practically speaking, you know, machines, they're not like humans, they can be run all the time. And there's a ton of downtime, both in advance of like questions being asked also like After questions have been asked too. So I think beyond just scaling at t…”
Charles Packer Apr 21, 2025 ▶ 1:02 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
Insight
Packer: True AI agents run continuously rather than waiting for triggers
“And I think that's another aspect of like what makes something agentic, like not having to have a user send an event to trigger the machine to turn on, just allowing these machines to run all the time.”
Charles Packer Apr 21, 2025 ▶ 22:37 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
Prediction Not checkable as stated
Packer: Background sleep-time agent architectures will be standard within two years
“I think those two, yeah, I think similar to memgpt, I think they're definitely like very good reference designs for just what's coming next. I think this sort of thing is, is just gonna be like the norm in like a year or two years.”
Charles Packer Apr 21, 2025 ▶ 30:44 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
Insight
Packer: Sleep-time compute re-represents token state into easily queryable formats
“In the test time compute setting, you know, here, the state is tokens and the kind of like sleep time, like indexing process is a re-representation of those tokens into something that is like more easily queryable and more flexible.”
Charles Packer Apr 21, 2025 ▶ 7:40 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
Insight
Packer: Stateful AI agents require an LLM OS to maintain state
“To have a stateful agent, you need like an LMOS because you need something other than the LM to kind of maintain state.”
Charles Packer Apr 21, 2025 ▶ 3:56 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
Assertion Supported
Packer: Most test-time compute benchmarks assume stateless context delivery
“Most of the evaluations and work in this area, they kind of assume you get all of your context at test time. You get a math problem, you get like the entire setup and the question at test time.”
Charles Packer Apr 21, 2025 ▶ 5:23 Sleep-Time Compute — Letta AI (Charles Packer, Charlie Snell, Kevin Lin)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.