why aren't all 14 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Packer: Sleep-time compute during idle downtime is a major missed opportunity
“And practically speaking, you know, machines, they're not like humans, they can be run all the time. And there's a ton of downtime, both in advance of like questions being asked also like After questions have been asked too. So I think beyond just scaling at t…”
Insight
Packer: True AI agents run continuously rather than waiting for triggers
“And I think that's another aspect of like what makes something agentic, like not having to have a user send an event to trigger the machine to turn on, just allowing these machines to run all the time.”
Prediction Not checkable as stated
Packer: Background sleep-time agent architectures will be standard within two years
“I think those two, yeah, I think similar to memgpt, I think they're definitely like very good reference designs for just what's coming next. I think this sort of thing is, is just gonna be like the norm in like a year or two years.”
Insight
Packer: Sleep-time compute re-represents token state into easily queryable formats
“In the test time compute setting, you know, here, the state is tokens and the kind of like sleep time, like indexing process is a re-representation of those tokens into something that is like more easily queryable and more flexible.”
Assertion Supported
Packer: Sleep-time compute offers Pareto improvements across Claude 3.7 and DeepSeek
“It's like pretty consistent across like both 3.7 deep seek, three mini, which all like the way you actually scale the x-axis here is fundamentally quite different in each case with 3.7 extended thinking mode. The parameter you provide to scale it is different …”
Insight
Packer: Stateful AI agents require an LLM OS to maintain state
“To have a stateful agent, you need like an LMOS because you need something other than the LM to kind of maintain state.”
Assertion Supported
Packer: Most test-time compute benchmarks assume stateless context delivery
“Most of the evaluations and work in this area, they kind of assume you get all of your context at test time. You get a math problem, you get like the entire setup and the question at test time.”
Insight
Packer: System-level memory analogies outperform cognitive metaphors for AI architecture
“I think the cognitive analogies are almost like a subset, kind of like Kevin was saying, of the systems level analogies. And I think the system level analogies are just a lot sharper because, you know, at the end of the day, like with tokens, it's like a memor…”
Prediction Held up
Packer: ChatGPT will likely use sleep-time compute to learn offline
“Like if you activate sleep time compute on a chatbot like ChatGPT, it can like learn about you as you're not on ChatGPT.com. I think that's, you know, kind of what they're probably going to try to do. That's the direction they're going in.”
Assertion Not checkable as stated
Packer: Nobody in the AI industry is actively scaling sleep-time compute
“How much we really can kind of take advantage of sleep time compute, which is effectively completely on mine today. Like nobody's really Scaling in the sleep time compute direction.”
Insight
Packer: Sleep-time compute cannot be brute-forced due to diminishing returns
“There's definitely the aspect of, there's diminishing returns. So depending on what you're trying to do, you know, you will kind of reach a limit of how much you can re-represent the context. I think in this case, you know, with these like GSM, AK style questi…”
Insight
Packer: Sleep-time compute avoids the user latency costs of test-time compute
“Assigning higher costs to tokens that come at test time, because once you're at test time, it kind of implies that a user is waiting. Something is waiting. It's either another process, a user, an event, and there is real cost to like every single token or like…”
Disclosure
Packer: MemGPT v2 uses background agents for aggressive memory management
“We have two different ones we're releasing. One is like more chat focused. So it's basically like memgptv two. So chatting with an agent, but the main agent isn't actually managing the memory and you have like a very aggressive like background process that can…”
Insight
Packer: Persistent memory is fundamentally required to scale sleep-time compute
“I think with the stateful agents thing in particular, I think one interesting thing about scaling sleep time compute is you need memory for it. You just fundamentally cannot do this sort of scaling without memory. Memory is like, you know, kind of table stakes…”