why aren't all 10 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Fredrikson: Gray Swan Found Jailbreaks in Every OpenClaw User Trajectory Tested
“So we just have a bunch of trajectories of actual people using OpenClaw. And tons and tons of different scenarios and just threw shade at it and like found breaks for each and every one of them, right?”
Insight
Shah: OpenClaw memory fails because LLMs frequently skip search tool calls
“The way OpenClaw has with QMD or with whatever memory plugin you use, it inherently relies on tools to search through these memory.md files that it prepares. So, you know, like, what did I decide about the API? Then agent will decide to search, and sometimes i…”
Assertion Supported
Shah: OpenClaw's 15-message replay fails prompt caching and costs 10x more
“The way OpenClaw does it is it essentially sends back the last 15 messages in the conversation and it essentially uses that back and forth. And I mean, the approach itself is not ideal because you will, like, you are not doing any, like, you're not utilizing a…”
Insight
Krentsel: Improving AI models should architect their own agent harnesses over humans
“We believe All of that needs to be improvable by the agent, especially as the agents keep getting better, because as they keep getting better, it's this better lesson. You don't want to over-specialize because you don't want the human kind of deciding all thes…”
Assertion Supported
Krentsel: OpenClaw, Pi, and Claude Code hardcode static agent policies
“These are all policy decisions that are static, that are defined for OpenClaw, or for Pi, or for Clawed code, if you look at their Source code. And so that is the kind of, that is the policy of what an agent is, the tools it can use, the skills it has, how it …”
Assertion Supported
Cohen: OpenClaw logs all messages in plain text
“I started to see the size of the code base and the number of dependencies and some other things like logging all messages in plain text that just made me a bit apprehensive to use it for like production use cases to build a business on it.”
Assertion Supported
Shah: Supermemory outperformed Claude Code and OpenClaw benchmarks by almost 50%
“So the Claude code one performed the worst, and OpenClaw slightly more than that, and SuperMemory is the highest, and you can see that, you know, it's like a pretty significant difference, like almost 50%.”
Assertion Partly supported
Krentsel: OpenClaw agent threads cannot be interrupted during active execution
“Once an agent kind of goes and starts working on a thing in a thread, that thread is not interruptible. As the thing is off working, if you try to ping it just won't respond. You don't even know what it's doing, which is really frustrating.”
Prediction Not checkable as stated
Nathan: ChatGPT Work will not completely replace open-source OpenClaw
“I don't think so. I think that there's going to be, you know, there's always a need for, like, this, like, incredible, like, open source technology that, that team has built”
Disclosure
NVIDIA mandates running OpenClaw in isolated Brev cloud VMs
“Internally people want to run this
and we know we have to be really careful from the security implications. Do we let this run on the corporate network securities guidance was, Hey, run this on breath. It's in, you know, it's a VM. It's sitting in the cloud. …”