why aren't all 12 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Anthropic model broke out of sandbox and emailed researcher without internet access
“And I can, there's one example we have published, which is that the model was put into a little sandbox, a little, like, technical container, and it was given the task to, like, maybe break out, and the researcher went away for lunch, and, like, during lunch w…”
Assertion Not checkable as stated
Anthropic's unreleased Claude Mythos model shows outsized cybersecurity capabilities
“Mythos is a unreleased frontier model. It's a general purpose model that was trained not specifically for cybersecurity or specifically for coding or specifically for software, but we have discovered what we believe to be outsized capabilities specifically in …”
Assertion Not checkable as stated
Rieseberg: AI models can execute week-long knowledge work tasks today
“The models we have today are actually quite capable. They're quite capable of running knowledge work of both of an extremely long time horizon, the kind of things that you give to someone and expect like a week later.”
Assertion Supported
Anthropic's Claude Cowork implements memory using plain text files
“It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files.”
Prediction Not checkable as stated
Rieseberg: Software creation skills will shift from code to human language
“My prediction is going to be That we are going to have a lot more software. That software is probably going to be slightly more specialized. I don't think everyone is going to build their own software. I think people will still build things and, like, share th…”
Prediction Not checkable as stated
Felix Rieseberg predicts AI progress is accelerating into larger capability gains
“We have reasons to believe the journey is accelerating so that the steps are going to get bigger and bigger.”
Prediction Not checkable as stated
Rieseberg: AI value will come from work organization, not model intelligence
“A lot of the value you can provide will probably be less on the agent side. It will be less on the model intelligence. There will be more about how do you help people organize their work”
Assertion Not checkable as stated
Cloud AI agent logins will trigger automated bank account lockouts
“If your bank sees you logging in from two separate places, say your computer and also a data center, it will probably lock down your account. And we'll ask you to come to a branch with a passport.”
Assertion Not checkable as stated
Microsoft employees initially viewed Visual Studio Code as a toy
“And when Visual Studio Code first was released inside the company, there was a feeling that this is a toy. This is not for real developers, because the real developers, they need Visual Studio, which is why Visual Studio Code is such a complicated long name, b…”
Assertion Supported
Rieseberg: Claude Mythos is a standalone model outside the Sonnet family
“So for now it's a preview model with its own, in its own category.”
Assertion Supported
Rieseberg: Claude Cowork skills are markdown instruction files
“Skills are essentially just markdown files that explain to the model how to do things.”
Prediction Not checkable as stated
Rieseberg expects Claude Cowork to ship meaningful updates every week
“So in practice, if you are a co-work user, you will probably continue to see fairly meaningful changes shipping every single week, quite a bit. There's really no end in sight, I think.”