why aren't all 12 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Krieger: Knowledge work AI progress follows code's exponential curve on a delay
“And I think of knowledge work as being on a similar sort of exponential as code has been on, but just time shifted, right?”
Assertion Not checkable as stated
Sonnet 4.5 ran autonomously for 30 hours versus 7 for Opus 4
“So this, you know, we had a customer and internally, we also got like a 30 hour plus kind of execution versus I think Opus four was seven hours.”
Opinion
Krieger: Claude Sonnet 4.5 Outperforms Opus at Generating 3D Games
“This is like officially good. It's like better than Opus at this. It's like, It generated this, like, great split-screen stereoscopic thing, three-dimensional, like, thing.”
Opinion
Krieger: Current vision models lack the precision of skilled visual designers
“The models don't see as well as they could. They see, okay, you know, you ask them analyze a complex photo and they're able to do it, but I want them to be as persnickety as a like really good visual designer. Like, no, that looks, the baseline looks a little …”
Insight
Krieger: AI Agents Must Support Both MCP and Visual Computer Use
“And that thing's never gonna have an MCP around it. Like, it's just like, who knows if the company created is even around much less like ready to sort of expose their kind of underlying constructs as API. So I think you will need to be able to do both.”
Assertion Not checkable as stated
Krieger: Sonnet beat Opus on SWE-bench before users felt it was better
“Even when it was already outperforming Opus, for example, on sweet bench, people still didn't feel it was better, but then it continued to train and it was like now better than Opus and people don't want to switch back.”
Assertion Not checkable as stated
Krieger: Claude Sonnet 4.5 day-one traffic eclipsed Sonnet 4
“We have more traffic on Sonnet 4.5 than we had on Sonnet four. So basically it's already eclipsed Sonnet four.”
Assertion Not checkable as stated
Krieger: Sonnet 4.5 First Anthropic Model With Upstream Product Input
“What was interesting about this model in particular is that it was really the first one where Product was upstream of research and downstream of research.”
Disclosure
Krieger: Anthropic generates internal dashboards on demand with Claude
“Even internally we have found that having Cloud generate UI on demand for internal dashboards is useful, not just like an interesting demo.”
Opinion
Krieger: Figma Make effectively bridges design models with LLMs
“I actually think Figma did a really good job with Make in that they you can tell that there's much more of a bridge between the sort of underlying design model and what the LLMs are doing in kind of coordination between Sonnet”
Disclosure
Krieger: Anthropic to integrate agent harness into Claude AI workflows
“There's, Cloud AI, and I think you'll see us start bringing that harness into Cloud AI more and more for things like document creation, advanced research, and so it'll power a lot more of the agentic workflows in there.”
Disclosure
Krieger: Anthropic plans hosted computation options for Claude Agent SDK
“There's Cloud Code, which is directly built on top of the Cloud Agent SDK, and then there's external companies building on top of the Cloud Agent SDK, and I think we'll also offer ways in which if you want to run an agent off of that SDK, but have a lot of the…”