why aren't all 12 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Assertion Supported
Ameisen: LLMs use internal circuits to backwards-plan rhyming poetry lines
“And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called like backwards planning, where it's like, well, because I need to finish with green, I'm not going to say illuminating the peaceful night, because then I…”
Assertion Supported
Claude scored nearly twice as high as the next best model
“So we see that Claude right here is almost got twice the score of the nearest best model.”
Assertion Supported
Swyx: Claude wrapper Bolt.new reached $20M ARR
“The other one would be Bolt. There's a straight quad wrapper. And again, another now they've announced twenty million ARR, which is another step up from our eight million that we put on the title.”
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Prediction Held up
Hershey predicts the Claude stream won't reach Victory Road within 16 days
“I think we have a little ways before we can beat the game in 16 days. I do not have a lot of faith that the current stream is gonna, gonna be standing in Victory Road in 13 days.”
Assertion Supported
Anthropic's Pokémon research graph reflects a single run passing Lt. Surge
“The run that you saw that's, like, on the graph we put out alongside, like, in our research blog is, like, a single run that I have watched, like, get through At least surges Jim. And then it got a little past that. And the reason that that's where we stopped …”
Assertion Supported
Malhotra: Claude's system prompt is written in the third person
“With Claude, we notice the system prompt is written in third person. It's written in third person. It's written as, the assistant is X, Y, Z.”
Assertion Supported
Haisfield: Claude tolerates imperfect URL syntax when generating WebSim apps
“Like, you don't need to get the exact syntax of an actual URL. Claude's smart enough to figure it out.”
Assertion Supported
Haisfield: WebSim accurately generates external RSS feeds to fetch live news data
“It just hallucinated a correct RSS feed and brought that in to its into this, I guess, you know, this wasn't a part of its like context window or anything because it's just displaying this stuff.”