why aren't all 10 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Anthropic Deleted 80% of Claude Code System Prompt for Opus 5
“So yeah, we deleted 80% of the system prompt.”
Assertion Supported
Poetiq scored 55% on Humanity's Last Exam, outperforming Claude Opus 4.6
“AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state of the art. Which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55% on it.”
Assertion Supported
SemiAnalysis reports Claude Code generates four percent of all public commits
“There were some other stuff from like semi-analysis that four percent of all public commits are made by quad code.”
Prediction Held up
Kaplan: AI training and inference will see 3x-10x yearly efficiency gains
“I think that over time, as AI becomes more and more widespread, I think that we're going to really drive down the cost of inference and training dramatically from where we are right now. Sort of, three X to 10 X gains algorithmically, and in sort of scaling up…”
Assertion Supported
Ries: Anthropic created a perpetual purpose trust at Series C
“I think it was a series C when they finally established this thing called the long-term benefit trust, which is not a nonprofit foundation. It's actually what's called a perpetual purpose trust, which is a different legal category, but the same idea outside tr…”
Assertion Supported
Kaplan: Claude 4 will store and retrieve memories across context windows
“Claude IV can blow through its context window with a very complex task, but can also, ah, store memories as files or records, retrieve them in order to sort of keep doing work across many, many, many context windows.”
Assertion Supported
Habib: Anthropic matched RLHF performance using AI-generated feedback
“Anthropic had this very exciting paper just a couple of weeks ago where actually we're able to get similar results to RLHF without the H. So just actually having a second model provide the evaluation feedback as well. And that's obviously a lot more scalable.”
Assertion Supported
Hu: Anthropic Opus 5 achieved 30% on ARC-AGI
“You guys got, took Arc AGI three to 30%, which is incredible.”
Assertion Supported
French-Owen: Claude Code spawns Haiku sub-agents with dedicated context windows
“When you ask Claude code to do something, it will typically spawn and explore sub agent or like multiple ones. And basically each of those are running haiku to traverse the file system and kind of like explore what's there. And they're doing it in their own co…”
Assertion Supported
Joseph: Public estimates placed GPT-3 training cost at $5 million
“Like the public estimates for GP three, I remember, were that it cost five million dollars to train, which you're like, on the one hand, five million is kind of a lot, but it's like a lot for an individual person. It's not really a lot from like a company pers…”