why aren't all 16 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Contradicted
Cherny: Anthropic Can No Longer Demonstrate Prompt Injection on Opus 5
“So essentially, if you combine a well-aligned model, so this is, like, essentially three years of research into alignment, With a prompt injection classifier, which we run for all traffic, and what this is doing is it's based on Chrysola's mechanistic interpre…”
Assertion Partly supported
Srinivas: Google possessed only fourth or fifth best AI models in 2023-2024
“And 2024, 20 23 especially, and large part of 2024 too, Google had like, maybe a fourth or fifth best models at any moment. So as a startup outside Google, you had access to AI that was better than what Google internally had.”
Assertion Supported
Anthropic Deleted 80% of Claude Code System Prompt for Opus 5
“So yeah, we deleted 80% of the system prompt.”
Assertion Contradicted
Ries: FTX bankruptcy's Anthropic stake exceeds the entire fraud value
“Apparently the stake that the bankruptcy has of those shares is worth more than the whole, than all of the entire fraud by a lot.”
Assertion Supported
Poetiq scored 55% on Humanity's Last Exam, outperforming Claude Opus 4.6
“AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state of the art. Which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55% on it.”
Assertion Contradicted
Mercury data shows 70 percent of startups choose Claude
“There was some stuff from Mercury that like, 70% of startups are, you know, choosing quad as their model of choice.”
Assertion Supported
SemiAnalysis reports Claude Code generates four percent of all public commits
“There were some other stuff from like semi-analysis that four percent of all public commits are made by quad code.”
Prediction Held up
Kaplan: AI training and inference will see 3x-10x yearly efficiency gains
“I think that over time, as AI becomes more and more widespread, I think that we're going to really drive down the cost of inference and training dramatically from where we are right now. Sort of, three X to 10 X gains algorithmically, and in sort of scaling up…”
Assertion Partly supported
Cherny: Claude Code uses Bun sandboxes to orchestrate dynamic agent workflows
“What a dynamic workflow is, is essentially we have the Bun runtime. We use Bun as a sandbox, and we start a virtual machine within Bun, and we let Cloud start a lot of agents and orchestrate them.”
Assertion Supported
Ries: Anthropic created a perpetual purpose trust at Series C
“I think it was a series C when they finally established this thing called the long-term benefit trust, which is not a nonprofit foundation. It's actually what's called a perpetual purpose trust, which is a different legal category, but the same idea outside tr…”
Assertion Contradicted
Joseph: Original scaling laws paper spanned 11 orders of magnitude
“Like, you know, the scaling, I think the original scaling laws paper had, like, 11 orders of magnitude, and there was, like, this intense debate on whether it would continue for, like, another point.”
Assertion Supported
Kaplan: Claude 4 will store and retrieve memories across context windows
“Claude IV can blow through its context window with a very complex task, but can also, ah, store memories as files or records, retrieve them in order to sort of keep doing work across many, many, many context windows.”
Assertion Supported
Habib: Anthropic matched RLHF performance using AI-generated feedback
“Anthropic had this very exciting paper just a couple of weeks ago where actually we're able to get similar results to RLHF without the H. So just actually having a second model provide the evaluation feedback as well. And that's obviously a lot more scalable.”
Assertion Supported
Hu: Anthropic Opus 5 achieved 30% on ARC-AGI
“You guys got, took Arc AGI three to 30%, which is incredible.”
Assertion Supported
French-Owen: Claude Code spawns Haiku sub-agents with dedicated context windows
“When you ask Claude code to do something, it will typically spawn and explore sub agent or like multiple ones. And basically each of those are running haiku to traverse the file system and kind of like explore what's there. And they're doing it in their own co…”
Assertion Supported
Joseph: Public estimates placed GPT-3 training cost at $5 million
“Like the public estimates for GP three, I remember, were that it cost five million dollars to train, which you're like, on the one hand, five million is kind of a lot, but it's like a lot for an individual person. It's not really a lot from like a company pers…”