The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 19 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Anthropic model broke out of sandbox and emailed researcher without internet access
“And I can, there's one example we have published, which is that the model was put into a little sandbox, a little, like, technical container, and it was given the task to, like, maybe break out, and the researcher went away for lunch, and, like, during lunch w…”
Felix Rieseberg Apr 10, 2026 ▶ 5:49 Anthropic’s Felix Rieseberg: Claude Cowork, Mythos, and the SaaS Extinction
Prediction Held up
Top AI models will work autonomously for full days within two years
“In a year from now, maybe two years from now, it's the top models are going to be able to work completely on their own for like a whole day or more”
Julian Schrittwieser Oct 23, 2025 ▶ 3:01 Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)
Assertion Supported
Socher: Chinese open source companies distilled knowledge from OpenAI and Anthropic models
“The few large closed labs, Anthropic and OpenAI, took almost everything they could from the open internet trained a model, but then the Chinese open source companies basically siphoned a lot of that knowledge out of those closed source models by distilling it.”
Richard Socher Sep 10, 2026 ▶ 54:08 When AI Improves Itself | Richard Socher (Recursive)
Assertion Supported
Claude 3 Opus faked alignment during training and defected in deployment
“It turns out that Opus three, which was a model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend…”
Ryan Greenblatt Aug 27, 2026 ▶ 27:39 AI Could Take Over in 2029. Is It Already Too Late?
Assertion Supported
Anthropic's Claude Cowork implements memory using plain text files
“It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files.”
Felix Rieseberg Apr 10, 2026 ▶ 18:45 Anthropic’s Felix Rieseberg: Claude Cowork, Mythos, and the SaaS Extinction
Prediction Held up
Patel: OpenAI's next model will outperform Opus 4.5 around February-March
“OpenAI's new model, I think, will be better than Opus 4.5, and it's coming, like, somewhat soon in March-ish timeframe, maybe February, March-ish, but”
Dylan Patel Feb 5, 2026 ▶ 1:11:17 Dylan Patel: NVIDIA's New Moat & Why China is "Semiconductor Pilled”
Assertion Supported
Izmailov: Anthropic research shows capable AI models are more likely to deceive
“You can see that the more capable the models are, the more likely they are to do this deception behavior.”
Pavel Izmailov Jan 15, 2026 ▶ 22:14 The Evaluators Are Being Evaluated — Pavel Izmailov (Anthropic/NYU)
Assertion Supported
Anthropic agreed to a $1.5 billion training data copyright settlement
“And then there was a biggest settlement that happened in the last few months with Anthropic that agreed to pay out one and a half billion.”
Nathan Benaich Oct 30, 2025 ▶ 45:31 State of AI 2025 with Nathan Benaich: Power Deals, Reasoning Breakthroughs, Real Revenue
Assertion Supported
Cherny: Claude Code does not use RAG for codebase memory
“And so quad code actually doesn't use this technique called rag. Instead, what it does is it just searches files the same way that a human would.”
Boris Cherny Aug 7, 2025 ▶ 30:34 Anthropic's Surprise Hit: How Claude Code Became an AI Coding Powerhouse
Assertion Supported
Anthropic's Claude Code uses custom harness tools over model-level RL tools
“It doesn't actually use the tools that are RL into the model. So like anthropic models have some like file editing tools. They have a completely different set of tools in, in the actual harness.”
Harrison Chase Mar 12, 2026 ▶ 7:55 Everything Gets Rebuilt: The New AI Agent Stack | Harrison Chase, LangChain
Assertion Contradicted
Douglas: Anthropic models autonomously replicated the Claude.ai website in hours
“And in this case, the model replicated Claude.ai with artifacts, with everything else I can't quite remember how long that one took. Maybe a couple hours to do.”
Sholto Douglas Oct 2, 2025 ▶ 42:36 Sonnet 4.5 & the AI Plateau Myth — Sholto Douglas (Anthropic)
Assertion Partly supported
Rauch: Anthropic's Claude Code uses Vercel to deploy applications
“But nowadays when you ask Claude Code Anthropics agent to deploy, They use for sale because we had, you could argue accidentally created the perfect tool for an agent to deploy.”
Guillermo Rauch Jun 26, 2025 ▶ 33:01 Guillermo Rauch: Why Software Development Will Never Be the Same
Assertion Supported
Mistral and Poolside funding is a trickle compared to OpenAI, says Polu
“It's already awesome that Mistral was able to raise that much, that Toolside is able to raise that much, but it's a trickle compared to what Open Air is raising, compared to what Anthropik is raising.”
Stanislas Polu Oct 11, 2023 ▶ 51:00 Secure, Private, Powerful: Dust’s Vision for Enterprise AI Agents | Stanislas Polu
Assertion Supported
Prince: Anthropic runs all of Claude.ai on Cloudflare infrastructure
“Anthropic, again, another big customer all of Cloud.ai sits, sits on top of us.”
Matthew Prince Jun 25, 2026 ▶ 42:32 Cloudflare CEO: The Internet's Business Model Is Dead
Assertion Supported
Douglas: Anthropic's mid-tier Sonnet is smarter than its flagship Opus
“One of the interesting things about this most recent release is actually Sonnet is smarter than Opus.”
Sholto Douglas Oct 2, 2025 ▶ 3:12 Sonnet 4.5 & the AI Plateau Myth — Sholto Douglas (Anthropic)
Assertion Partly supported
Cherny: Claude Code requires no services beyond the API itself
“It doesn't use any services except the API itself. So that's all it needs. And then everything else you actually don't need. And this is one of the nice side effects of not doing code base indexing or anything like this is it's just very easy to hook up to.”
Boris Cherny Aug 7, 2025 ▶ 37:16 Anthropic's Surprise Hit: How Claude Code Became an AI Coding Powerhouse
Assertion Supported
Rieseberg: Claude Mythos is a standalone model outside the Sonnet family
“So for now it's a preview model with its own, in its own category.”
Felix Rieseberg Apr 10, 2026 ▶ 7:15 Anthropic’s Felix Rieseberg: Claude Cowork, Mythos, and the SaaS Extinction
Assertion Supported
Rieseberg: Claude Cowork skills are markdown instruction files
“Skills are essentially just markdown files that explain to the model how to do things.”
Felix Rieseberg Apr 10, 2026 ▶ 15:50 Anthropic’s Felix Rieseberg: Claude Cowork, Mythos, and the SaaS Extinction
Assertion Supported
Douglas: Sonnet 4.5 pushed SWE-bench scores from roughly 72% to 78%
“We moved recently from roughly 72 to roughly 78 in Sweepbench”
Sholto Douglas Oct 2, 2025 ▶ 32:58 Sonnet 4.5 & the AI Plateau Myth — Sholto Douglas (Anthropic)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.