The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 14 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Lukas Petersson Jun 4, 2026 ▶ 46:27 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Doug O'Laughlin Feb 24, 2026 ▶ 33:33 Claude Code for Finance + The Global Memory Shortage: Doug O'Laughlin, SemiAnalysis
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Shreya Rajpal Sep 25, 2025 ▶ 23:21 ⚡️Snowglobe: Simulations for your AI
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Supported
Ameisen: LLMs use internal circuits to backwards-plan rhyming poetry lines
“And two, this plan doesn't just control, like, what you're gonna rhyme with. It's also doing what's called like backwards planning, where it's like, well, because I need to finish with green, I'm not going to say illuminating the peaceful night, because then I…”
Emmanuel Ameisen Jun 6, 2025 ▶ 1:18:59 The Utility of Interpretability — Emmanuel Amiesen
Assertion Supported
Claude scored nearly twice as high as the next best model
“So we see that Claude right here is almost got twice the score of the nearest best model.”
Jack Hopkins Apr 27, 2025 ▶ 26:32 ⚡️Factorio Learning Environment: the ultimate Game Agent Eval — Jack Hopkins
Assertion Supported
Swyx: Claude wrapper Bolt.new reached $20M ARR
“The other one would be Bolt. There's a straight quad wrapper. And again, another now they've announced twenty million ARR, which is another step up from our eight million that we put on the title.”
Shawn Wang Jan 1, 2025 ▶ 44:32 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Micah Hill-Smith Jan 9, 2026 ▶ 30:09 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Partly supported
Chroma research finds Claude models lead in long-context utilization
“And you know, one thing Chroma released this context rod paper recently about context utilization and the cloud models are actually the best at using kind of like longer context.”
Alessio Fanelli Aug 6, 2025 ▶ 37:36 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Prediction Held up
Hershey predicts the Claude stream won't reach Victory Road within 16 days
“I think we have a little ways before we can beat the game in 16 days. I do not have a lot of faith that the current stream is gonna, gonna be standing in Victory Road in 13 days.”
David Hershey Mar 4, 2025 ▶ 32:42 How Claude Plays Pokémon was made
Assertion Supported
Anthropic's Pokémon research graph reflects a single run passing Lt. Surge
“The run that you saw that's, like, on the graph we put out alongside, like, in our research blog is, like, a single run that I have watched, like, get through At least surges Jim. And then it got a little past that. And the reason that that's where we stopped …”
David Hershey Mar 4, 2025 ▶ 33:55 How Claude Plays Pokémon was made
Assertion Supported
Malhotra: Claude's system prompt is written in the third person
“With Claude, we notice the system prompt is written in third person. It's written in third person. It's written as, the assistant is X, Y, Z.”
Karan Malhotra Apr 27, 2024 ▶ 9:40 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Supported
Haisfield: Claude tolerates imperfect URL syntax when generating WebSim apps
“Like, you don't need to get the exact syntax of an actual URL. Claude's smart enough to figure it out.”
Rob Haisfield Apr 27, 2024 ▶ 36:58 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Assertion Supported
Haisfield: WebSim accurately generates external RSS feeds to fetch live news data
“It just hallucinated a correct RSS feed and brought that in to its into this, I guess, you know, this wasn't a part of its like context window or anything because it's just displaying this stuff.”
Rob Haisfield Apr 27, 2024 ▶ 51:16 This World Does Not Exist — Joscha Bach, Karan Malhotra, Rob Haisfield (WorldSim, WebSim, Liquid AI)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.