The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 235 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 3 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Hershey: Anthropic Study Found Claude Treats Named Characters Better
“Anthropic actually did like a blinded study of like named characters versus unnamed characters in different settings, and Claude like actually does clearly prefer and is nicer to named characters, which is an interesting thing.”
David Hershey Apr 5, 2025 ▶ 18:16 Claude Plays Pokémon Hackathon: Escape from Mt. Moon!
Opinion
Swyx: Frontier Models Exist Primarily to Distill Smaller, Usable Models
“Even GPT 4.5 is too expensive. Normally it's really gonna use it in, in any reasonable quantity. Like, you know, Claude 3.5 Opus, like if it does exist, still not like, you know, the thing that we actually use is Sonnet, right? So like, it's almost like a depl…”
Shawn Wang Mar 23, 2025 ▶ 3:58 The Magic of LLM Distillation — Rishabh Agarwal, Google DeepMind
Opinion
Snipd CEO: Claude is the best model at phrasing and personality
“Like, in my opinion, Claude is the best one when it comes to the way it formulates things.”
Kevin Ben-Smith Mar 14, 2025 ▶ 51:20 Snipd: The AI Podcast App for Learning — with CEO Kevin Ben-Smith
Disclosure
Nguyen: Claude 2's distinct personality was unintentional until Claude 3
“People said, like, Cloud II is, like, so much better at, like, writing and, like, has a certain personality, even though it was, like, unintentional at all. And we did not pay that much attention and didn't know even how to, like, productionize this property o…”
Karina Nguyen Feb 1, 2025 ▶ 26:16 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Opinion
Swix: Bearish on computer-use AI agents due to cost, speed, and accuracy
“I have been very bearish in computer use because they're slow. They're expensive. They're imprecise. Like the accuracy is horrible. Still, even with Anthropix new stuff, I'm really waiting to see what opening I might do to change my opinions.”
Shawn Wang Feb 1, 2025 ▶ 55:02 The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI
Assertion Supported
Swyx: Claude wrapper Bolt.new reached $20M ARR
“The other one would be Bolt. There's a straight quad wrapper. And again, another now they've announced twenty million ARR, which is another step up from our eight million that we put on the title.”
Shawn Wang Jan 1, 2025 ▶ 44:32 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Opinion
Neubig: Claude Is The Best Agent Model, Open Models Lag Behind
“I still am under the impression that Claude is the best. The other closed models are, you know, not quite as good, and then the open models are a little bit behind that.”
Graham Neubig Dec 25, 2024 ▶ 15:16 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Opinion
Neubig: GPT Loops On Errors While Claude Tries New Approaches
“So, like, GPT doesn't have very good air recovery ability. And so, because of this, it will go into loops and do the same thing over and over and over again, whereas Claude does not do this.”
Graham Neubig Dec 25, 2024 ▶ 14:25 Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
Assertion Supported
Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60%
“And the opening I spend at the beginning, at the end of last year in November of 23 was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume.”
Pranav Reddy Dec 21, 2024 ▶ 5:08 The State of AI Startups in 2024 [LS Live @ NeurIPS]
Assertion Not checkable as stated
Schluntz: String replacement is the most reliable file-editing tool for LLMs
“We did a few different experiments with like different ways to specify how to edit a file and string replace. Basically the model has to write out the existing version of the string and then a new version, and that just gets swapped in. We found that to be the…”
Erik Schluntz Nov 28, 2024 ▶ 24:55 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Prediction Not checkable as stated
Schluntz: Production AI agent applications will be bespoke, not off-the-shelf
“You know, I think that might be useful for hobbyists and demos, but the ultimate end applications are going to be bespoke. And so we just want to make sure that the model's great at any tool that it uses”
Erik Schluntz Nov 28, 2024 ▶ 44:42 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Disclosure
Anthropic: Tool engineering mattered more than prompt engineering for SWE-bench
“I would say actually we did more engineering of the tools than the overall prompt.”
Erik Schluntz Nov 28, 2024 ▶ 22:53 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Insight
Crivello: Tasks capable of using APIs must stay API-driven over computer use
“My philosophy about it is anything that can be done with an API must be done by an API or should be done by an API for a very long time.”
Florent Crivello Nov 15, 2024 ▶ 38:16 Agents @ Work: Lindy.ai (with live demo!)
Assertion Supported
Polu: Claude Sonnet executes an unpublicized chain-of-thought step during function calling
“They kind of innovated in an interesting way, which was never quite publicized, but it's that they have that kind of chain of thoughts step whenever you use a Clouds model or Sonnet model with function calling. That chain of service step doesn't exist when you…”
Stanislas Polu Nov 11, 2024 ▶ 42:20 Agents @ Work: Dust.tt — with Stanislas Polu
Opinion
Liu: Claude 3 Haiku outperforms OpenAI models at function calling
“Overall, I'm like super happy with the anthropic models compared to the OpenAM models. Like, Sonnet is very cost effective. Haiku is, in function calling, it's actually better.”
Jason Liu Apr 24, 2024 ▶ 21:57 High Agency Pydantic over VC Backed Frameworks — with Jason Liu of Instructor
Insight
Lambert: Claude's constitution dictates output priorities, not model beliefs
“If you look at Claude's constitution, like, that doesn't mean the model believes these things. It's just trying Trained and to prioritize these things.”
Nathan Lambert Jan 11, 2024 ▶ 13:06 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Dylan Patel Dec 5, 2023 ▶ 1:06:37 The State of Silicon and the GPU Poors - with Dylan Patel of SemiAnalysis
Assertion Not checkable as stated
Swyx: Top AI agent labs receive secret discounts from model providers
“Agent Labs get discounts from every model provider, and that's also very interesting when people compare public pricing of, like, a discounted cloud code from Anthopic versus what Anthopic does with Model Labs, with Agent Labs”
Shawn Wang Jul 10, 2026 ▶ 26:36 Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv
Prediction Not checkable as stated
Swyx: Frontier AI labs will not provide bespoke enterprise integration support
“The labs do not have 200 people dedicated to like, you know, being on call with you with Goldman Sachs going like, okay guys, what do you need? We got it. You need the Microsoft Teams zero integration. Got it. You don't use GitHub. You use this like weird org …”
Shawn Wang Jul 10, 2026 ▶ 24:20 Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv
Disclosure
Midha: AMP Foundry invested hundreds of millions into Anthropic this year
“We put a few hundred million dollars into Anthropic from our fund earlier this year.”
Anjney Midha Jun 18, 2026 ▶ 14:14 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Opinion
Midha: Anthropic's velocity came from standardizing on the transformer architecture
“Like, one of the reasons Anthropic has had extraordinary sort of velocity is because they picked the transform architecture and said, this is simple, let's double down on it, right? And now, luckily, there's enough investment going into space that we can affor…”
Anjney Midha Jun 18, 2026 ▶ 26:35 Why AI Labs With Unlimited GPUs Still Fail — Anjney Midha, AMP
Assertion Not checkable as stated
Awais: Claude tolerates tool errors and self-corrects, unlike open models
“Claude is actually really, really lenient for tool calls. So even if, you know, your coding agent harness messes up, it can figure out that, oh, I'm being sent this error and can fix itself. Not the case with you know open models”
Ahmad Awais Jun 6, 2026 ▶ 27:12 ⚡️Making DeepSeek v4 outperform Opus 4.7 with Taste — @AhmadAwais , CommandCode.ai
Disclosure
Chatbase customer model usage is split 50% OpenAI, 50% Anthropic and Google
“Maybe 50% is still on Okunai. Yeah. Yeah. And then 50 on everything else. Yeah. But everything else is like mainly Anthropic and Google.”
Yasser Elsaid May 2, 2026 ▶ 13:30 ⚡️ Competing with ChatGPT and Sierra, building a $10M ARR company — Yasser Elsaid, Founder, Chatbase
Assertion Not checkable as stated
Claude Sonnet Crashed on Duplicate Tool Names While OpenAI Handled the Error
“Sonic couldn't handle two tools with the same name in OpenAI, GPT, 5.2. It was like, ah, I can figure this out. So that was an interesting one that we learned by accident through a SEV.”
Sarah Sachs Apr 15, 2026 ▶ 55:35 Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
Insight
Rieseberg: Prompt Opus by stating goals, not specifying exact execution steps
“Honestly though, like I see that you're using Opus 4.6, right? Like my recommendation for people is increasingly don't worry about it anymore. Just like tell it what you want it to do. And it's probably going to figure out a way to do it.”
Felix Rieseberg Mar 17, 2026 ▶ 53:13 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Insight
Rieseberg: Cowork's most impressed users find unexpected capabilities
“Every single person who's like most amazed is usually amazed about a thing that I didn't even expect Cowork would be good at.”
Felix Rieseberg Mar 17, 2026 ▶ 2:06 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Assertion Supported
Rieseberg: Claude Cowork is Claude Code running in a sandboxed virtual machine
“Cowork is cloud code running in a virtual machine with a little bit of padding, a little bit more guardrails, making it a little safer, a little bit more convenient for people who don't want to first open up the terminal when they go to work.”
Felix Rieseberg Mar 17, 2026 ▶ 3:35 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Insight
Rieseberg: Claude Cowork skills can be as simple as a text message
“One thing that is very fun for me about skills in particular is that they're so easy to make. Like anyone can make a skill, like a text message could be a skill and they can be so hyper-personalized to you.”
Felix Rieseberg Mar 17, 2026 ▶ 29:06 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Insight
Rieseberg: A labs team should only tackle ideas no one else would
“The sort of the idea of a Labs team is that it should only work on things that make really no sense for anyone else to work on.”
Felix Rieseberg Mar 17, 2026 ▶ 1:27:04 Anthropic’s Felix Rieseberg on AI Coworkers, Local-First Agents, and the Future of Knowledge Work
Assertion Not checkable as stated
Eskildsen: Anthropic, Notion, and Cursor use Turbopuffer across three deployment models
“You can run Turbo Puffer either in SAS, right? That's what cursor does. You can run it in a single tenant cluster. So it's just you. That's what Notion does. And then you can run it in, in, in BYOC where everything is inside the customer's VPC. That's what, fo…”
Simon Eskildsen Mar 12, 2026 ▶ 37:25 Retrieval After RAG: Hybrid Search, Agents, and Database Design — Simon Eskildsen of Turbopuffer
Assertion Not checkable as stated
Wang: Anthropic Claude Cowork automates complex customer cohort data analysis
“Like for the first time you can actually get one shot data analysis, right? Which, you know, if you're going to do a customer database, analyze a cohort retention, right? That's just stuff that you had to do by hand before. And our team, the other, it was like…”
Sarah Wang Feb 19, 2026 ▶ 27:34 Inside AI’s $10B+ Capital Flywheel — Martin Casado & Sarah Wang of a16z
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Micah Hill-Smith Jan 9, 2026 ▶ 30:09 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Disclosure
Anthropic remains fully committed to MCP following its foundation donation
“Like the commitment of Anthropic is the same, right? I'm still, We still have the same people I'm helping with the SDKs. We're still super committed in our products to MCP. I'm still the lead core maintainer. Nothing has actually changed.”
David Soria Parra Dec 28, 2025 ▶ 1:01:23 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Disclosure
Anthropic created MCP so rapidly expanding internal teams could build integrations independently
“MCP before we even open source it was born of the idea of like, I'm in a company that is growing crazy. I'm in the development side of things, development tooling side of things. I will grow slower than the rest. How can I build something that they can all bui…”
David Soria Parra Dec 28, 2025 ▶ 28:18 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Supported
Block's Goose was the first open-source agent to integrate MCP
“Goose was the first open source agent interface or agent that reached out to us and worked with us to integrate MCP. And I think Rad is actually like technically the first non-anthropic contributor to MCP ever on like day two or something like that, like very,…”
David Soria Parra Dec 28, 2025 ▶ 1:12:42 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Supported
Google, Microsoft, Amazon, OpenAI, and Anthropic joined AAIF as platinum members
“You have Google, Microsoft, Amazon Block, Bloomberg, Cloudflare, OpenAI, Anthropic. Just a platinum member, create a foundation.”
David Soria Parra Dec 28, 2025 ▶ 1:35:53 One Year of MCP — with David Soria Parria and AAIF leads from OpenAI, Goose, Linux Foundation
Assertion Not checkable as stated
Yegge: Anthropic is hiring over 100 people for Claude Code
“They're hiring like a hundred plus people for cloud code in the next, I don't know, month. I mean, like they're going wild and that's just cloud code.”
Steve Yegge Dec 26, 2025 ▶ 30:13 Steve Yegge's Vibe Coding Manifesto: Why Claude Code Isn't It & What Comes After the IDE
Disclosure
Superhuman builds dynamic on-the-fly aggregation lambdas with Anthropic
“We're working right now with Anthropic to basically do kind of like a building on the fly, small, kind of a key component of lambdas that will build the code to do the aggregation.”
Loïc Houssier Dec 11, 2025 ▶ 25:26 The Future of Email: Superhuman CTO on Your Inbox As the Real AI Agent (Not ChatGPT) — Loïc Houssier
Assertion Supported
Anthropic Maintains an 80% One-Year Employee Retention Rate
“I'm referring to the exact same article where I think their retention, one year retention on employees is the 80%, which in AI world is, is quite wild.”
Deedy Das Nov 14, 2025 ▶ 21:26 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Insight
Enterprise LLM Churn Is Low Due to Long-Term Compute Commitments
“In terms of enterprises, often what will happen is they'll buy up large chunks of long-term compute and dedicated instances, in which case you just don't churn, right? Like this is what you use.”
Deedy Das Nov 14, 2025 ▶ 26:26 Anthropic, Glean & OpenRouter: How AI Moats Are Built with Deedy Das of Menlo Ventures
Prediction Open · timeframe Nov 2030
Anthropic and OpenAI will never open-source their high-performance inference kernels
“The high performance inference kernels that sort of drive a lot of, you know, anthropic and open AI and stuff, their models, those aren't open source. They're not going to be open source.”
Quentin Anthony Nov 3, 2025 ▶ 40:52 How Zyphra went all-in on AMD + Why Devs feel faster with AI but are slower — with Quentin Anthony
Assertion Supported
Merrill: Dario Amodei highlighted Terminal-Bench on the Claude model card
“I think one of the really key moments for us was getting onto the Claude IV model card. Being one of two benchmarks that Dario actually mentioned while releasing the model.”
Mike Merrill Oct 18, 2025 ▶ 3:58 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Insight
Krieger: AI Agents Must Support Both MCP and Visual Computer Use
“And that thing's never gonna have an MCP around it. Like, it's just like, who knows if the company created is even around much less like ready to sort of expose their kind of underlying constructs as API. So I think you will need to be able to do both.”
Mike Krieger Sep 30, 2025 ▶ 13:00 ⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
Opinion
Krieger: Claude Sonnet 4.5 Outperforms Opus at Generating 3D Games
“This is like officially good. It's like better than Opus at this. It's like, It generated this, like, great split-screen stereoscopic thing, three-dimensional, like, thing.”
Mike Krieger Sep 30, 2025 ▶ 4:24 ⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
Assertion Supported
Martin: Claude Code operates entirely without codebase indexing
“Clock code doesn't do any indexing. It's just doing, quote unquote, agentic retrieval, just using simple tool calls, for example, using grep, to kind of poke around your files, no indexing whatsoever, and obviously works extremely well.”
Lance Martin Sep 11, 2025 ▶ 17:02 Context Engineering for Agents - Lance Martin, LangChain
Assertion Partly supported
Chroma research finds Claude models lead in long-context utilization
“And you know, one thing Chroma released this context rod paper recently about context utilization and the cloud models are actually the best at using kind of like longer context.”
Alessio Fanelli Aug 6, 2025 ▶ 37:36 The AI Agenda: GPT5 leaks and the business of AI News — Steph Palazzolo, The Information
Opinion
McCloy: Claude users represent an exceptionally valuable audience for companies
“Claude, which is important for, you know, not necessarily huge in terms of raw number of users, but the people who do use Claude tend to be like a very valuable audience, especially for some types of company.”
Robert McCloy Jul 23, 2025 ▶ 7:28 AI is Eating Search
Assertion Supported
Anthropic finds a single LLM judge outperforms five specialized judges
“They initially started with five LLM as judges. So each one of these points had their own LLM as a judge. They tested the ability and accuracy of that LLM as judge collective to judge, and it actually didn't perform as well as one. So they replaced all of thos…”
Dylan Davis Jul 5, 2025 ▶ 11:03 ⚡️Anthropic vs Cognition on Multi-Agents: A Breakdown with Dylan Davis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.