why aren't all 91 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Douglas: Examining training data and simple tweaks yields massive model gains
“A really high ROI use on, on your time would probably be to just go and look at the data and think hard about what, like the model is learning or doing and like make some tweaks that makes it like, there's so many, like you could even just The simplest things …”
Assertion Not checkable as stated
Douglas: Even AI pioneer Noam Shazeer only sees 10% of ideas work
“I once asked this question of Noam Chazier. And he was like, yeah, maybe like 10% of my ideas work, and that's not, right? You know, one of the, you know, an absolute genius, one of the best in the field. So if only 10% of his ideas work, then I think that, yo…”
Disclosure
Douglas: Anthropic intentionally deprioritized math reasoning unlike OpenAI and DeepMind
“You know, one thing that Anthropik, like, You noticeably hasn't focused on compared to DeepMind and to OpenAI is is mathematical reasoning, right? DeepMind and OpenAI have been pursuing mathematical reasoning because of the implications for science and for sci…”
Assertion Contradicted
Douglas: Anthropic models autonomously replicated the Claude.ai website in hours
“And in this case, the model replicated Claude.ai with artifacts, with everything else I can't quite remember how long that one took. Maybe a couple hours to do.”
Insight
Douglas: Coding is uniquely tractable for AI because execution is verifiable
“Coding is a uniquely tractable problem in some respects for the techniques that we have in terms, the data exists in many ways. You can containerize and run things in parallel. You can run unit tests, and so you can verify.”
Assertion Not checkable as stated
Cherny: Most code at Anthropic is now written using Claude Code
“You know, at this point, most code at Anthropic is written using quad code and almost everyone at Anthropic is using it every day.”
Disclosure
Cherny: Claude Code hit its stride with Sonnet 4 and Opus 4
“We started building quad code when it was still Sonnet at 3.5 and it was okay. And then with, you know, 3.6 and 3.7, it was still, it was fine. It was pretty good. But then when Sonnet four and Opus four came out, that's when it really hit its stride and we fe…”
Assertion Not checkable as stated
Cherny: Claude Code cut Anthropic engineering onboarding from weeks to days
“And what we saw is adanthropic technical onboarding used to take a few weeks, but now engineers are usually productive within the first few days.”
Assertion Partly supported
Rauch: Anthropic's Claude Code uses Vercel to deploy applications
“But nowadays when you ask Claude Code Anthropics agent to deploy, They use for sale because we had, you could argue accidentally created the perfect tool for an agent to deploy.”
Insight
Evans: Memory features in AI models create user stickiness, not network effects
“I think that stickiness, I don't think it's a network effect.”
Assertion Supported
Mistral and Poolside funding is a trickle compared to OpenAI, says Polu
“It's already awesome that Mistral was able to raise that much, that Toolside is able to raise that much, but it's a trickle compared to what Open Air is raising, compared to what Anthropik is raising.”
Assertion Supported
Prince: Anthropic runs all of Claude.ai on Cloudflare infrastructure
“Anthropic, again, another big customer all of Cloud.ai sits, sits on top of us.”
Disclosure
Anthropic team sprinted for 10 days to launch Claude Cowork
“My team sprinted on this for, I think, the last 10 days or so, which is accurate.”
Opinion
Rieseberg: Model Context Protocol (MCP) connectors are underrated
“MCP connectors are underrated.”
Disclosure
Rieseberg: Claude Code's genesis was pure local UX on existing models
“The very genesis of it was, what if Cloud, but instead of, like, in the cloud, it's running on your computer in your terminal. That, that is almost entirely UX. It's the same model. It's the same core capabilities.”
Insight
Rieseberg: AI models are grown rather than built, making capabilities unpredictable
“We often say that models are more grown than built out of the nature of how these language models are being made. So you don't always know ahead of time necessarily what are they going to be very good at, what are they maybe going to be bad at.”
Disclosure
Anthropic built Claude Cowork by giving Claude Code a virtual machine
“What we've done is we've taken cloud code and we've given cloud code a virtual machine That Claude can use to run its own code.”
Disclosure
Anthropic currently runs over 100 internal AI application prototypes
“We probably have easily a hundred different prototypes of like various applications inside the company right now.”
Insight
Rieseberg: Developers should avoid over-building custom AI agent infrastructure
“I would probably recommend first not to actually build your, not to build too much of your own infrastructure and use a product that we've launched today called Cloud Managed Agents that make this particular case very useful.”
Disclosure
Rieseberg: Claude Cowork team maintains maximum one-month product roadmap
“We tend to plan not more than a month out.”
Insight
Patel: AI model performance is a lagging indicator of prior hardware CapEx
“Ultimately the capex that Microsoft spent in 2024 for OpenAI is what results in 2025 for OpenAI or CoreWeaver or whoever is what results in their models being so good this year. Same with Anthropic and Amazon Google and their models now being so good now is th…”
Insight
Douglas: AI releases accelerating due to dual-paradigm scaling
“There's now this two paradigm regime where previously you did free training scaling and reinforcement learning scaling, and now we're in a mix of the two basically. And so I think that gives you more opportunities to update models because it means that you can…”
Assertion Supported
Douglas: Anthropic's mid-tier Sonnet is smarter than its flagship Opus
“One of the interesting things about this most recent release is actually Sonnet is smarter than Opus.”
Assertion Not checkable as stated
Cherny: Most Anthropic employees became daily active users of Claude Code
“Pretty soon after we launched it, most of Anthropic was a daily active user”
Assertion Partly supported
Cherny: Claude Code requires no services beyond the API itself
“It doesn't use any services except the API itself. So that's all it needs. And then everything else you actually don't need. And this is one of the nice side effects of not doing code base indexing or anything like this is it's just very easy to hook up to.”
Opinion
Socher: Anthropic's Claude models excel at legal reasoning tasks
“One very concrete example is Anthropics, quite good for legal types of reasoning and like just kind of connected to some of their morals and ethics and so on. And so that those questions are often better routed to a Claude-like model from Anthropics.”
Insight
Real-time AI apps require abstraction layers to handle LLM provider outages
“Yeah, I think you have to create a layer on top of these at a minimum, right in the end, whether Entropic or OpenAI, you know, they're companies, right, they sometimes have stability issues, right, if they're down, so what they're gonna do about that? Since ou…”
Disclosure
Moonhub was used by Anthropic and Inflection as pre-launch clients
“Before we launched, we worked with about a hundred companies helping them hire and scale their teams. So some of our sort of early customers pre-launch were companies like Inflection, Anthropic, Sandbox, U.com public companies out there as well as some early s…”
Assertion Supported
Rieseberg: Claude Mythos is a standalone model outside the Sonnet family
“So for now it's a preview model with its own, in its own category.”
Assertion Supported
Rieseberg: Claude Cowork skills are markdown instruction files
“Skills are essentially just markdown files that explain to the model how to do things.”
Disclosure
Izmailov: Interpretability tools are growing more useful internally at Anthropic
“So we are still pretty far from the dream that we will Fully understand everything that happens in the model, but these tools are becoming increasingly more useful internally at Anthropic in particular, and also there is constant progress, and it's pretty fasc…”
Disclosure
Douglas: Top AI labs give researchers months to explore unproven ideas
“It's one of the things that at both Anthropic and at DeepMind we really tried to build, which is like a culture of safe experimentation where People were trusted to explore ideas for a long time out in the wild. Cause you often need months of independent resea…”
Disclosure
Douglas: Anthropic built persistent memory files into its AI agentic harness
“One of the things that we're pretty happy with about the recent launches, we've finally taught the models to use what's called memory. And so, and we've built that into the agentic harness. So it's able to create a markdown file of to-dos and things that it th…”
Assertion Supported
Douglas: Sonnet 4.5 pushed SWE-bench scores from roughly 72% to 78%
“We moved recently from roughly 72 to roughly 78 in Sweepbench”
Disclosure
Cherny: Anthropic is developing single-binary Claude Code without Node.js dependencies
“We landed native Windows support recently, and we're working on single file distribution, so you don't need a Node.js anymore.”
Disclosure
Cherny: Anthropic only releases Claude Code features used daily internally
“Generally our bar is if we find ourselves really happy with it and we find ourselves using it every day, then we release it to everyone.”
Disclosure
Cherny: Claude Code will feature multi-agent systems with agents managing agents
“Expect a lot more agents. So, you know, be able to launch agents, agents managing agents and a lot more kind of freedom, freedom this way”
Assertion Not checkable as stated
Cherny: Anthropic data scientists, designers, and PMs all use Claude Code
“The data scientists at Anthropic all use quad code to write their queries, and designers use it to build small prototypes, and product managers use it to manage tasks.”
Disclosure
Anthropic released Claude Code to learn how to make AI agents safe in the wild
“One of the reasons that we release cloud code is just to learn how people use it, and to learn how to make this thing safe, to learn how it behaves in the wild, so that we know what to do next.”
Disclosure
Cherny: Claude Code offers subscription pricing from $20 to $200 monthly
“For Quadco, there's two pricing models today. One is you can get a subscription. There's Pro, and then there's Max, and this is I think it's 20 bucks a month, hundred bucks a month, or 200 bucks a month, and it has very, very generous rate limits.”
Disclosure
Levie: Box officially supports Gemini, Anthropic, and OpenAI models
“We officially support Gemini family. We support anthropic, you know, variety of cloud models and then support open AI and the GPTs.”
Disclosure
ZoomInfo selected Anthropic over OpenAI and Google for call summarization
“We are using Anthropic, we are using OpenAI, we are using Google, but for that one, we released Anthropic.”
Disclosure
Sternberg: Notion uses models from both OpenAI and Anthropic
“While those are the two that I can say are, yes, publicly, like, we have leveraged both of those partners”