why aren't all 235 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 3 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Hershey: Anthropic Study Found Claude Treats Named Characters Better
“Anthropic actually did like a blinded study of like named characters versus unnamed characters in different settings, and Claude like actually does clearly prefer and is nicer to named characters, which is an interesting thing.”
Opinion
Swyx: Frontier Models Exist Primarily to Distill Smaller, Usable Models
“Even GPT 4.5 is too expensive. Normally it's really gonna use it in, in any reasonable quantity. Like, you know, Claude 3.5 Opus, like if it does exist, still not like, you know, the thing that we actually use is Sonnet, right? So like, it's almost like a depl…”
Opinion
Snipd CEO: Claude is the best model at phrasing and personality
“Like, in my opinion, Claude is the best one when it comes to the way it formulates things.”
Disclosure
Nguyen: Claude 2's distinct personality was unintentional until Claude 3
“People said, like, Cloud II is, like, so much better at, like, writing and, like, has a certain personality, even though it was, like, unintentional at all. And we did not pay that much attention and didn't know even how to, like, productionize this property o…”
Opinion
Swix: Bearish on computer-use AI agents due to cost, speed, and accuracy
“I have been very bearish in computer use because they're slow. They're expensive. They're imprecise. Like the accuracy is horrible. Still, even with Anthropix new stuff, I'm really waiting to see what opening I might do to change my opinions.”
Assertion Supported
Swyx: Claude wrapper Bolt.new reached $20M ARR
“The other one would be Bolt. There's a straight quad wrapper. And again, another now they've announced twenty million ARR, which is another step up from our eight million that we put on the title.”
Opinion
Neubig: Claude Is The Best Agent Model, Open Models Lag Behind
“I still am under the impression that Claude is the best. The other closed models are, you know, not quite as good, and then the open models are a little bit behind that.”
Opinion
Neubig: GPT Loops On Errors While Claude Tries New Approaches
“So, like, GPT doesn't have very good air recovery ability. And so, because of this, it will go into loops and do the same thing over and over and over again, whereas Claude does not do this.”
Assertion Supported
Reddy: OpenAI's share of enterprise LLM spend dropped from 90% to 60%
“And the opening I spend at the beginning, at the end of last year in November of 23 was close to 90% of total volume. And today, less than a year later, it's closer to 60% of total volume.”
Assertion Not checkable as stated
Schluntz: String replacement is the most reliable file-editing tool for LLMs
“We did a few different experiments with like different ways to specify how to edit a file and string replace. Basically the model has to write out the existing version of the string and then a new version, and that just gets swapped in. We found that to be the…”
Prediction Not checkable as stated
Schluntz: Production AI agent applications will be bespoke, not off-the-shelf
“You know, I think that might be useful for hobbyists and demos, but the ultimate end applications are going to be bespoke. And so we just want to make sure that the model's great at any tool that it uses”
Disclosure
Anthropic: Tool engineering mattered more than prompt engineering for SWE-bench
“I would say actually we did more engineering of the tools than the overall prompt.”
Insight
Crivello: Tasks capable of using APIs must stay API-driven over computer use
“My philosophy about it is anything that can be done with an API must be done by an API or should be done by an API for a very long time.”
Assertion Supported
Polu: Claude Sonnet executes an unpublicized chain-of-thought step during function calling
“They kind of innovated in an interesting way, which was never quite publicized, but it's that they have that kind of chain of thoughts step whenever you use a Clouds model or Sonnet model with function calling. That chain of service step doesn't exist when you…”
Opinion
Liu: Claude 3 Haiku outperforms OpenAI models at function calling
“Overall, I'm like super happy with the anthropic models compared to the OpenAM models. Like, Sonnet is very cost effective. Haiku is, in function calling, it's actually better.”
Insight
Lambert: Claude's constitution dictates output priorities, not model beliefs
“If you look at Claude's constitution, like, that doesn't mean the model believes these things. It's just trying Trained and to prioritize these things.”
Assertion Supported
Patel: Google, OpenAI, and Anthropic are developing multi-datacenter training
“One of the big bottlenecks is how much power and how many chips you can get into a single data center. So, like, A, Google and OpenAI and Anthropic are working on this, right?”
Assertion Not checkable as stated
Swyx: Top AI agent labs receive secret discounts from model providers
“Agent Labs get discounts from every model provider, and that's also very interesting when people compare public pricing of, like, a discounted cloud code from Anthopic versus what Anthopic does with Model Labs, with Agent Labs”
Prediction Not checkable as stated
Swyx: Frontier AI labs will not provide bespoke enterprise integration support
“The labs do not have 200 people dedicated to like, you know, being on call with you with Goldman Sachs going like, okay guys, what do you need? We got it. You need the Microsoft Teams zero integration. Got it. You don't use GitHub. You use this like weird org …”
Disclosure
Midha: AMP Foundry invested hundreds of millions into Anthropic this year
“We put a few hundred million dollars into Anthropic from our fund earlier this year.”
Opinion
Midha: Anthropic's velocity came from standardizing on the transformer architecture
“Like, one of the reasons Anthropic has had extraordinary sort of velocity is because they picked the transform architecture and said, this is simple, let's double down on it, right? And now, luckily, there's enough investment going into space that we can affor…”
Assertion Not checkable as stated
Awais: Claude tolerates tool errors and self-corrects, unlike open models
“Claude is actually really, really lenient for tool calls. So even if, you know, your coding agent harness messes up, it can figure out that, oh, I'm being sent this error and can fix itself. Not the case with you know open models”
Disclosure
Chatbase customer model usage is split 50% OpenAI, 50% Anthropic and Google
“Maybe 50% is still on Okunai. Yeah. Yeah. And then 50 on everything else. Yeah. But everything else is like mainly Anthropic and Google.”
Assertion Not checkable as stated
Claude Sonnet Crashed on Duplicate Tool Names While OpenAI Handled the Error
“Sonic couldn't handle two tools with the same name in OpenAI, GPT, 5.2. It was like, ah, I can figure this out. So that was an interesting one that we learned by accident through a SEV.”
Insight
Rieseberg: Prompt Opus by stating goals, not specifying exact execution steps
“Honestly though, like I see that you're using Opus 4.6, right? Like my recommendation for people is increasingly don't worry about it anymore. Just like tell it what you want it to do. And it's probably going to figure out a way to do it.”
Insight
Rieseberg: Cowork's most impressed users find unexpected capabilities
“Every single person who's like most amazed is usually amazed about a thing that I didn't even expect Cowork would be good at.”
Assertion Supported
Rieseberg: Claude Cowork is Claude Code running in a sandboxed virtual machine
“Cowork is cloud code running in a virtual machine with a little bit of padding, a little bit more guardrails, making it a little safer, a little bit more convenient for people who don't want to first open up the terminal when they go to work.”
Insight
Rieseberg: Claude Cowork skills can be as simple as a text message
“One thing that is very fun for me about skills in particular is that they're so easy to make. Like anyone can make a skill, like a text message could be a skill and they can be so hyper-personalized to you.”
Insight
Rieseberg: A labs team should only tackle ideas no one else would
“The sort of the idea of a Labs team is that it should only work on things that make really no sense for anyone else to work on.”
Assertion Not checkable as stated
Eskildsen: Anthropic, Notion, and Cursor use Turbopuffer across three deployment models
“You can run Turbo Puffer either in SAS, right? That's what cursor does. You can run it in a single tenant cluster. So it's just you. That's what Notion does. And then you can run it in, in, in BYOC where everything is inside the customer's VPC. That's what, fo…”
Assertion Not checkable as stated
Wang: Anthropic Claude Cowork automates complex customer cohort data analysis
“Like for the first time you can actually get one shot data analysis, right? Which, you know, if you're going to do a customer database, analyze a cohort retention, right? That's just stuff that you had to do by hand before. And our team, the other, it was like…”
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Disclosure
Anthropic remains fully committed to MCP following its foundation donation
“Like the commitment of Anthropic is the same, right? I'm still, We still have the same people I'm helping with the SDKs. We're still super committed in our products to MCP. I'm still the lead core maintainer. Nothing has actually changed.”
Disclosure
Anthropic created MCP so rapidly expanding internal teams could build integrations independently
“MCP before we even open source it was born of the idea of like, I'm in a company that is growing crazy. I'm in the development side of things, development tooling side of things. I will grow slower than the rest. How can I build something that they can all bui…”
Assertion Supported
Block's Goose was the first open-source agent to integrate MCP
“Goose was the first open source agent interface or agent that reached out to us and worked with us to integrate MCP. And I think Rad is actually like technically the first non-anthropic contributor to MCP ever on like day two or something like that, like very,…”
Assertion Supported
Google, Microsoft, Amazon, OpenAI, and Anthropic joined AAIF as platinum members
“You have Google, Microsoft, Amazon Block, Bloomberg, Cloudflare, OpenAI, Anthropic. Just a platinum member, create a foundation.”
Assertion Not checkable as stated
Yegge: Anthropic is hiring over 100 people for Claude Code
“They're hiring like a hundred plus people for cloud code in the next, I don't know, month. I mean, like they're going wild and that's just cloud code.”
Disclosure
Superhuman builds dynamic on-the-fly aggregation lambdas with Anthropic
“We're working right now with Anthropic to basically do kind of like a building on the fly, small, kind of a key component of lambdas that will build the code to do the aggregation.”
Assertion Supported
Anthropic Maintains an 80% One-Year Employee Retention Rate
“I'm referring to the exact same article where I think their retention, one year retention on employees is the 80%, which in AI world is, is quite wild.”
Insight
Enterprise LLM Churn Is Low Due to Long-Term Compute Commitments
“In terms of enterprises, often what will happen is they'll buy up large chunks of long-term compute and dedicated instances, in which case you just don't churn, right? Like this is what you use.”
Prediction Open · timeframe Nov 2030
Anthropic and OpenAI will never open-source their high-performance inference kernels
“The high performance inference kernels that sort of drive a lot of, you know, anthropic and open AI and stuff, their models, those aren't open source. They're not going to be open source.”
Assertion Supported
Merrill: Dario Amodei highlighted Terminal-Bench on the Claude model card
“I think one of the really key moments for us was getting onto the Claude IV model card. Being one of two benchmarks that Dario actually mentioned while releasing the model.”
Insight
Krieger: AI Agents Must Support Both MCP and Visual Computer Use
“And that thing's never gonna have an MCP around it. Like, it's just like, who knows if the company created is even around much less like ready to sort of expose their kind of underlying constructs as API. So I think you will need to be able to do both.”
Opinion
Krieger: Claude Sonnet 4.5 Outperforms Opus at Generating 3D Games
“This is like officially good. It's like better than Opus at this. It's like, It generated this, like, great split-screen stereoscopic thing, three-dimensional, like, thing.”
Assertion Supported
Martin: Claude Code operates entirely without codebase indexing
“Clock code doesn't do any indexing. It's just doing, quote unquote, agentic retrieval, just using simple tool calls, for example, using grep, to kind of poke around your files, no indexing whatsoever, and obviously works extremely well.”
Assertion Partly supported
Chroma research finds Claude models lead in long-context utilization
“And you know, one thing Chroma released this context rod paper recently about context utilization and the cloud models are actually the best at using kind of like longer context.”
Opinion
McCloy: Claude users represent an exceptionally valuable audience for companies
“Claude, which is important for, you know, not necessarily huge in terms of raw number of users, but the people who do use Claude tend to be like a very valuable audience, especially for some types of company.”
Assertion Supported
Anthropic finds a single LLM judge outperforms five specialized judges
“They initially started with five LLM as judges. So each one of these points had their own LLM as a judge. They tested the ability and accuracy of that LLM as judge collective to judge, and it actually didn't perform as well as one. So they replaced all of thos…”