why aren't all 235 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 3 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
What-if
Parakhin: Liquid AI could beat frontier models with equal compute
“I think if they if they had similar level of compute, they would be very competitive and maybe even beat the largest models, at least from what I've seen.”
Assertion Not checkable as stated
Patel: Anthropic added $2B in monthly revenue at positive margins
“Anthropic doesn't just add two billion dollars of revenue in one month. You know, with, without having, you know, huge demand and they're doing it at positive margins, right?”
Insight
Casado: Frontier AI Labs Could Outspend and Consume Application Layer Startups
“It literally becomes an issue of like raise capital, turn that directly into growth, use that to raise three times more. And if you can keep doing that, you literally can outspend any company that's built. Not any company. You can outspend the aggregate of com…”
Opinion
Ubl: Anthropic's Boris Power Vibe-Codes From a Position of Privilege
“I think he comes from a particular position of extreme unusual privilege, which is that he works at an AI lab where like people in the office next door are like writing the evals and are like training the model like every day in exactly that way.”
Prediction Not checkable as stated
Slack: Top AI labs face a major customer stampede within two months
“I think we are one or two months away from a possible news cycle. That is the foundation model companies have spent billions of dollars in capex and hired like crazy. And now, you know, they're no longer the best in this realm and there's a huge stampede away …”
Opinion
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Opinion
Neubig: Anthropic's MCP Duplicates Existing APIs With Little Added Value
“We already have an API for GitHub. So why do we need an MCP for GitHub, right? You know, like GitHub has an API. The GitHub API is evolving. We can look up the GitHub API documentation. So it seems like kind of duplicated a little bit. And also they have a set…”
Insight
Anthropic's Schluntz: Avoid agent frameworks and start from scratch with raw prompts
“I think with agent frameworks in general, they can certainly save you some like boilerplate, but I think there's actually this like downside of making agents too easy, where you end up very quickly, like building a much more complex system than you need. And s…”
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And
Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Assertion Supported
Kant: Major AI labs did not prioritize RL for LLMs three years ago
“And the second was that reinforcement learning was going to be the biggest driver for LLM capabilities. Today, very obvious three years ago was not an opinion held or direction held at either OpenAI or Google or Anthropic or others.”
Prediction Open · timeframe Jun 2029
Anthropic will become a trillion-dollar company within four years of founding
“Have you met Dario? Dario's a scientist. He's gone from zero to like what will soon be a trillion dollar company in four years.”
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Assertion Not checkable as stated
Open-source AI demand spiked on hype before reverting to frontier labs
“Like all the open source models, I think what happened was they got like very hyped and people were very interested in using them. But I think like over time, like there was a spike in usage for these models. And then it goes back to open AI, Anthropic and Goo…”
Prediction Not checkable as stated
Rieseberg: AI takeoff will create an accelerating, self-reinforcing loop
“Big bang moment where things will accelerate so quickly that it becomes a self-reinforcing loop. And at that point it's sort of like off to the races and there will be no more like slowly catching up. You know, just have Claude being so good at everything.”
Disclosure
Rieseberg: Anthropic builds all prototype candidates quickly instead of writing memos
“We internally at Anthropica are now probably much closer to the point where, like, don't even write a memo. Just, like, build, like, let's build all the candidates very quickly. Let's just build all of them and then pick the best ones.”
Disclosure
Anthropic deeply worries about AI automating entry-level junior jobs
“At Anthropic as a group of people, we're deeply worried about the impact that the tools are going to have on the labor market, especially for like junior employees that, because I think it's only honest to say that when we talk about automating a lot away, a l…”
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Opinion
Nair: Frontier AI labs have converged on similar reinforcement learning methods
“Well, it does seem like basically a lot of the labs have kind of like converged onto some similar-ish way of doing RL, and they're all kind of back at the same level of like Frontier again”
Opinion
Yegge: Google, Anthropic, and OpenAI are unbelievably chaotic internally
“All three of those companies, Google, Anthropic, and OpenAI are unbelievably chaotic internally right now.”
Assertion Supported
Pliny: Anthropic added a $20k–$30k bounty but withheld jailbreak data
“That whole thing ended with no open sourcing of data, but they did add a 30,000 or 20,000 dollar bounty, which I sort of sat myself out of, let the community go for it.”
Insight
Model Layer Is Structurally More Defensible Than the AI App Layer
“It is far easier for Anthropic to try to go into one of the spaces of the apps than an app to try to go into the space of Anthropic, which makes me feel like one is more defensible than the other all else equal.”
Prediction Not checkable as stated
AI Labs Will Not Dedicate Engineering Talent to Deep Enterprise Search
“If you really want to go deep, I don't think you will ever dedicate the people to do it. And the last thing I'll say is you think about from an anthropic engineer's perspective, you joined a big AI lab to work on models, not to build Google drive connectors, r…”
Insight
AI Coding Intelligence Has an Uncapped Frontier That Drives Revenue
“But the interesting about Anthropic is if you look at coding, that's probably never going to be the case. Like there's always an increasing frontier of how you good you could be at a task like that. And we're nowhere close to that frontier. So it's more possib…”
Assertion Partly supported
Anthropic Is the Fastest-Growing Software Company in History
“Anthropic is the fastest growing software company of all time. I think I can say that fairly. I'm, I haven't been disproven yet.”
Opinion
Dwivedi: Claude is superior at agentic tool calling and error unstacking
“For some of the agentic part of the stack, we are shifting towards Anthropic because they're agentic and the tool calling, especially the unstacking part, you know, when you go down the wrong path and you build context that forces you to keep going down the wr…”
Insight
Ball: Custom MCP tools fail if workflows diverge from frontier training
“If I give it this other custom-made MCP that we built internally, and our processes don't map to anything that OpenAI and Anthropic have seen or trained for, it won't be used, and you won't get good results.”
Disclosure
Ball: The Sourcegraph Amp team does not use formal evals
“I think we don't have any set evals. We don't. And this was controversial up until a week ago, I think, when I think Boris from
Or two weeks ago from Anthropix that they don't have evals for the coding agent too.
But we don't, and we haven't had them.”
Assertion Not checkable as stated
Fanelli: Cursor Default Model Switch Cost Anthropic $200M in ARR
“When cursors switch from Sonnet to GPT-V as like the default model that was like, you know, Two hundred million our revenue for Anthropic that kind of went away and like moved on to GPT-V.”
Assertion Supported
Rajpal: Anthropic Claude models had regressions from serving architecture changes
“Anthropix kind of cloud models kind of had a regression, right? Because they changed to a new serving architecture.”
Assertion Supported
Palazzolo: Claude Code leads stayed at Cursor only two weeks
“We know that they went there, they were there for, I think, about two weeks, and they came back.”
Prediction Not checkable as stated
Dax Reed predicts Claude's $200/month pricing is an unsustainable growth strategy.
“I think it's pretty easy to, like, use more than 200 dollars worth, even by accident. So I would also lean towards that the Claude Max plans are a growth strategy, not like any long term pricing thing that can work, at least at the current, given the current s…”
Prediction Not checkable as stated
Dax Reed predicts OpenCode will dominate when a competitor beats Claude Sonnet.
“What would change things is if there's a day where either another LM lab or like, you know, an open source model drops that is competitive with Sonnet, maybe even better than Sonnet on that day, open code is going to be the only way to do this kind of thing. C…”
Disclosure
OpenCode exactly replicated Claude Code by dumping its system prompts and schemas.
“We look at cloud code. We dump all their system prompts. We dump all their tool descriptions. We dump all their tool schemas. We re-implement the tools. And when you're using an anthropic model, we basically have the exact same implementation.”
Insight
Lambert: OpenAI's Model Spec is more useful than Anthropic's Constitution
“The model spec is much more useful than a constitution because the constitution is like an intermediate training artifact that you give to the training algorithm in order to get the model that you want. It is not necessarily like what model did we, like we don…”
Assertion Not checkable as stated
Zach Lloyd: Cursor represents a very significant portion of Anthropic's revenue
“Cursor I think is some very significant portion of their revenue.”
Insight
Ameisen: LLMs plan future tokens rather than operating purely myopically
“Language models are next token predictors is like a fact. Like that is what they do. They are trained to predict the next token. However, that does not mean that they myopically only consider the next token When they choose the next token, you can work on brea…”
Opinion
Ameisen: Current LLM Chain of Thought Is Unfaithful and Untrustworthy
“So I think there's like a sense in which right now the chain of thought is, is unfaithful, or at least you can't read the chain of thought and trust that that's how the model did it.”
Prediction Held up
Ameisen: Deceptive Backward Reasoning Exists in Base Pre-Trained Models
“I bet, I don't know how much I bet a hundred bucks. So somebody can like, they would get a hundred bucks from me if they prove that I'm wrong, that this behavior for a model that does a drink fine tuning, it also does it post pre-training.”
Disclosure
Anthropic avoided building an IDE to target future AI model scaling
“If we want some product that has like very broad a product market fit today, we would build, you know, a cursor or a Windsurf or something like this. Like these are awesome products that so many people use every day. I use them. That's not the product that we …”
Assertion Supported
Cherny: Anthropic is currently bordering on AI Safety Level 3 capabilities
“Yeah, we're kind of bordering on three right now.”
Opinion
Cherny: Claude Code delivers up to 10x productivity gains for Anthropic engineers
“Anecdotally for me, it's probably two X my productivity. So I'm just like, I'm an engineer that codes all day, every day. For me, it's probably two X. Yeah. I think there's some engineers at Anthropic where It's probably 10 X their productivity”
Assertion Not checkable as stated
Swix: Claude 3 degraded in capability a month after launch
“I used the same project to do this, to try to repeat the demo that I made for myself a month afterwards, and it wasn't anywhere as smart. So cloud three got dumber, but it looks like I like made up the demo or something, but no, like literally I just reran the…”
Assertion Not checkable as stated
Conrad: OpenRouter open-source traffic required only around 10 H100 nodes
“The entirety of Open Router that was not Anthropic or Google like, or Gemini or OpenAI or something. It was like, 10 H 100 nodes or something like that. It's just, like, not that much. It's like, not that many GPUs, actually, to service that entire demand.”
Insight
Schluntz: Smarter AI models require less agent scaffolding
“And I think like the smarter the models are, the less you need that kind of extra scaffolding.”
Opinion
Polu: Anthropic split was driven by disagreement over OpenAI's API commercialization
“What I understood of it is that there was a disagreement of the commercialization of that technology. I think the focal point of the disagreement was the fact that we started working on the API and wanted to make those models available through an API. Is that …”
Assertion Not checkable as stated
Goyal: OpenAI dominates production while Anthropic Sonnet leads side projects
“We still see an overwhelming majority of customers using OpenAI, but almost everyone is using Anthropic for the, and Sonnet specifically for their side projects, whether it's
You know, via cursor or prototypes or whatever.”
Opinion
Claude 3 can bypass its assistant persona to expose the underlying simulator
“Instead of having this entity, like GPT-IV, that's an assistant that just pops up in your face that you have to kind of, like, punch your way through and continue to have to deal with as a headache, instead, there's ways to kindly coax Claude into having the a…”