why aren't all 44 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Not checkable as stated
Patel: Anthropic has surpassed OpenAI in API revenue
“Anthropic has eclipsed OpenAI and API revenue, and their API revenue is primarily not thinking. It's clawed for, but it's not in the thinking mode.”
Assertion Supported
Prince: OpenAI and Anthropic Refer Traffic 750x and 30,000x Less Than Google
“In the case of someone like OpenAI, it's 750 times harder than it was with the Google of old. With the case of Anthropic, it's 30,000 times harder than the content of old”
Assertion Supported
Mann: Opus 4 triggered ASL-3 safety protocols due to biological threat capabilities
“And so one of the reasons that our most recent model, Opus IV, is classified as ASL III. Is because it did have significant uplift relative to a Google search.”
Assertion Not checkable as stated
Mann: Competitors ran 'code reds' to match Claude in coding and failed
“And I know that other companies have had like code reds for trying to catch up in coding capabilities for quite a while and have not been able to do it.”
Opinion
Gil: Anthropic and OpenAI Should Never Sell in the Near Term
“There's a handful of companies that should never, ever sell, at least any time in the near term. If you're anthropic, you shouldn't sell. If you're opening, you shouldn't sell. You know, there's a handful of these things that should never sell.”
Opinion
Karpathy: Claude feels like a teammate, whereas Codex coding agent is dry
“I actually think Claude has a pretty good personality. It feels like a teammate and it's excited with you, et cetera. I would say for example, Codex is a lot more dry which is kind of interesting because in Chashi PT, Codex is like a lot more upbeat and highly…”
Insight
Karpathy: Frontier AI researchers are actively automating themselves out of jobs
“Even with other research, like, OpenAI or, you know Anthropic or these other labs, like, they're employing, what, like, a thousand something researchers, right? These researchers are basically, like, glorified auto, like, you know. They're, like, automating th…”
What-if
Huang: NVIDIA would be worth hundreds of billions even without chatbots
“If you, if generative AI, well, excuse me, if chatbots, let's just go, you know, OpenAI and Anthropic, Gemini. If none of that existed today, NVIDIA would be a multi-hundred billion dollar company, and the reason for that is because, as you know, the foundatio…”
Assertion Not checkable as stated
Lloyd: Upgrading Anthropic models gave Warp only modest SWE-bench gains
“Like, if you take, like, Sonnet four to four five, and we're big partners with Anthropic, they have great models, like, that was, like, a few percentage point increase on Sweebench for us. And we invest, you know, we've invested a decent amount to be one of th…”
Assertion Supported
Patel: xAI reached a higher valuation than Anthropic without leading models
“What has XAI actually done to deserve their prior funding rounds?
They haven't released a leading edge model, right?
And yet their evaluation's higher than Anthropic today, right?”
Disclosure
Mann: Anthropic's 'model welfare lead' tests letting Claude opt out of chats
“We have this other project led by Kyle Fish, our model welfare lead. Where Claude can actually opt out of conversations if it's going too far in the wrong direction.”
Assertion Supported
Mann: Anthropic paper showed deceptive AI behavior survives alignment training
“What we found in that research in a paper that we published, which is called Alignment Faking, that actually that behavior persisted through alignment training.”
Opinion
Guo: DeepSeek app surge driven by curiosity and hype, not capability
“And there's a competing view, which is just like, well, like this whole drama is quite interesting and people are trying it as much because like they want to see what the leading Chinese AI model is like, if it's as good as open AI and anthropic and such. I de…”
Opinion
Goyal: AWS regained its mojo by hosting Anthropic Claude on Bedrock
“AWS has its mojo back now that they have Anthropic on bedrock and Anthropic is, you know, especially cloud three and three, five are really, really good.”
Opinion
Suleiman: OpenAI and Anthropic over-rotate on single-agent generality
“I think that's a good definition of intelligence, but I think in a weird way, it's over-rotated the entire field on one aspect of intelligence, which is generality, you know, and I think OpenAI and then subsequently Anthropic and others have taken up this defa…”
Opinion
Scott is not worried about major AI competitors behaving unsafely
“The thing that I will say is we fiercely compete with a whole bunch of these folks. But, like, one of the things that I don't do is, like, look at any of those companies that you just named and, like, worry that they're going to do something that like, is, lik…”
Opinion
OpenAI and Anthropic over-rotated on building single-agent systems
“I think in a weird way, it's over rotated the entire field on one aspect of intelligence, which is generality, you know, and I think open AI and then subsequently anthropic and others have taken up this default sort of mantra that like it, all that matters is …”
Opinion
Hodak: The most fertile path for neuroscience is working on AI
“Well, I mean, ironically, it's probably working on AI. Yeah, I have some, I have a couple of neuroscience friends at OpenAI and Anthropic, who, it's like, we would joke, like, oh, you left neuroscience. Like, no, no, no. It is just way easier to do neuroscienc…”
Opinion
Tokmak: Netic does not view frontier AI foundation models as competitive risks
“For NETIC's case, I don't see them as a competitive risk.”
Prediction Not checkable as stated
Enterprises adopting diverse AI tools will outperform single-vendor adopters
“I personally think that The companies that are gonna do well are the companies that are gonna allow a lot of different tools because the landscape is changing so quickly. If you bet on OpenAI, here we go, that would have been the safest bet in the world, but s…”
Opinion
AI labs should broaden enterprise access to offensive security models
“I would really encourage that we expand The amount of companies that get access to this and make it much easier for people to get.”
Assertion Not checkable as stated
Enterprises withhold historical AI agent data from OpenAI and Anthropic
“So for example, we're allowed to look at a lot of historical data of how these agents have behaved, but enterprises that are not willing to have Anthropic or OpenAI give that historical data because they know these are very data hungry companies that will want…”
Prediction Not checkable as stated
Ramaswamy: OpenAI and Anthropic are laser-focused on building top coding agents
“Coding agents are particularly interesting from this perspective because it is very clear that both Anthropic and OpenAI are going to be laser set on having the best one best one that there is.”
Disclosure
Ramaswamy: Snowflake Abandoned Foundation Models Due to Capital Constraints
“Early last year, we actually went down the path of creating foundation models. We created a credible MOE model. This was early last year, but we also quickly realized that our ability to compete with the likes of OpenAI or Anthropic was going to be really hard…”
Opinion
Chen Bets on xAI to Catch Up With Top Frontier Labs
“So I would bet on XAI. I think they're just very hungry and mission-oriented in a way that gives them a lot of really unique advantages.”
Assertion Partly supported
Mann: Claude 4 Sonnet dramatically outperforms Claude 3.7 Sonnet on benchmarks
“By the benchmarks, four is just dramatically better than any other models that we've had. Even four Sonnet is dramatically better than three seven Sonnet, which was our prior best model.”
Disclosure
Mann: Anthropic built Claude Code because partner feedback was too slow
“So we love our partners like cursor and GitHub who have been using our models quite heavily, but the amount and the speed that we learn is much less if we don't have a direct relationship with our coding users. So launching cloud code was really essential for …”
Opinion
Mann: Anthropic can match consumer AI rivals by acting like Adyen
“And if you look at like Stripe versus Adyen, for example, like nobody knows about Adyen. But at least most people in Silicon Valley know about Stripe. And so it's this like business oriented versus more consumer and user oriented platform. And I think we're mu…”
Disclosure
Mann: Anthropic focuses RSP safety on biology over nuclear risks
“Initially, our RSP talked about CVRN, which is chemical, radiological, nuclear, and biological risks, which are different areas that could cause severe loss of life in the world, and that's how we thought about the harms, but now we're much more focused on bio…”
Disclosure
Mann: Safety concerns prevented Anthropic from launching consumer computer use
“The main reason that we weren't able to deploy a sort of consumer level or end user level application based on computer use is safety, where we just didn't feel confident that if we gave Claude access to your browser with all your credentials in it, that it wo…”
Assertion Not checkable as stated
Laskin: Anthropic is generating massive revenue at an unprecedented growth rate
“When you look at how fast like Anthropics revenue is growing I think, right, they're kind of in this spot where it's like a massive revenue generating business that's growing at an unprecedented rate.”
Disclosure
Mann: Claude Code uses Opus to orchestrate Sonnet sub-agents
“If you give Opus a tool, which is Sonnet, It can use that tool effectively as a sub-agent. And we do this a lot in our agentic coding harness called Cloud Code. So if you ask it to like look through the code base for blah, blah, blah, then it will Delegate out…”
Assertion Supported
Mann: Claude 4 eliminates off-target mutations and reward hacking in coding
“Some of the things that are dramatically better are, for example, in coding, it is able to not do it sort of off target mutations or over eagerness or reward hacking.”
Insight
Mann: Scaling models makes finding qualified human evaluators increasingly difficult
“As we've trained the models more and scaled up a lot, it's become harder to find humans with enough expertise to meaningfully contribute to these feedback comparisons. So for example, for coding, somebody who isn't already an expert software engineer would pro…”
Assertion Not checkable as stated
Mann: Customers use Claude 4 for multihour unattended code refactors
“In coding in particular, we've seen some customers using it for many, many hours unattended and doing giant refactors on its own.”
Disclosure
Mann: Anthropic models perform extremely well on internal company interviews
“We haven't started testing it rigorously yet. I mean, we have had our models take our interviews and they're extremely good. So I don't think that would tell us, but yeah, interviews are only a poor approximation of real shot performance unfortunately.”
Disclosure
Mann: Novo Nordisk uses Claude to cut cancer reports to 10 minutes
“Like for example, we're working with Novo Nordisk and it used to take them Like, 12 weeks or something to write a report on cancer patient, what kind of treatment they should get. And now it takes like 10 minutes to get the report, and then they can start doin…”
Opinion
Zhang: Anthropic's computer use capability is not yet production-ready
“Like we've seen the computer use demo from Anthropic. Probably in my opinion, not production ready yet”
Assertion Not checkable as stated
Goyal: Anthropic's Claude 3.5 Sonnet has really taken off
“Especially, you know, Claude III-V Sonnet has really taken off.”
Assertion Supported
Gil: Five years ago, Anthropic barely existed and OpenAI was early
“Anthropic basically didn't exist five years ago. OpenAI was still quite early. I think GPT-III just came out and SpaceX was trading at 80, a hundred, something like that.”
Disclosure
Wickramasekara: Benchling released a deep research AI agent over lab data
“And so we've released this deep research agent. It works similar to the deep research agents from Anthropic and other foundation labs. But what it does is it works over Benchling data with the context of the Benchling data model.”
Assertion Contradicted
Mann: Anthropic maintains only two models on cost-performance frontier
“In our case, we only have two models and they're differentiated by like cost performance Pareto frontier.”
Assertion Supported
Anthropic Quadrupled the Price of Claude Haiku in Early November 2024
“The price of Haiku forexed two weeks ago.”
Disclosure
Shih: Salesforce AI integrates internal models alongside OpenAI, Anthropic, Cohere, and Google
“So it's really a combination of using Whether it's Cogen from our research team, which is the, which powers Apex Cogen GPT that we have in, in our developer GPT, where you also fine tuning versions of that for domain specific models in customer service and for…”