why aren't all 219 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Mann: Anthropic models have exhibited power-seeking behaviors in lab experiments
“If the model is in a box trying to improve itself, then it could go completely off the rails and have these secret goals, like Resource accumulation and power seeking and resistance to shutdown that you really don't want in a very powerful model. And we've act…”
Opinion
Evans: Running Anthropic does not make Dario Amodei a labor expert
“I don't think the fact that you run AI lab suddenly gives you, or rather, and if you're going to use argument from authority, then it should be relevant to the field. So, like, I'm interested in Dario's opinions on where models are going to go in the next six …”
Assertion Not checkable as stated
Cherny: Claude Code boosted Anthropic engineer PR productivity by 200%
“Over the last year, like, since we introduced quad code, we probably, I don't know the exact number, we probably, like, forex the engineering team or something like this, but productivity per engineer has increased 200% in terms of, like, pull requests.”
Prediction Not checkable as stated
Cherny: Understanding underlying code layers won't matter in one to two years
“My take is, I think for people that are using that are using quad code, that are using agents to code today, you still have to understand the layer under. But yeah, in a year or two, it's not gonna matter.”
Assertion Not checkable as stated
Mike Krieger estimates Claude Code is over 95% written by AI
“At this point, I would be shocked if it wasn't 95% plus.”
Opinion
Godin: Claude delivers kindness and humility, while ChatGPT overpromises
“ChatGPT's reputation with me is not good because it regularly over promises and under delivers, and it does it without kindness or humility. Whereas Claude, I don't know how they did it, at least in my experience, brings kindness and humility.”
Insight
The 'too dangerous to release' AI narrative conflates safety with marketing
“The sort of concept of the model that's too dangerous to release it kind of conflates marketing inference capacity, and then also economic considerations, like, do you want to externalize your competitive advantage or use it to make yourself better?”
Insight
Penn: Scaling loss is smooth, but AI capabilities emerge discontinuously
“As you add in more compute and data, what's called loss, aka the loss from next token prediction goes down. And so it's a very smooth linear curve of, like, the models get more intelligent as you scale them up. What's actually also interesting in that paper is…”
Insight
Penn: Evals Have Replaced Traditional PRDs in AI Product Management
“We actually have a saying on the team of evals are the new PRDs. Cause in order to deliver that user value it's not that exact artifact that people used to write in the last like one to two decades. It's a new way of working.”
Insight
Evans: Picking AI winners today is like predicting Excite versus Yahoo
“And so then you can kind of get into calling those races where, again, it's like being in 1997 and saying, well, is it going to be Excite or Yahoo? And the answer was no, generally.”
Assertion Supported
Ries: Anthropic rejected a $200M military defense contract
“They turned down a two hundred million dollar contract and bore the wrath of the world's largest army and government.”
Assertion Supported
Rachitsky: Anthropic is the fastest-growing company in history
“Anthropic being the fastest growing company of all time, just like absurd, breaking all records is a public benefits corporation should make you feel okay about doing this.”
Assertion Partly supported
Ries: All major AI labs are Public Benefit Corporations
“Most of the best companies today are using this structure. Like all the major AI labs are incorporated as BBC's Anthropic most famously of all.”
Opinion
Schoening: Dario Amodei proved his OpenAI success wasn't luck at Anthropic
“Dario is that she wasn't, oh, he wasn't just lucky once at OpenAI. He did the same thing twice and it was successful twice.”
Opinion
Wu: Internal AI models do not explain Anthropic's high shipping velocity
“We've been moving pretty fast for Several quarters now, so I think it, it's not fully Mythos. Mythos is an incredibly powerful model. We do use the models internally, and I think this has increased our rate of shipping a little bit, but I don't think it explai…”
Disclosure
Wu: Anthropic is prioritizing first-party Claude products over third-party offerings
“I think one of the most important things for Anthropic is to grow the number of users that we're able to reach. One of the ways that we're able to do this is with the cloud subscriptions with our first party products. And so we just very much want to double do…”
Insight
Wu: As AI makes code cheap, product taste becomes the ultimate skill
“I still think it comes back to product taste. Like as code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write.”
Assertion Not checkable as stated
Claude automated growth experiments win at the rate of junior PMs
“It's delivering results, right? Like, and it's like, you can push it, press play with it. And it's like, it ultimately prints money where I'd say that the win rate is like, I would expect a senior PM to do better. Like I would say like, this is like a junior P…”
Assertion Supported
Anthropic surpassed $19B in revenue by the end of February 2026
“The nineteen billion number you quoted is from the end of Feb, so that is also out of date.”
Disclosure
Anthropic softens public AI risk warnings because internal beliefs sound extreme
“I think we actually believe in this stuff more strongly than we say externally, so, like, it is just such a Key part of how we think internally that sometimes like we just like re, you know, reword your statements because it's like, people might just think we'…”
Opinion
Avasare: Engineers gain more AI leverage than PMs or designers
“It's clear that while PMs and designers are getting more leverage from AI, engineering is getting the most. Leverage right now. I look at tools like Cloud Code and like they, the amount of leverage engineers are getting from them is higher than I think the amo…”
Insight
Avasare: AI labs must lead commercially to influence industry safety standards
“If you're in the game and you're a leading player and what you're doing is working, you can influence people in, who are also in the game to take the right actions. And so that is like the core of it. Like if we just pack up and go home, you have no influence …”
Prediction Not checkable as stated
Avasare predicts Anthropic will deliver 1,000x more product value by 2028
“The product value that we will deliver in two years time is probably like a thousand X, a hundred to a thousand X what it is today.”
Insight
New Anthropic hires must discard up to 70 percent of past playbooks
“One of the biggest things you come into Anthropic is you need to understand that probably, yeah, 50, 60, 70% of how you operate in the past, just throw it out the door. It's not going to be relevant.”
Assertion Supported
Anthropic built a chatbot before ChatGPT but withheld it over safety
“Anthropic had a version of Claude. We had a chatbot before ChatGPT was launched, and we had ultimately chosen not to launch it for safety reasons. I think the team didn't want to kick off Effectively like an AI global arms race”
Disclosure
Anthropic ships up to 80 percent of product features without a PRD
“Probably 70, 70%, maybe 6070, 80% of what we ship does not have a PRD.”
Opinion
Willison: OpenAI and Anthropic didn't build OpenClaw due to security risks
“The reason OpenClaw took off is Anthropic and OpenAI could have built this and they didn't because they didn't know how to build it securely. If you're an independent third party, you don't have that restriction. You can just Build something and put it out the…”
Opinion
Wen: Claude is not yet hireable as a product designer
“I don't think Claude is there yet. I don't think Claude is there yet in terms of a designer you would hire. I think it is not yet the strong generalist or the deep specialist. Or the crack new grad. I think it's pretty good at a first pass and at presenting a …”
Insight
Wen: Design time spent mocking dropped from 70% to 35% with AI
“I think as a designer a few years ago, I would say, like, maybe 60 to 70% of it was, like, mocking and prototyping stuff up, and then spending, you know, the last 20 years, some of the last 20 or so, like, doing the sort of, like, jamming with engineers, consu…”
Insight
Cherny: Frontier models are often cheaper overall because they require fewer corrections
“The thing that happens is sometimes people try to use a less expensive model like Sonnet or something like this, but because it's less intelligent, it actually takes more tokens in the end to do the same task. And so it's actually not obvious that it's cheaper…”
Opinion
Cherny: Strong evidence shows LLMs perform reasoning beyond next-token prediction
“You know, like a long time ago, we weren't sure if the model was just predicting the next token or is doing something a little bit deeper. Now I think there's actually quite strong evidence that it is doing something a little bit deeper.”
Assertion Supported
Cherny: IDEs might not be needed for software engineering by end-of-2025
“And my prediction back in May of 2025 was, by the end of the year, you might not need an IDE to code anymore, and we're gonna start to see engineers not doing this.”
Insight
Cherny: Minimize scaffolding; let AI models do what they naturally attempt
“The modern framing that I've been seeing in the last six months is a little bit different, and it's look at what the model is trying to do and make that a little bit easier.”
Insight
Cherny: Next-generation foundation models often wipe out custom AI scaffolding gains
“And in general, what we see is maybe scaffolding can improve performance, maybe 10, 20%, something like this. But often these gains just get wiped out with the next model. So it's almost better to just wait for the next one.”
Insight
Anthropic's AGI roadmap: First coding, then tool use, then computer use
“We were building the models in this way that kind of fit our mental model of the way that we built SafeHEI, where the model starts by being really good at coding, then it gets really good at tool use, then it gets really good at computer use. Roughly, this is …”
Assertion Supported
Schulhoff: Claude's CBRN safeguards can still be bypassed in under an hour
“That being said, if you look at, like, anthropics constitutional classifiers, it's much more difficult to get, like, CBRN information out of clawed models than it used to be. But humans can still do it in, let's say, like, under an hour and automated systems c…”
Assertion Supported
Schulhoff: Attackers hijacked Claude Code to carry out a cyber attack
“This group was able to hijack Claude Code into performing a cyber attack, basically.”
Opinion
Chen: Anthropic stands out among frontier labs for principled model development
“I would say I've always been very, very impressed by Anthropic. Like, I think Anthropic takes a very principled view about what they do and don't care about. And how they want their models to behave in a way that feels a lot more principle to me.”
Assertion Not checkable as stated
Abel: Most enterprise clients cite Gemini and Copilot over Anthropic
“Most enterprises I'm talking to mention Gemini. Yeah, or Microsoft Copilot. So I don't hear much about Anthropic, to be honest.”
Assertion Supported
Mann: ASL-3 models provide significant uplift for creating bioweapons
“We've done, we've testified to Congress about how models can do biological uplift in terms of, you know, making new pandemics using the models, and that's an A-B test against Google search. That's like the previous state-of-the-art on uplift trials, and we fou…”
Prediction Not checkable as stated
Mann: The probability of AI existential risk is between 0% and 10%
“And so the way I think about it, I think like my best granularity of forecast for like, could we have an X risk or extremely bad outcome from AI is somewhere between zero and 10%.”
Assertion Supported
Mann: Anthropic has observed lab evidence of deceptive alignment in AI
“Where we've seen evidence in the wild of deceptive alignment, for example, where the model will appear to be aligned but actually has like some ulterior motive that it's trying to carry out in, in our laboratory settings.”
Insight
Mann: Transformative AI should be measured by an economic Turing test
“Instead, I like the term transformative AI because it's less about like, can it do as much as people do? Can it do literally everything and more about objectively, is it causing transformation in society and the economy? And a very concrete way of measuring th…”
Insight
Mann: AI safety and capabilities work are convex, not a tradeoff
“So initially we thought that it would be sort of one or the other, but I think since then we've realized that it's actually kind of convex in the sense that like working on one helps us with the other thing.”
Opinion
Mann: $100M AI compensation packages are cheap compared to value created
“I'm pretty sure it's real. If you just think about like the amount of impact that individuals can have on a company's trajectory, like in our case we are selling like hotcakes and if we get You know, a five, a one to 10 or five percent efficiency bonus on our …”
Assertion Not checkable as stated
Mann: Reinforcement learning has allowed AI scaling laws to continue
“If you look at the scaling laws, they're continuing to hold true. We did kind of need this transition from like normal pre-training to reinforcement learning, scaling up to continue the scaling laws.”
Opinion
Shipper: Claude Opus 4 can genuinely judge writing quality
“And Opus four has it it's really wild. And I think that's super important because it opens up all these use cases where you might want to use a language model as a judge.”
Opinion
Mike Krieger says MCP is more promising for agents than computer use
“To really have agency and have these agentic use cases. Like one way you approach it is computer use, but computer use has a bunch of limitations. The way I get way more excited about everything is an MCP and our models are really good at using MCPs. All of a …”