why aren't all 91 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 3 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Contradicted
Cherny: Anthropic Can No Longer Demonstrate Prompt Injection on Opus 5
“So essentially, if you combine a well-aligned model, so this is, like, essentially three years of research into alignment, With a prompt injection classifier, which we run for all traffic, and what this is doing is it's based on Chrysola's mechanistic interpre…”
Assertion Not checkable as stated
Cherny: Claude routines do the maintenance work of hundreds of engineers
“And so now we have every day, maybe 20 or 30 of these routines. It's running across all of our code bases and It's not totally there yet, but we're on the path to fully automating the maintenance of our apps by doing this. And this is, again, hundreds of agent…”
Opinion
Garry Tan: Claude in Chrome MCP is among worst software ever used
“Claude in Chrome MCP is one of the worst pieces of software I've ever used. You know, every time it would try to do an action, it would think and think and think. There was crazy context bloat. Often it wouldn't even do anything, but it would take two to three…”
Assertion Not checkable as stated
Claude Code increased Anthropic engineer productivity by 150 percent
“And since Quad Code came out, productivity per engineer at Anthropic has grown a 150%.”
Assertion Not checkable as stated
Claude generates between 70 and 90 percent of Anthropic's codebase
“If you look at Anthropic overall, it ranges between, like, 70 to 90% you know, depending on the team. For a lot of teams, it's also, like, a hundred percent. For a lot of people, it's a hundred percent.”
Assertion Partly supported
Srinivas: Google possessed only fourth or fifth best AI models in 2023-2024
“And 2024, 20 23 especially, and large part of 2024 too, Google had like, maybe a fourth or fifth best models at any moment. So as a startup outside Google, you had access to AI that was better than what Google internally had.”
Disclosure
Cherny: Anthropic uses Claude routines to maintain its own apps via Slack
“We actually have Cloud maintaining itself now. And the way we do this is we have a Slack channel where we just had Cloud start a bunch of different routines to maintain its own code base. And we actually do this for the CLI, for the iOS app, for the Android ap…”
Opinion
Cherny: Claude 3.5 Sonnet Is Terrible by Modern Standards
“This was like Sonnet 3.5. At the time, that was an incredible coding model. That was like the best coding model that exists. Nowadays, it's, you know, a pretty terrible coding model by modern standards. But I think that was like the first great coding model th…”
Assertion Open · timeframe Jul 2027
Cherny: Claude Opus 5 Can Run Autonomously for Months Without Scaffolding
“For five, one example of something it does that I think no other model has done is it runs for a very long period of time. And especially when you combine Opus Five with auto mode, it's just like incredible. Like it can go for days, weeks, months at a time. It…”
Assertion Not checkable as stated
Claude Code's plugin feature was built entirely by an autonomous swarm
“I think the first kind of big example where it worked is our plugins feature was entirely built by a swarm over, over a weekend. It just ran for like a few days. There wasn't really human intervention and plugins is pretty much in the form that it was when it …”
Prediction Not checkable as stated
Non-engineering roles at Anthropic already write code, a trend spreading everywhere
“Every single function on our team codes, like our PM's code, our designers code, our EM codes, our like everyone, our finance guy codes, like everyone on our team codes. We're going to start to see this everywhere.”
Disclosure
Anthropic's path to safe AGI centers on coding and tool use
“You know, as Anthropic, I think for Ant, the bet has been coding
For a long time.
And the bet has been the path to save, to safe AGI is through coding.
And this is, this has kind of always been the idea.
And the way you get there is you teach the model how to …”
Assertion Not checkable as stated
Hu: Anthropic surpasses OpenAI as top API for YC applicants
“And, shockingly, in this batch, the number one API is actually Anthropic. It came out a bit more than OpenAI, which who would have thought?”
Insight
Joseph: LLM training requires collaborative infrastructure work over publishable research papers
“And to do a project like training a large language model requires a lot of people to collaborate on like a really complicated piece of infrastructure that isn't going to be a paper, right? Like you're not going to publish like, oh, I got a slightly, I got five…”
Insight
Nick Joseph: Third parties can steer frontier AI labs by publishing evals
“Like, it is the case that, like, the labs right now are really driven by getting good eval scores. And it's hard to make them, and anyone can do it. There's no comparative advantage to having the model to making an eval. So I do think it's actually, like, an i…”
Insight
Anthropic's Nick Joseph: Frontier pre-training teams primarily need engineers, not researchers
“The thing we, like, most need is engineers. Almost always, like, throughout, like, the entire history of this field. It's, like, the case that you throw more compute, the thing kind of works. The challenge is, like, actually doing that.”
Prediction Not checkable as stated
Anthropic's Joseph: Scaling alone likely will not achieve AGI without further paradigm shifts
“Like I think the sort of shift towards more RL is like one paradigm shift in the field, and I think it's, I think there will probably be more. I think a lot of people sort of argue about like, oh, it's like, you know, current paradigm's enough to get us to EGI…”
Insight
Joseph: Sparsely linked long-tail data may be most valuable for frontier AI
“And it might be that like, that data ends up more valuable because you, everything that's linked to a lot, you've already got. Like at some point, you're maybe like going for the tails, or you're going for the stuff that no one's ever, like, you know, it's onl…”
Insight
Nick Joseph: Scaling standard models is easier and more reliable than inventing novel architectures
“It's just that scale is easier, and it's more reliable, and I think you, we're still seeing really big gains to that.”
Prediction Not checkable as stated
Anthropic's Joseph: Certain AI Alignment Pieces Will Move to Pre-Training
“I do think at some point there will be, like, some pieces of alignment that, like, you do want to export back into pre-training because that might be a way to, like, Put them in with more strength, like, more robustness, kind of, or more core to the intelligen…”
Insight
Joseph: Training purely on raw LLM generations cannot produce a better model
“Theoretically, I shouldn't be able to train a better model than that. Like, I'm just going to get the same thing out. So I think that's-”
Insight
Anthropic's Joseph: Compute matters far more than pre-training objective details
“I think that, like, the one sort of general intuition I have is, like, compute is the thing that matters. So, like, I think if you throw enough compute at any of these objectives, you're gonna get something that's probably pretty good, and can kind of be fine …”
Prediction Not checkable as stated
Kaplan: AI scaling curves point smoothly toward human-level AGI
“I think that scaling Really suggests a kind of smooth curve towards what I expect is kind of human level AI or AGI.”
Insight
Kaplan: Broken scaling laws usually signal flawed training setups, not fundamental limits
“If scaling laws are failing, it's because we've screwed up AI training in some way. Maybe we got, ah, we got the architecture of the neural network wrong, or there's some bottleneck in training that we don't see, or there's some problem with Precision and the …”
Insight
Kaplan: Founders should build products that current AI cannot quite support
“One is I think it's really a good idea to build things that don't quite work yet.
This is probably always a good idea.
We always want to have ambition, but I think specifically AI models right now are getting better very, very quickly.
And I think that's going…”
Prediction Open · timeframe Jul 2028
Srinivas: OpenAI and Anthropic Will Both Build Their Own Browsers
“That said, I'm fully working with the assumption that OpenAI will also build its own browser. Anthropic will also try to build its own browser.”
Opinion
Google has better consumer product taste than both OpenAI and Anthropic
“Out of the list you mentioned, Google is the only company that actually has the product taste to do this.”
Insight
Cherny: Claude Code Users Should Delete CLAUDE.md Scaffolding Every Six Months
“Every six months, delete your Cloud MD. Delete your skills. Delete your hooks. See what the model does, and it might surprise you. And actually for Opus Five, this is something we really do recommend, is just try deleting all of these things, because the model…”
Assertion Not checkable as stated
Cherny: Claude still struggles with systems code, distributed systems, and UI verification
“So coding is solved for the kind of coding that I do. It's not solved for everyone. You know, there's still code bases that are like super deep systems code bases where quad still struggles. There's distributed systems where quad still struggles. There's reall…”
Assertion Supported
Anthropic Deleted 80% of Claude Code System Prompt for Opus 5
“So yeah, we deleted 80% of the system prompt.”
Assertion Not checkable as stated
Cherny: Claude autonomously created a Slack channel to live-blog task progress
“And actually in this case, Quad also decided to live blog it. So what it did is it created a Slack channel internally, and it started just posting screenshots every few minutes of its progress.”
Assertion Contradicted
Ries: FTX bankruptcy's Anthropic stake exceeds the entire fraud value
“Apparently the stake that the bankruptcy has of those shares is worth more than the whole, than all of the entire fraud by a lot.”
Insight
Mukund Jha: Coding is only 20% of shipping software to production
“Our view is that I think the coding aspect is only 20% of the job, right? I think like taking an app to production is like really, really hard.”
Assertion Supported
Poetiq scored 55% on Humanity's Last Exam, outperforming Claude Opus 4.6
“AI hasn't passed it yet, but we got to 55%, which is almost two percentage points higher than the previous state of the art. Which came out just last week from Anthropic with Claude Opus 4.6. They got 53.1%, and we got 55% on it.”
Assertion Contradicted
Mercury data shows 70 percent of startups choose Claude
“There was some stuff from Mercury that like, 70% of startups are, you know, choosing quad as their model of choice.”
Disclosure
Anthropic uses Claude Agent SDK to automate reviews and deployments
“We use the Quad Agent SDK to automate pretty much every part of development. It automates code review, security review it labels all of our issues. It shepherds things to production. It does pretty much everything for us.”
Assertion Not checkable as stated
Anthropic built the Claude desktop app in 10 days using AI
“And they built it in, I think something like 10 days. It was just like a hundred percent written by quad code.”
Insight
Anthropic designs tools for AI models expected six months in the future
“At Anthropic, the way that we thought about it is we don't build for the model of today. We build for the model six months from now.”
Assertion Not checkable as stated
Claude Code autonomously messages Anthropic engineers on Slack via git blame
“A really common pattern is I ask claw to build something. It'll look in the code base. It'll see some engineer touch something in the git flame and then it'll message that engineer on Slack. Just like asking a clarifying question. And then once it gets the ans…”
Assertion Supported
SemiAnalysis reports Claude Code generates four percent of all public commits
“There were some other stuff from like semi-analysis that four percent of all public commits are made by quad code.”
Disclosure
No code in the Claude Code codebase is over six months old
“All of CloudCode has just been written and rewritten and rewritten and rewritten over and over and over. We unship tools every couple weeks. We add new tools every couple weeks. There's no part of CloudCode that was around six months ago. It's just constantly …”
Assertion Not checkable as stated
Friedman: Vast majority of founder AI use cases are not coding
“The vast majority of the use cases people are using it for though is not coding.”
Insight
Diana Hu: Series B AI startups are orchestrating and arbitraging models
“Startups are doing as well, they are actually arbitraging a lot of the models. I had some conversations with a number of founders where before they might have been loyalists to, let's say, OpenAI models or Anthropic, and I just had some conversations recently …”
Insight
Joseph: Reinforcement learning exhibits scaling laws where compute yields better models
“You can get pretty big wins from RL. You sort of have another set of scaling laws. It's like you put more and more compute into RL, you can get better and better models out of that.”
Disclosure
Anthropic built custom distributed training to scale beyond Facebook's infrastructure
“We don't want to outsource this to some package because A, we're about to go to a bigger scale, like PyTorch, for instance, they had a package for doing this. But we were going to go to a bigger scale than Facebook had been to. And you don't want to have a dep…”
Disclosure
Anthropic's rate limits are caused by short-notice compute shortages
“Anthropic has rate limits constantly, and people complain about it a lot, and like the reason is like, there's only so much compute we can get on short notice, so you, like, making your inference more efficient is like the way you can serve more users.”
Assertion Open · timeframe Sep 2026
Joseph: Adversaries actively publish web data designed to poison AI models
“There are people who are, like, trying to put stuff out that is, like, as damaging as possible for the model, you know, how can I make it past the filter and get into the model would be totally like secretly useless.”
Insight
Nick Joseph: Solving coding interview questions proved to be shockingly narrow, not AGI
“I used to think that if you had an AI that could solve coding interview questions, it would probably be AGI. I was like, that's what I did to get my job, I could probably do the job. And it turns out like, nope, nope, you solve those, it's shockingly narrow, a…”