why aren't all 55 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 1 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Contradicted
Cherny: Anthropic Can No Longer Demonstrate Prompt Injection on Opus 5
“So essentially, if you combine a well-aligned model, so this is, like, essentially three years of research into alignment, With a prompt injection classifier, which we run for all traffic, and what this is doing is it's based on Chrysola's mechanistic interpre…”
Assertion Supported
Cherny: Bun codebase was rewritten to Rust in 11 days using Claude
“And he had the model rewrite it from Zig to Rust. It was one prompt. It was a dynamic workflow, and a dynamic workflows are a feature in quad code that essentially let you orchestrate, you know, dozens, 100,000 of agents to do work productively. And it ran for…”
Assertion Not checkable as stated
Cherny: Claude routines do the maintenance work of hundreds of engineers
“And so now we have every day, maybe 20 or 30 of these routines. It's running across all of our code bases and It's not totally there yet, but we're on the path to fully automating the maintenance of our apps by doing this. And this is, again, hundreds of agent…”
Assertion Not checkable as stated
Claude Code increased Anthropic engineer productivity by 150 percent
“And since Quad Code came out, productivity per engineer at Anthropic has grown a 150%.”
Disclosure
Cherny uninstalled his IDE and lands 20 AI-written PRs daily
“For me personally, it's been a hundred percent for like since Opus 4.5. I just, I uninstalled my IDE. I don't edit a single line of code by hand. It's just a hundred percent quad code in Opus. And you know, I land, you know, like 20 pairs a day, every day.”
Assertion Not checkable as stated
Claude generates between 70 and 90 percent of Anthropic's codebase
“If you look at Anthropic overall, it ranges between, like, 70 to 90% you know, depending on the team. For a lot of teams, it's also, like, a hundred percent. For a lot of people, it's a hundred percent.”
Prediction Not checkable as stated
Cherny predicts AI will generally solve coding for everyone
“So continuing to trace the exponential, I think what will happen is coding will be generally solved for everyone. And I think today coding is practically solved, you know, for me, and I think it'll be the case for everyone you know, regardless of domain.”
Prediction Not checkable as stated
Cherny: The software engineer title will disappear as roles become builders
“I think we're gonna start to see the title software engineer go away, and I think it's just gonna be maybe builder, maybe product manager, maybe we'll keep the title as kind of a vestigial thing, but the work that people do, it's not just gonna be coding, it's…”
Assertion Open · timeframe Jul 2027
Cherny: Claude Opus 5 Can Run Autonomously for Months Without Scaffolding
“For five, one example of something it does that I think no other model has done is it runs for a very long period of time. And especially when you combine Opus Five with auto mode, it's just like incredible. Like it can go for days, weeks, months at a time. It…”
Assertion Not checkable as stated
Cherny: Claude Is More Intelligent Without System Prompts
“And what's interesting is that the model is actually a little bit more intelligent without these prompts. That's something that we've been finding.”
Opinion
Cherny: Claude 3.5 Sonnet Is Terrible by Modern Standards
“This was like Sonnet 3.5. At the time, that was an incredible coding model. That was like the best coding model that exists. Nowadays, it's, you know, a pretty terrible coding model by modern standards. But I think that was like the first great coding model th…”
Assertion Not checkable as stated
Cherny: Modern models can rewrite essentially any codebase into another language
“One example is the model can now rewrite essentially any code base from one language to a different language.”
Disclosure
Cherny: Anthropic uses Claude routines to maintain its own apps via Slack
“We actually have Cloud maintaining itself now. And the way we do this is we have a Slack channel where we just had Cloud start a bunch of different routines to maintain its own code base. And we actually do this for the CLI, for the iOS app, for the Android ap…”
Disclosure
Anthropic's path to safe AGI centers on coding and tool use
“You know, as Anthropic, I think for Ant, the bet has been coding
For a long time.
And the bet has been the path to save, to safe AGI is through coding.
And this is, this has kind of always been the idea.
And the way you get there is you teach the model how to …”
Insight
Scaffolding gains get wiped out by next-generation AI models
“There's this other general principle that I think is maybe interesting where you can build for the model and then you can build scaffolding around the model in order to improve performance a little bit. And depending on the domain, you can improve performance,…”
Insight
Delete oversized CLAUDE.md files and start fresh rather than compacting them
“If you hit this, my recommendation would be delete your QuadMD and just start fresh.”
Assertion Not checkable as stated
Claude Code's plugin feature was built entirely by an autonomous swarm
“I think the first kind of big example where it worked is our plugins feature was entirely built by a swarm over, over a weekend. It just ran for like a few days. There wasn't really human intervention and plugins is pretty much in the form that it was when it …”
Prediction Not checkable as stated
Non-engineering roles at Anthropic already write code, a trend spreading everywhere
“Every single function on our team codes, like our PM's code, our designers code, our EM codes, our like everyone, our finance guy codes, like everyone on our team codes. We're going to start to see this everywhere.”
Assertion Supported
Anthropic Deleted 80% of Claude Code System Prompt for Opus 5
“So yeah, we deleted 80% of the system prompt.”
Insight
Cherny: Claude Code Users Should Delete CLAUDE.md Scaffolding Every Six Months
“Every six months, delete your Cloud MD. Delete your skills. Delete your hooks. See what the model does, and it might surprise you. And actually for Opus Five, this is something we really do recommend, is just try deleting all of these things, because the model…”
Insight
Cherny: Only Add System Prompt Instructions When Models Repeatedly Stumble
“The thing that you want to do is you want to run it. And if it's like a custom agentic product that you're building, you want to kind of run the products. You want to see where it fails with the model. You want to see what it does well. If you're using quad co…”
Insight
Cherny: AI evals saturate and must be discarded every few generations
“I think evals, they outlive the harness a little bit, but not quite that much. Like, an eval might live for maybe one, two, three model generations, but nowadays the, you know, we're on the exponential. The model is improving so quickly, very often we just sat…”
Insight
Cherny: Startups Are Missing Massive AI Product Overhang Opportunities
“I think that nowadays, with modern models, there is so much product overhang that I, I'm not seeing startups capture. And I think there's people thinking about these problems, but there's just a huge amount, amount of opportunity to elicit these behaviors from…”
Assertion Supported
Cherny: Claude Code runs in production on Bun's AI-rewritten Rust codebase
“This is in production now. This is what quad code uses now when you're running it.”
Insight
Cherny: Verification loops are the most important overlooked aspect of AI prompting
“I think the skill nowadays is less about prompt engineering and more about figuring out how do you give Cloud a hard task that seems a little bit too hard? And then how do you make it possible for Cloud to verify its work along the way? And the verification, I…”
Assertion Not checkable as stated
Cherny: Claude autonomously created a Slack channel to live-blog task progress
“And actually in this case, Quad also decided to live blog it. So what it did is it created a Slack channel internally, and it started just posting screenshots every few minutes of its progress.”
Insight
Cherny: Veteran engineers fail with AI models by over-specifying tasks
“When I look at engineers that have been, you know, coding for a long time, you know, like for years or for decades, this is a really, really common failure mode is trying to over specify and it's trying to be overly specific and then, you know, get the model t…”
Assertion Not checkable as stated
Cherny: Claude still struggles with systems code, distributed systems, and UI verification
“So coding is solved for the kind of coding that I do. It's not solved for everyone. You know, there's still code bases that are like super deep systems code bases where quad still struggles. There's distributed systems where quad still struggles. There's reall…”
Insight
Anthropic designs tools for AI models expected six months in the future
“At Anthropic, the way that we thought about it is we don't build for the model of today. We build for the model six months from now.”
Disclosure
Claude Code stayed a CLI because graphical UIs obsolete too quickly
“And really, I think that's why we stayed in the CLI is because we felt there is no UI we could build that would still be relevant in six months because the model was improving so quickly.”
Insight
Cherny: AI makes traditional senior engineering opinions less relevant
“In my old job at a big company, when I hired like architects and this kind of a type of engineer, you look for people that have a lot of experience and really strong opinions, but it actually turns out a lot of this stuff just isn't relevant anymore. And a lot…”
Disclosure
Anthropic uses Claude Agent SDK to automate reviews and deployments
“We use the Quad Agent SDK to automate pretty much every part of development. It automates code review, security review it labels all of our issues. It shepherds things to production. It does pretty much everything for us.”
Assertion Not checkable as stated
Claude Code autonomously messages Anthropic engineers on Slack via git blame
“A really common pattern is I ask claw to build something. It'll look in the code base. It'll see some engineer touch something in the git flame and then it'll message that engineer on Slack. Just like asking a clarifying question. And then once it gets the ans…”
Assertion Not checkable as stated
Claude Code enables building 20 UI prototypes in a few hours
“You can write these prototypes and you can just do like 20 prototypes back to back, see which one you like, and then ship that, and the whole thing takes maybe a couple hours. Whereas in the past, what you would have had to do is like learn to use origami or f…”
Disclosure
No code in the Claude Code codebase is over six months old
“All of CloudCode has just been written and rewritten and rewritten and rewritten over and over and over. We unship tools every couple weeks. We add new tools every couple weeks. There's no part of CloudCode that was around six months ago. It's just constantly …”
Assertion Contradicted
Mercury data shows 70 percent of startups choose Claude
“There was some stuff from Mercury that like, 70% of startups are, you know, choosing quad as their model of choice.”
Assertion Supported
SemiAnalysis reports Claude Code generates four percent of all public commits
“There were some other stuff from like semi-analysis that four percent of all public commits are made by quad code.”
Assertion Supported
NASA JPL used Claude Code to navigate the Perseverance Mars rover
“It plotted the course for perseverance, like for like the Mars Rover.”
Assertion Not checkable as stated
Anthropic built the Claude desktop app in 10 days using AI
“And they built it in, I think something like 10 days. It was just like a hundred percent written by quad code.”
Insight
Cherny: Prompt modern models with high-level guardrails, not micromanaged steps
“And for modern models, that's actually really not the way to do it. You want to go a little bit higher level. You want to describe the task. You want to describe the guardrails. You want to describe, like, the exit criteria, and then just go with the model coo…”
Assertion Partly supported
Cherny: Claude Code uses Bun sandboxes to orchestrate dynamic agent workflows
“What a dynamic workflow is, is essentially we have the Bun runtime. We use Bun as a sandbox, and we start a virtual machine within Bun, and we let Cloud start a lot of agents and orchestrate them.”
Assertion Not checkable as stated
Claude Code autonomously wrote a tool to debug a memory leak
“And quad code, like took the heap dump. It wrote a little tool for itself to like analyze the heap dump. And then it found the leak faster than I did.”
Insight
Cherny: The most effective AI engineers are hyper-specialists or hyper-generalists
“When I look at engineers on the team that I think are the most effective, there's essentially two, it's very bimodal. There's one side where it's extreme specialists and so, like, I named Jared before, like, he's a really good example of this, and kind of the …”
Opinion
Cherny believes most agents today are prompted as Claude subagents
“I haven't pulled the data on this, but I would bet the majority of agents are actually prompted by quad today in the form of subagents.”
Prediction Not checkable as stated
Cherny predicts plan mode in AI tools has a limited lifespan
“I think plan mode probably has a limited lifespan.”
Opinion
Cherny claims recent Claude Opus models execute plans without babysitting
“And nowadays what I find with Opus 4.5, I think it started with 4.6, it got really good. Once the plan is good, it just stays on track, and it'll just do the thing exactly right almost every time.”
Prediction Not checkable as stated
Cherny predicts Claude will increasingly discover product ideas by analyzing logs
“And I think Quad is going to get increasingly good at kind of figuring out these kind of product ideas for you, just because it can look at feedback, it can look at debug logs, it can kind of figure this out.”
Insight
Cherny: Dev tool startups should enable what AI models naturally attempt
“The way I would frame it is think about the thing that the model wants to do and figure out how do you make that easier? And that's something that we saw, you know, like when I first started hacking on Claude code, I realized like this thing just wants to use …”