Jul 27, 2026 · 35m · y-combinator

Boris Cherny: We Cut 80% of Claude Code’s Prompt · Y Combinator

Boris Cherny · 24m spoken Diana Hu · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Y Combinator Startup School 2026, Boris Cherny, Head of Claude Code at Anthropic, discusses the technical breakthroughs of Opus 5, explaining how deleting prompt scaffolding, unhobbling models, and enabling long-running autonomous agent workflows are fundamentally transforming software development.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The partners hold 18.3% of the talking time here. How this is scored →

The partners as informed peer 3.7 Guest teaching 3.7 Guest disagreement 0.8 The partners pushing back 0.4
05100:0010:0020:0030:000:00–3:21 · The partners as informed peer 4/10 Y Combinator Startup School 2026 Title Sequence Diana opens the session citing benchmark metrics like ARC-AGI scores and introduces Opus 5. Boris elaborates on long-running tasks and explains how mechanistic interpretability prevents prompt injection in a collaborative, informative tone.3:21–6:55 · The partners as informed peer 4/10 Reducing Claude Code System Prompts by 80% Diana synthesizes how Claude Code radically departs from traditional product engineering by wiping prompts and harnesses. Boris gently clarifies that they don't wipe the entire codebase but run systematic ablations to remove bloat.6:55–9:38 · The partners as informed peer 3/10 Empirical Prompt Building and Biological AI Mindset Diana asks practical questions about reconstructing prompts after wiping them. Boris explains the empirical mindset needed, treating the LLM like an organic creature requiring observation rather than traditional upfront software architecture.9:38–14:46 · The partners as informed peer 5/10 Managing Evals and Avoiding Saturated Metrics When Diana suggests evals are the one stable constant to keep appending to, Boris pushes back and corrects her, explaining that models rapidly saturate evals and require them to be rewritten. He then introduces the concepts of unhobbling and product overhang.14:46–19:44 · The partners as informed peer 3/10 Giving High-Level Goals and the 11-Day Bun Rewrite Boris shares the case study of Bun being rewritten from Zig to Rust over 11 days. When Diana asks if it was accomplished in a single shot, Boris clarifies that it required steering rather than pure zero-shot execution.19:44–23:10 · The partners as informed peer 3/10 Verification Loops and Long-Running Autonomous Tasks Diana engages the audience on whether anyone has run multi-week autonomous agent tasks. Boris explains that verification feedback loops, rather than complex prompting scaffolding, are the true key to long-running tasks.23:10–25:19 · The partners as informed peer 3/10 Unlearning Over-Engineering and Developing Empirical Intuition Boris dismisses conventional online prompt engineering advice and explains how experienced software engineers must unlearn over-specification habits to treat models more like autonomous co-workers.25:19–30:16 · The partners as informed peer 4/10 Dynamic Workflows, Routines, and Self-Maintaining Codebases Diana asks how power users orchestrate thousands of agents, prompting Boris to detail dynamic workflows, functional programming abstractions for agent orchestration, and automated routines maintaining Anthropic's repositories.30:16–34:46 · The partners as informed peer 4/10 The Changing Role of Software Engineering and Advice for Students Diana presses Boris on his previous claim that coding is solved. Boris immediately adds nuance by specifying the exact domains where models still fail, before giving practical advice to students on learning applied problem solving.0:00–3:21 · Guest teaching 4/10 Y Combinator Startup School 2026 Title Sequence Diana opens the session citing benchmark metrics like ARC-AGI scores and introduces Opus 5. Boris elaborates on long-running tasks and explains how mechanistic interpretability prevents prompt injection in a collaborative, informative tone.3:21–6:55 · Guest teaching 3/10 Reducing Claude Code System Prompts by 80% Diana synthesizes how Claude Code radically departs from traditional product engineering by wiping prompts and harnesses. Boris gently clarifies that they don't wipe the entire codebase but run systematic ablations to remove bloat.6:55–9:38 · Guest teaching 4/10 Empirical Prompt Building and Biological AI Mindset Diana asks practical questions about reconstructing prompts after wiping them. Boris explains the empirical mindset needed, treating the LLM like an organic creature requiring observation rather than traditional upfront software architecture.9:38–14:46 · Guest teaching 5/10 Managing Evals and Avoiding Saturated Metrics When Diana suggests evals are the one stable constant to keep appending to, Boris pushes back and corrects her, explaining that models rapidly saturate evals and require them to be rewritten. He then introduces the concepts of unhobbling and product overhang.14:46–19:44 · Guest teaching 3/10 Giving High-Level Goals and the 11-Day Bun Rewrite Boris shares the case study of Bun being rewritten from Zig to Rust over 11 days. When Diana asks if it was accomplished in a single shot, Boris clarifies that it required steering rather than pure zero-shot execution.19:44–23:10 · Guest teaching 3/10 Verification Loops and Long-Running Autonomous Tasks Diana engages the audience on whether anyone has run multi-week autonomous agent tasks. Boris explains that verification feedback loops, rather than complex prompting scaffolding, are the true key to long-running tasks.23:10–25:19 · Guest teaching 4/10 Unlearning Over-Engineering and Developing Empirical Intuition Boris dismisses conventional online prompt engineering advice and explains how experienced software engineers must unlearn over-specification habits to treat models more like autonomous co-workers.25:19–30:16 · Guest teaching 3/10 Dynamic Workflows, Routines, and Self-Maintaining Codebases Diana asks how power users orchestrate thousands of agents, prompting Boris to detail dynamic workflows, functional programming abstractions for agent orchestration, and automated routines maintaining Anthropic's repositories.30:16–34:46 · Guest teaching 4/10 The Changing Role of Software Engineering and Advice for Students Diana presses Boris on his previous claim that coding is solved. Boris immediately adds nuance by specifying the exact domains where models still fail, before giving practical advice to students on learning applied problem solving.0:00–3:21 · Guest disagreement 0/10 Y Combinator Startup School 2026 Title Sequence Diana opens the session citing benchmark metrics like ARC-AGI scores and introduces Opus 5. Boris elaborates on long-running tasks and explains how mechanistic interpretability prevents prompt injection in a collaborative, informative tone.3:21–6:55 · Guest disagreement 1/10 Reducing Claude Code System Prompts by 80% Diana synthesizes how Claude Code radically departs from traditional product engineering by wiping prompts and harnesses. Boris gently clarifies that they don't wipe the entire codebase but run systematic ablations to remove bloat.6:55–9:38 · Guest disagreement 0/10 Empirical Prompt Building and Biological AI Mindset Diana asks practical questions about reconstructing prompts after wiping them. Boris explains the empirical mindset needed, treating the LLM like an organic creature requiring observation rather than traditional upfront software architecture.9:38–14:46 · Guest disagreement 2/10 Managing Evals and Avoiding Saturated Metrics When Diana suggests evals are the one stable constant to keep appending to, Boris pushes back and corrects her, explaining that models rapidly saturate evals and require them to be rewritten. He then introduces the concepts of unhobbling and product overhang.14:46–19:44 · Guest disagreement 1/10 Giving High-Level Goals and the 11-Day Bun Rewrite Boris shares the case study of Bun being rewritten from Zig to Rust over 11 days. When Diana asks if it was accomplished in a single shot, Boris clarifies that it required steering rather than pure zero-shot execution.19:44–23:10 · Guest disagreement 0/10 Verification Loops and Long-Running Autonomous Tasks Diana engages the audience on whether anyone has run multi-week autonomous agent tasks. Boris explains that verification feedback loops, rather than complex prompting scaffolding, are the true key to long-running tasks.23:10–25:19 · Guest disagreement 1/10 Unlearning Over-Engineering and Developing Empirical Intuition Boris dismisses conventional online prompt engineering advice and explains how experienced software engineers must unlearn over-specification habits to treat models more like autonomous co-workers.25:19–30:16 · Guest disagreement 0/10 Dynamic Workflows, Routines, and Self-Maintaining Codebases Diana asks how power users orchestrate thousands of agents, prompting Boris to detail dynamic workflows, functional programming abstractions for agent orchestration, and automated routines maintaining Anthropic's repositories.30:16–34:46 · Guest disagreement 2/10 The Changing Role of Software Engineering and Advice for Students Diana presses Boris on his previous claim that coding is solved. Boris immediately adds nuance by specifying the exact domains where models still fail, before giving practical advice to students on learning applied problem solving.0:00–3:21 · The partners pushing back 0/10 Y Combinator Startup School 2026 Title Sequence Diana opens the session citing benchmark metrics like ARC-AGI scores and introduces Opus 5. Boris elaborates on long-running tasks and explains how mechanistic interpretability prevents prompt injection in a collaborative, informative tone.3:21–6:55 · The partners pushing back 1/10 Reducing Claude Code System Prompts by 80% Diana synthesizes how Claude Code radically departs from traditional product engineering by wiping prompts and harnesses. Boris gently clarifies that they don't wipe the entire codebase but run systematic ablations to remove bloat.6:55–9:38 · The partners pushing back 0/10 Empirical Prompt Building and Biological AI Mindset Diana asks practical questions about reconstructing prompts after wiping them. Boris explains the empirical mindset needed, treating the LLM like an organic creature requiring observation rather than traditional upfront software architecture.9:38–14:46 · The partners pushing back 1/10 Managing Evals and Avoiding Saturated Metrics When Diana suggests evals are the one stable constant to keep appending to, Boris pushes back and corrects her, explaining that models rapidly saturate evals and require them to be rewritten. He then introduces the concepts of unhobbling and product overhang.14:46–19:44 · The partners pushing back 1/10 Giving High-Level Goals and the 11-Day Bun Rewrite Boris shares the case study of Bun being rewritten from Zig to Rust over 11 days. When Diana asks if it was accomplished in a single shot, Boris clarifies that it required steering rather than pure zero-shot execution.19:44–23:10 · The partners pushing back 0/10 Verification Loops and Long-Running Autonomous Tasks Diana engages the audience on whether anyone has run multi-week autonomous agent tasks. Boris explains that verification feedback loops, rather than complex prompting scaffolding, are the true key to long-running tasks.23:10–25:19 · The partners pushing back 0/10 Unlearning Over-Engineering and Developing Empirical Intuition Boris dismisses conventional online prompt engineering advice and explains how experienced software engineers must unlearn over-specification habits to treat models more like autonomous co-workers.25:19–30:16 · The partners pushing back 0/10 Dynamic Workflows, Routines, and Self-Maintaining Codebases Diana asks how power users orchestrate thousands of agents, prompting Boris to detail dynamic workflows, functional programming abstractions for agent orchestration, and automated routines maintaining Anthropic's repositories.30:16–34:46 · The partners pushing back 1/10 The Changing Role of Software Engineering and Advice for Students Diana presses Boris on his previous claim that coding is solved. Boris immediately adds nuance by specifying the exact domains where models still fail, before giving practical advice to students on learning applied problem solving.

speaking balance: gold is the partners, purple is the guest (3 minute bins)

0:00 · the partners 23.5% · guest 76.5%0:00 · the partners 23.5% · guest 76.5%3:00 · the partners 24.4% · guest 75.6%3:00 · the partners 24.4% · guest 75.6%6:00 · the partners 20% · guest 80%6:00 · the partners 20% · guest 80%9:00 · the partners 22.5% · guest 77.5%9:00 · the partners 22.5% · guest 77.5%12:00 · the partners 28.3% · guest 71.7%12:00 · the partners 28.3% · guest 71.7%15:00 · the partners 7.7% · guest 92.3%15:00 · the partners 7.7% · guest 92.3%18:00 · the partners 14.1% · guest 85.9%18:00 · the partners 14.1% · guest 85.9%21:00 · the partners 22.1% · guest 77.9%21:00 · the partners 22.1% · guest 77.9%24:00 · the partners 16% · guest 84%24:00 · the partners 16% · guest 84%27:00 · the partners 0% · guest 100%27:00 · the partners 0% · guest 100%30:00 · the partners 23.4% · guest 76.6%30:00 · the partners 23.4% · guest 76.6%33:00 · the partners 18.7% · guest 81.3%33:00 · the partners 18.7% · guest 81.3%
Sharpest disagreement ▶ 9:53 Boris rejects the idea that evals are constant

Boris directly rejects Diana's premise that evals remain a permanent constant across model generations, emphasizing that exponential capability growth quickly saturates and invalidates them.

Hardest push from the partners ▶ 30:15 Diana challenges Boris on declaring coding solved

Diana holds Boris accountable to his past provocative statements by asking what distinguishes builders if software engineering is truly solved.

Biggest teaching moment ▶ 9:53 Explaining eval saturation and deprecation cycles

Boris educates the audience and host on why evals must be regularly thrown out due to exponential model improvement rather than accumulated indefinitely.

The partners hold their own ▶ 13:54 Diana frames Claude Code's terminal access as unhobbling

Diana synthesizes the core insight of product overhang, connecting Sonnet 3.5's architectural breakthrough directly to eliminating IDE scaffolding.

the scores for every segment, with the reasoning behind each
ChapterTopicThe partners as informed peerGuest teachingGuest disagreementThe partners pushing backWhy
Y Combinator Startup School 2026 Title Sequence 4400 Diana opens the session citing benchmark metrics like ARC-AGI scores and introduces Opus 5. Boris elaborates on long-running tasks and explains how mechanistic interpretability prevents prompt injection in a collaborative, informative tone.
Reducing Claude Code System Prompts by 80% 4311 Diana synthesizes how Claude Code radically departs from traditional product engineering by wiping prompts and harnesses. Boris gently clarifies that they don't wipe the entire codebase but run systematic ablations to remove bloat.
Empirical Prompt Building and Biological AI Mindset 3400 Diana asks practical questions about reconstructing prompts after wiping them. Boris explains the empirical mindset needed, treating the LLM like an organic creature requiring observation rather than traditional upfront software architecture.
Managing Evals and Avoiding Saturated Metrics 5521 When Diana suggests evals are the one stable constant to keep appending to, Boris pushes back and corrects her, explaining that models rapidly saturate evals and require them to be rewritten. He then introduces the concepts of unhobbling and product overhang.
Giving High-Level Goals and the 11-Day Bun Rewrite 3311 Boris shares the case study of Bun being rewritten from Zig to Rust over 11 days. When Diana asks if it was accomplished in a single shot, Boris clarifies that it required steering rather than pure zero-shot execution.
Verification Loops and Long-Running Autonomous Tasks 3300 Diana engages the audience on whether anyone has run multi-week autonomous agent tasks. Boris explains that verification feedback loops, rather than complex prompting scaffolding, are the true key to long-running tasks.
Unlearning Over-Engineering and Developing Empirical Intuition 3410 Boris dismisses conventional online prompt engineering advice and explains how experienced software engineers must unlearn over-specification habits to treat models more like autonomous co-workers.
Dynamic Workflows, Routines, and Self-Maintaining Codebases 4300 Diana asks how power users orchestrate thousands of agents, prompting Boris to detail dynamic workflows, functional programming abstractions for agent orchestration, and automated routines maintaining Anthropic's repositories.
The Changing Role of Software Engineering and Advice for Students 4421 Diana presses Boris on his previous claim that coding is solved. Boris immediately adds nuance by specifying the exact domains where models still fail, before giving practical advice to students on learning applied problem solving.

Statements from this episode (22)

Assertion Supported
Hu: Anthropic Opus 5 achieved 30% on ARC-AGI
“You guys got, took Arc AGI three to 30%, which is incredible.”
Diana Hu Jul 27, 2026 ▶ 0:33
Assertion Open · timeframe Jul 2027
Cherny: Claude Opus 5 Can Run Autonomously for Months Without Scaffolding
“For five, one example of something it does that I think no other model has done is it runs for a very long period of time. And especially when you combine Opus Five with auto mode, it's just like incredible. Like it can go for days, weeks, months at a time. It…”
Boris Cherny Jul 27, 2026 ▶ 1:25
Assertion Contradicted
Cherny: Anthropic Can No Longer Demonstrate Prompt Injection on Opus 5
“So essentially, if you combine a well-aligned model, so this is, like, essentially three years of research into alignment, With a prompt injection classifier, which we run for all traffic, and what this is doing is it's based on Chrysola's mechanistic interpre…”
Boris Cherny Jul 27, 2026 ▶ 2:46
Assertion Supported
Anthropic Deleted 80% of Claude Code System Prompt for Opus 5
“So yeah, we deleted 80% of the system prompt.”
Boris Cherny Jul 27, 2026 ▶ 4:30
Assertion Not checkable as stated
Cherny: Claude Is More Intelligent Without System Prompts
“And what's interesting is that the model is actually a little bit more intelligent without these prompts. That's something that we've been finding.”
Boris Cherny Jul 27, 2026 ▶ 5:06
Assertion Not checkable as stated
Cherny: Most Claude Code Harness Code Is Safety and Permissions
“If you look at actually the code that's in the cloud code harness today, almost all of it is about safety and permissions and static analysis, and there's a bunch of UI code, and we've actually unshipped a lot of the other code already.”
Boris Cherny Jul 27, 2026 ▶ 6:24
Insight
Cherny: Claude Code Users Should Delete CLAUDE.md Scaffolding Every Six Months
“Every six months, delete your Cloud MD. Delete your skills. Delete your hooks. See what the model does, and it might surprise you. And actually for Opus Five, this is something we really do recommend, is just try deleting all of these things, because the model…”
Boris Cherny Jul 27, 2026 ▶ 6:56
Insight
Cherny: Only Add System Prompt Instructions When Models Repeatedly Stumble
“The thing that you want to do is you want to run it. And if it's like a custom agentic product that you're building, you want to kind of run the products. You want to see where it fails with the model. You want to see what it does well. If you're using quad co…”
Boris Cherny Jul 27, 2026 ▶ 7:55
Insight
Cherny: AI evals saturate and must be discarded every few generations
“I think evals, they outlive the harness a little bit, but not quite that much. Like, an eval might live for maybe one, two, three model generations, but nowadays the, you know, we're on the exponential. The model is improving so quickly, very often we just sat…”
Boris Cherny Jul 27, 2026 ▶ 10:01
Opinion
Cherny: Claude 3.5 Sonnet Is Terrible by Modern Standards
“This was like Sonnet 3.5. At the time, that was an incredible coding model. That was like the best coding model that exists. Nowadays, it's, you know, a pretty terrible coding model by modern standards. But I think that was like the first great coding model th…”
Boris Cherny Jul 27, 2026 ▶ 12:17
Insight
Cherny: Startups Are Missing Massive AI Product Overhang Opportunities
“I think that nowadays, with modern models, there is so much product overhang that I, I'm not seeing startups capture. And I think there's people thinking about these problems, but there's just a huge amount, amount of opportunity to elicit these behaviors from…”
Boris Cherny Jul 27, 2026 ▶ 13:34
Insight
Cherny: Prompt modern models with high-level guardrails, not micromanaged steps
“And for modern models, that's actually really not the way to do it. You want to go a little bit higher level. You want to describe the task. You want to describe the guardrails. You want to describe, like, the exit criteria, and then just go with the model coo…”
Boris Cherny Jul 27, 2026 ▶ 15:14
Assertion Not checkable as stated
Cherny: Modern models can rewrite essentially any codebase into another language
“One example is the model can now rewrite essentially any code base from one language to a different language.”
Boris Cherny Jul 27, 2026 ▶ 15:46
Assertion Supported
Cherny: Bun codebase was rewritten to Rust in 11 days using Claude
“And he had the model rewrite it from Zig to Rust. It was one prompt. It was a dynamic workflow, and a dynamic workflows are a feature in quad code that essentially let you orchestrate, you know, dozens, 100,000 of agents to do work productively. And it ran for…”
Boris Cherny Jul 27, 2026 ▶ 17:24
Assertion Supported
Cherny: Claude Code runs in production on Bun's AI-rewritten Rust codebase
“This is in production now. This is what quad code uses now when you're running it.”
Boris Cherny Jul 27, 2026 ▶ 18:12
Insight
Cherny: Verification loops are the most important overlooked aspect of AI prompting
“I think the skill nowadays is less about prompt engineering and more about figuring out how do you give Cloud a hard task that seems a little bit too hard? And then how do you make it possible for Cloud to verify its work along the way? And the verification, I…”
Boris Cherny Jul 27, 2026 ▶ 20:13
Assertion Not checkable as stated
Cherny: Claude autonomously created a Slack channel to live-blog task progress
“And actually in this case, Quad also decided to live blog it. So what it did is it created a Slack channel internally, and it started just posting screenshots every few minutes of its progress.”
Boris Cherny Jul 27, 2026 ▶ 22:48
Insight
Cherny: Veteran engineers fail with AI models by over-specifying tasks
“When I look at engineers that have been, you know, coding for a long time, you know, like for years or for decades, this is a really, really common failure mode is trying to over specify and it's trying to be overly specific and then, you know, get the model t…”
Boris Cherny Jul 27, 2026 ▶ 24:10
Assertion Partly supported
Cherny: Claude Code uses Bun sandboxes to orchestrate dynamic agent workflows
“What a dynamic workflow is, is essentially we have the Bun runtime. We use Bun as a sandbox, and we start a virtual machine within Bun, and we let Cloud start a lot of agents and orchestrate them.”
Boris Cherny Jul 27, 2026 ▶ 25:37
Disclosure
Cherny: Anthropic uses Claude routines to maintain its own apps via Slack
“We actually have Cloud maintaining itself now. And the way we do this is we have a Slack channel where we just had Cloud start a bunch of different routines to maintain its own code base. And we actually do this for the CLI, for the iOS app, for the Android ap…”
Boris Cherny Jul 27, 2026 ▶ 28:12
Assertion Not checkable as stated
Cherny: Claude routines do the maintenance work of hundreds of engineers
“And so now we have every day, maybe 20 or 30 of these routines. It's running across all of our code bases and It's not totally there yet, but we're on the path to fully automating the maintenance of our apps by doing this. And this is, again, hundreds of agent…”
Boris Cherny Jul 27, 2026 ▶ 29:40
Assertion Not checkable as stated
Cherny: Claude still struggles with systems code, distributed systems, and UI verification
“So coding is solved for the kind of coding that I do. It's not solved for everyone. You know, there's still code bases that are like super deep systems code bases where quad still struggles. There's distributed systems where quad still struggles. There's reall…”
Boris Cherny Jul 27, 2026 ▶ 30:40
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.