Jul 26, 2026 · 1h 33m · lennys-podcast

Why AI is going vertical (again) | Dianne Penn (Anthropic)

Dianne Penn · 58m spoken Lenny Rachitsky · 25m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Anthropic's Head of Product for AI Research and Labs, Dianne Penn, joins Lenny Rachitsky to discuss how frontier models, purpose-built harnesses like Claude Code, and eval-driven product management are transforming software development. She shares insider lessons on Anthropic's rapid experimentation culture, navigating exponential capability jumps, and why human judgment remains essential in an AI-driven world.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 30.2% of the talking time here. How this is scored →

Lenny as informed peer 4.1 Guest teaching 4.0 Guest disagreement 0.1 Lenny pushing back 0.3
05100:0020:0040:001:00:001:20:000:00–2:32 · Lenny as informed peer 1/10 Episode Preview: Key Themes and High-Velocity AI Insights Lenny introduces Dianne Penn and previews the episode themes alongside brief promotional snippets from the conversation.2:36–7:45 · Lenny as informed peer 3/10 Anthropic's Humble Beginnings and the Golden Gate Claude Inflection Lenny recalls Anthropic's early days competing against OpenAI, while Dianne explains the culture and the Golden Gate Claude experiment.7:46–12:38 · Lenny as informed peer 4/10 Sponsor Message: WorkOS Enterprise Solutions Following an ad read, Lenny asks about inflection points, and Dianne breaks down how Opus 3 and long-form coding differentiation established Anthropic's position.12:38–17:22 · Lenny as informed peer 5/10 The Synergy of Opus 4.5 and Claude Code Lenny notes Dario Amodei's accurate coding predictions and the exponential curve, prompting Dianne to discuss adaptability and first-principles reasoning.17:22–20:02 · Lenny as informed peer 5/10 Discontinuous Capabilities, Scaling Laws, and Product Overhang Lenny summarizes model adaptability, and Dianne refines the explanation using scaling law graphs and discontinuous emergent capabilities.20:03–23:33 · Lenny as informed peer 4/10 Token Spending, Communal Experimentation, and Working in Public Lenny brings up Garry Tan's token-spending thesis, which Dianne gently reframes from raw spend input to communal experimentation output.23:33–27:30 · Lenny as informed peer 4/10 Anthropic Labs: Incubating Discontinuous Bets Outside the Core Roadmap Lenny inquires about Anthropic Labs' operational structure, and Dianne outlines how small pods incubate high-risk discontinuous bets.27:30–31:35 · Lenny as informed peer 4/10 Inside the AI Research Loop: Translating User Feedback into Evals Lenny asks what AI researchers do all day, and Dianne provides a detailed operational breakdown of translating vague user complaints into rigorous evals.31:36–35:19 · Lenny as informed peer 4/10 Traits of World-Class Researchers and Designing for Forward Compatibility Lenny asks about researcher traits, and Dianne describes how forward compatibility and first-principles thinking define top AI talent.35:19–38:25 · Lenny as informed peer 4/10 Frontier Model Safety, Safeguard Packages, and Responsible Deployment Lenny highlights increasing safety scrutiny and restricted releases; Dianne sets boundaries regarding policy but explains fallback UX systems.38:25–44:16 · Lenny as informed peer 5/10 Sponsor Message: Mercury Banking with Command AI Operator After an ad break, Dianne explains how PM hiring criteria remained steady while the core PM artifact shifted from PRDs to evals.44:16–46:42 · Lenny as informed peer 5/10 Case Study: Resolving JSON Output Reliability Through Custom Evals Dianne illustrates an eval case study concerning JSON output compliance, which Lenny compares to test-driven development.46:42–50:00 · Lenny as informed peer 4/10 Balancing Evals and PRDs for Large-Scale Alignment and Vision Lenny checks if PRDs are truly dead; Dianne confirms they remain crucial for broad multi-team alignment and exploratory, ambiguous capabilities.50:01–54:05 · Lenny as informed peer 4/10 Hands-On AI Leadership: Why Managers Must Build and Ship Dianne emphasizes that engineering and product managers must actively build and ship to evaluate AI output effectively.54:06–58:21 · Lenny as informed peer 5/10 Collaborative Discovery and Deep Immersion in Specific AI Use Cases Lenny and Dianne discuss community survey insights showing that going deep on narrow use cases yields more satisfaction than superficial tinkering.58:22–1:03:53 · Lenny as informed peer 4/10 Augmenting EQ: Tactical Claude Workflows for Coaching and Thinking Dianne shares how she uses Claude skills for Crucial Conversations coaching, and addresses Lenny's query about cognitive reliance on AI.1:03:53–1:07:05 · Lenny as informed peer 5/10 The Power of Claude's Constitution: Why Pushback Drives Quality Lenny notes Claude's distinctive conversational personality, and Dianne explains how constitutional alignment creates a genuine thinking partner through pushback.1:07:05–1:11:40 · Lenny as informed peer 5/10 The Jagged Frontier of AI Writing and Future Improvements Lenny asks why LLMs struggle with natural prose writing; Dianne explains the jagged capability frontier and prioritizing agentic milestones.1:11:41–1:15:36 · Lenny as informed peer 4/10 Human Judgment, Lifelong Learning, and Developing an Inner Voice Dianne highlights enduring human strengths like judgment and persistence, relating them to parenting strategies in an AI-dominated era.1:15:37–1:21:53 · Lenny as informed peer 4/10 Sustaining High Performance: The Hive Mind and Team Resilience Dianne describes how the Anthropic team maintains resilience through collective ownership, shared coverage, and an ego-free hive mind.1:21:54–1:24:31 · Lenny as informed peer 5/10 Closing Reflections: The Enduring Need for User-Centric PMs Dianne summarizes why deeply user-centric PMs will remain vital, which Lenny enthusiastically validates.1:24:31–1:31:20 · Lenny as informed peer 3/10 Lightning Round: Books, Culture, and Trading Floor Lessons During the lightning round, Dianne shares recommendations and reflects on how her bond trading desk experience taught her to champion conviction and truth over hierarchy.0:00–2:32 · Guest teaching 0/10 Episode Preview: Key Themes and High-Velocity AI Insights Lenny introduces Dianne Penn and previews the episode themes alongside brief promotional snippets from the conversation.2:36–7:45 · Guest teaching 4/10 Anthropic's Humble Beginnings and the Golden Gate Claude Inflection Lenny recalls Anthropic's early days competing against OpenAI, while Dianne explains the culture and the Golden Gate Claude experiment.7:46–12:38 · Guest teaching 5/10 Sponsor Message: WorkOS Enterprise Solutions Following an ad read, Lenny asks about inflection points, and Dianne breaks down how Opus 3 and long-form coding differentiation established Anthropic's position.12:38–17:22 · Guest teaching 4/10 The Synergy of Opus 4.5 and Claude Code Lenny notes Dario Amodei's accurate coding predictions and the exponential curve, prompting Dianne to discuss adaptability and first-principles reasoning.17:22–20:02 · Guest teaching 5/10 Discontinuous Capabilities, Scaling Laws, and Product Overhang Lenny summarizes model adaptability, and Dianne refines the explanation using scaling law graphs and discontinuous emergent capabilities.20:03–23:33 · Guest teaching 4/10 Token Spending, Communal Experimentation, and Working in Public Lenny brings up Garry Tan's token-spending thesis, which Dianne gently reframes from raw spend input to communal experimentation output.23:33–27:30 · Guest teaching 4/10 Anthropic Labs: Incubating Discontinuous Bets Outside the Core Roadmap Lenny inquires about Anthropic Labs' operational structure, and Dianne outlines how small pods incubate high-risk discontinuous bets.27:30–31:35 · Guest teaching 6/10 Inside the AI Research Loop: Translating User Feedback into Evals Lenny asks what AI researchers do all day, and Dianne provides a detailed operational breakdown of translating vague user complaints into rigorous evals.31:36–35:19 · Guest teaching 4/10 Traits of World-Class Researchers and Designing for Forward Compatibility Lenny asks about researcher traits, and Dianne describes how forward compatibility and first-principles thinking define top AI talent.35:19–38:25 · Guest teaching 3/10 Frontier Model Safety, Safeguard Packages, and Responsible Deployment Lenny highlights increasing safety scrutiny and restricted releases; Dianne sets boundaries regarding policy but explains fallback UX systems.38:25–44:16 · Guest teaching 6/10 Sponsor Message: Mercury Banking with Command AI Operator After an ad break, Dianne explains how PM hiring criteria remained steady while the core PM artifact shifted from PRDs to evals.44:16–46:42 · Guest teaching 5/10 Case Study: Resolving JSON Output Reliability Through Custom Evals Dianne illustrates an eval case study concerning JSON output compliance, which Lenny compares to test-driven development.46:42–50:00 · Guest teaching 4/10 Balancing Evals and PRDs for Large-Scale Alignment and Vision Lenny checks if PRDs are truly dead; Dianne confirms they remain crucial for broad multi-team alignment and exploratory, ambiguous capabilities.50:01–54:05 · Guest teaching 4/10 Hands-On AI Leadership: Why Managers Must Build and Ship Dianne emphasizes that engineering and product managers must actively build and ship to evaluate AI output effectively.54:06–58:21 · Guest teaching 3/10 Collaborative Discovery and Deep Immersion in Specific AI Use Cases Lenny and Dianne discuss community survey insights showing that going deep on narrow use cases yields more satisfaction than superficial tinkering.58:22–1:03:53 · Guest teaching 4/10 Augmenting EQ: Tactical Claude Workflows for Coaching and Thinking Dianne shares how she uses Claude skills for Crucial Conversations coaching, and addresses Lenny's query about cognitive reliance on AI.1:03:53–1:07:05 · Guest teaching 4/10 The Power of Claude's Constitution: Why Pushback Drives Quality Lenny notes Claude's distinctive conversational personality, and Dianne explains how constitutional alignment creates a genuine thinking partner through pushback.1:07:05–1:11:40 · Guest teaching 4/10 The Jagged Frontier of AI Writing and Future Improvements Lenny asks why LLMs struggle with natural prose writing; Dianne explains the jagged capability frontier and prioritizing agentic milestones.1:11:41–1:15:36 · Guest teaching 4/10 Human Judgment, Lifelong Learning, and Developing an Inner Voice Dianne highlights enduring human strengths like judgment and persistence, relating them to parenting strategies in an AI-dominated era.1:15:37–1:21:53 · Guest teaching 4/10 Sustaining High Performance: The Hive Mind and Team Resilience Dianne describes how the Anthropic team maintains resilience through collective ownership, shared coverage, and an ego-free hive mind.1:21:54–1:24:31 · Guest teaching 4/10 Closing Reflections: The Enduring Need for User-Centric PMs Dianne summarizes why deeply user-centric PMs will remain vital, which Lenny enthusiastically validates.1:24:31–1:31:20 · Guest teaching 3/10 Lightning Round: Books, Culture, and Trading Floor Lessons During the lightning round, Dianne shares recommendations and reflects on how her bond trading desk experience taught her to champion conviction and truth over hierarchy.0:00–2:32 · Guest disagreement 0/10 Episode Preview: Key Themes and High-Velocity AI Insights Lenny introduces Dianne Penn and previews the episode themes alongside brief promotional snippets from the conversation.2:36–7:45 · Guest disagreement 0/10 Anthropic's Humble Beginnings and the Golden Gate Claude Inflection Lenny recalls Anthropic's early days competing against OpenAI, while Dianne explains the culture and the Golden Gate Claude experiment.7:46–12:38 · Guest disagreement 0/10 Sponsor Message: WorkOS Enterprise Solutions Following an ad read, Lenny asks about inflection points, and Dianne breaks down how Opus 3 and long-form coding differentiation established Anthropic's position.12:38–17:22 · Guest disagreement 0/10 The Synergy of Opus 4.5 and Claude Code Lenny notes Dario Amodei's accurate coding predictions and the exponential curve, prompting Dianne to discuss adaptability and first-principles reasoning.17:22–20:02 · Guest disagreement 1/10 Discontinuous Capabilities, Scaling Laws, and Product Overhang Lenny summarizes model adaptability, and Dianne refines the explanation using scaling law graphs and discontinuous emergent capabilities.20:03–23:33 · Guest disagreement 1/10 Token Spending, Communal Experimentation, and Working in Public Lenny brings up Garry Tan's token-spending thesis, which Dianne gently reframes from raw spend input to communal experimentation output.23:33–27:30 · Guest disagreement 0/10 Anthropic Labs: Incubating Discontinuous Bets Outside the Core Roadmap Lenny inquires about Anthropic Labs' operational structure, and Dianne outlines how small pods incubate high-risk discontinuous bets.27:30–31:35 · Guest disagreement 0/10 Inside the AI Research Loop: Translating User Feedback into Evals Lenny asks what AI researchers do all day, and Dianne provides a detailed operational breakdown of translating vague user complaints into rigorous evals.31:36–35:19 · Guest disagreement 0/10 Traits of World-Class Researchers and Designing for Forward Compatibility Lenny asks about researcher traits, and Dianne describes how forward compatibility and first-principles thinking define top AI talent.35:19–38:25 · Guest disagreement 1/10 Frontier Model Safety, Safeguard Packages, and Responsible Deployment Lenny highlights increasing safety scrutiny and restricted releases; Dianne sets boundaries regarding policy but explains fallback UX systems.38:25–44:16 · Guest disagreement 0/10 Sponsor Message: Mercury Banking with Command AI Operator After an ad break, Dianne explains how PM hiring criteria remained steady while the core PM artifact shifted from PRDs to evals.44:16–46:42 · Guest disagreement 0/10 Case Study: Resolving JSON Output Reliability Through Custom Evals Dianne illustrates an eval case study concerning JSON output compliance, which Lenny compares to test-driven development.46:42–50:00 · Guest disagreement 0/10 Balancing Evals and PRDs for Large-Scale Alignment and Vision Lenny checks if PRDs are truly dead; Dianne confirms they remain crucial for broad multi-team alignment and exploratory, ambiguous capabilities.50:01–54:05 · Guest disagreement 0/10 Hands-On AI Leadership: Why Managers Must Build and Ship Dianne emphasizes that engineering and product managers must actively build and ship to evaluate AI output effectively.54:06–58:21 · Guest disagreement 0/10 Collaborative Discovery and Deep Immersion in Specific AI Use Cases Lenny and Dianne discuss community survey insights showing that going deep on narrow use cases yields more satisfaction than superficial tinkering.58:22–1:03:53 · Guest disagreement 0/10 Augmenting EQ: Tactical Claude Workflows for Coaching and Thinking Dianne shares how she uses Claude skills for Crucial Conversations coaching, and addresses Lenny's query about cognitive reliance on AI.1:03:53–1:07:05 · Guest disagreement 0/10 The Power of Claude's Constitution: Why Pushback Drives Quality Lenny notes Claude's distinctive conversational personality, and Dianne explains how constitutional alignment creates a genuine thinking partner through pushback.1:07:05–1:11:40 · Guest disagreement 0/10 The Jagged Frontier of AI Writing and Future Improvements Lenny asks why LLMs struggle with natural prose writing; Dianne explains the jagged capability frontier and prioritizing agentic milestones.1:11:41–1:15:36 · Guest disagreement 0/10 Human Judgment, Lifelong Learning, and Developing an Inner Voice Dianne highlights enduring human strengths like judgment and persistence, relating them to parenting strategies in an AI-dominated era.1:15:37–1:21:53 · Guest disagreement 0/10 Sustaining High Performance: The Hive Mind and Team Resilience Dianne describes how the Anthropic team maintains resilience through collective ownership, shared coverage, and an ego-free hive mind.1:21:54–1:24:31 · Guest disagreement 0/10 Closing Reflections: The Enduring Need for User-Centric PMs Dianne summarizes why deeply user-centric PMs will remain vital, which Lenny enthusiastically validates.1:24:31–1:31:20 · Guest disagreement 0/10 Lightning Round: Books, Culture, and Trading Floor Lessons During the lightning round, Dianne shares recommendations and reflects on how her bond trading desk experience taught her to champion conviction and truth over hierarchy.0:00–2:32 · Lenny pushing back 0/10 Episode Preview: Key Themes and High-Velocity AI Insights Lenny introduces Dianne Penn and previews the episode themes alongside brief promotional snippets from the conversation.2:36–7:45 · Lenny pushing back 1/10 Anthropic's Humble Beginnings and the Golden Gate Claude Inflection Lenny recalls Anthropic's early days competing against OpenAI, while Dianne explains the culture and the Golden Gate Claude experiment.7:46–12:38 · Lenny pushing back 0/10 Sponsor Message: WorkOS Enterprise Solutions Following an ad read, Lenny asks about inflection points, and Dianne breaks down how Opus 3 and long-form coding differentiation established Anthropic's position.12:38–17:22 · Lenny pushing back 0/10 The Synergy of Opus 4.5 and Claude Code Lenny notes Dario Amodei's accurate coding predictions and the exponential curve, prompting Dianne to discuss adaptability and first-principles reasoning.17:22–20:02 · Lenny pushing back 1/10 Discontinuous Capabilities, Scaling Laws, and Product Overhang Lenny summarizes model adaptability, and Dianne refines the explanation using scaling law graphs and discontinuous emergent capabilities.20:03–23:33 · Lenny pushing back 0/10 Token Spending, Communal Experimentation, and Working in Public Lenny brings up Garry Tan's token-spending thesis, which Dianne gently reframes from raw spend input to communal experimentation output.23:33–27:30 · Lenny pushing back 0/10 Anthropic Labs: Incubating Discontinuous Bets Outside the Core Roadmap Lenny inquires about Anthropic Labs' operational structure, and Dianne outlines how small pods incubate high-risk discontinuous bets.27:30–31:35 · Lenny pushing back 0/10 Inside the AI Research Loop: Translating User Feedback into Evals Lenny asks what AI researchers do all day, and Dianne provides a detailed operational breakdown of translating vague user complaints into rigorous evals.31:36–35:19 · Lenny pushing back 0/10 Traits of World-Class Researchers and Designing for Forward Compatibility Lenny asks about researcher traits, and Dianne describes how forward compatibility and first-principles thinking define top AI talent.35:19–38:25 · Lenny pushing back 1/10 Frontier Model Safety, Safeguard Packages, and Responsible Deployment Lenny highlights increasing safety scrutiny and restricted releases; Dianne sets boundaries regarding policy but explains fallback UX systems.38:25–44:16 · Lenny pushing back 0/10 Sponsor Message: Mercury Banking with Command AI Operator After an ad break, Dianne explains how PM hiring criteria remained steady while the core PM artifact shifted from PRDs to evals.44:16–46:42 · Lenny pushing back 0/10 Case Study: Resolving JSON Output Reliability Through Custom Evals Dianne illustrates an eval case study concerning JSON output compliance, which Lenny compares to test-driven development.46:42–50:00 · Lenny pushing back 1/10 Balancing Evals and PRDs for Large-Scale Alignment and Vision Lenny checks if PRDs are truly dead; Dianne confirms they remain crucial for broad multi-team alignment and exploratory, ambiguous capabilities.50:01–54:05 · Lenny pushing back 0/10 Hands-On AI Leadership: Why Managers Must Build and Ship Dianne emphasizes that engineering and product managers must actively build and ship to evaluate AI output effectively.54:06–58:21 · Lenny pushing back 0/10 Collaborative Discovery and Deep Immersion in Specific AI Use Cases Lenny and Dianne discuss community survey insights showing that going deep on narrow use cases yields more satisfaction than superficial tinkering.58:22–1:03:53 · Lenny pushing back 1/10 Augmenting EQ: Tactical Claude Workflows for Coaching and Thinking Dianne shares how she uses Claude skills for Crucial Conversations coaching, and addresses Lenny's query about cognitive reliance on AI.1:03:53–1:07:05 · Lenny pushing back 0/10 The Power of Claude's Constitution: Why Pushback Drives Quality Lenny notes Claude's distinctive conversational personality, and Dianne explains how constitutional alignment creates a genuine thinking partner through pushback.1:07:05–1:11:40 · Lenny pushing back 1/10 The Jagged Frontier of AI Writing and Future Improvements Lenny asks why LLMs struggle with natural prose writing; Dianne explains the jagged capability frontier and prioritizing agentic milestones.1:11:41–1:15:36 · Lenny pushing back 0/10 Human Judgment, Lifelong Learning, and Developing an Inner Voice Dianne highlights enduring human strengths like judgment and persistence, relating them to parenting strategies in an AI-dominated era.1:15:37–1:21:53 · Lenny pushing back 0/10 Sustaining High Performance: The Hive Mind and Team Resilience Dianne describes how the Anthropic team maintains resilience through collective ownership, shared coverage, and an ego-free hive mind.1:21:54–1:24:31 · Lenny pushing back 0/10 Closing Reflections: The Enduring Need for User-Centric PMs Dianne summarizes why deeply user-centric PMs will remain vital, which Lenny enthusiastically validates.1:24:31–1:31:20 · Lenny pushing back 0/10 Lightning Round: Books, Culture, and Trading Floor Lessons During the lightning round, Dianne shares recommendations and reflects on how her bond trading desk experience taught her to champion conviction and truth over hierarchy.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 61.9% · guest 38.1%0:00 · Lenny 61.9% · guest 38.1%3:00 · Lenny 24.9% · guest 75.1%3:00 · Lenny 24.9% · guest 75.1%6:00 · Lenny 41% · guest 59%6:00 · Lenny 41% · guest 59%9:00 · Lenny 5.7% · guest 94.3%9:00 · Lenny 5.7% · guest 94.3%12:00 · Lenny 54.7% · guest 45.3%12:00 · Lenny 54.7% · guest 45.3%15:00 · Lenny 19.2% · guest 80.8%15:00 · Lenny 19.2% · guest 80.8%18:00 · Lenny 25.4% · guest 74.6%18:00 · Lenny 25.4% · guest 74.6%21:00 · Lenny 23.5% · guest 76.5%21:00 · Lenny 23.5% · guest 76.5%24:00 · Lenny 15.8% · guest 84.2%24:00 · Lenny 15.8% · guest 84.2%27:00 · Lenny 21% · guest 79%27:00 · Lenny 21% · guest 79%30:00 · Lenny 21.8% · guest 78.2%30:00 · Lenny 21.8% · guest 78.2%33:00 · Lenny 33.3% · guest 66.7%33:00 · Lenny 33.3% · guest 66.7%36:00 · Lenny 39.6% · guest 60.4%36:00 · Lenny 39.6% · guest 60.4%39:00 · Lenny 40.6% · guest 59.4%39:00 · Lenny 40.6% · guest 59.4%42:00 · Lenny 18.7% · guest 81.3%42:00 · Lenny 18.7% · guest 81.3%45:00 · Lenny 17.7% · guest 82.3%45:00 · Lenny 17.7% · guest 82.3%48:00 · Lenny 22.7% · guest 77.3%48:00 · Lenny 22.7% · guest 77.3%51:00 · Lenny 31.8% · guest 68.2%51:00 · Lenny 31.8% · guest 68.2%54:00 · Lenny 36.6% · guest 63.4%54:00 · Lenny 36.6% · guest 63.4%57:00 · Lenny 18.4% · guest 81.6%57:00 · Lenny 18.4% · guest 81.6%1:00:00 · Lenny 29.4% · guest 70.6%1:00:00 · Lenny 29.4% · guest 70.6%1:03:00 · Lenny 37% · guest 63%1:03:00 · Lenny 37% · guest 63%1:06:00 · Lenny 41% · guest 59%1:06:00 · Lenny 41% · guest 59%1:09:00 · Lenny 46% · guest 54%1:09:00 · Lenny 46% · guest 54%1:12:00 · Lenny 13.9% · guest 86.1%1:12:00 · Lenny 13.9% · guest 86.1%1:15:00 · Lenny 59.5% · guest 40.5%1:15:00 · Lenny 59.5% · guest 40.5%1:18:00 · Lenny 21.6% · guest 78.4%1:18:00 · Lenny 21.6% · guest 78.4%1:21:00 · Lenny 20.3% · guest 79.7%1:21:00 · Lenny 20.3% · guest 79.7%1:24:00 · Lenny 35.9% · guest 64.1%1:24:00 · Lenny 35.9% · guest 64.1%1:27:00 · Lenny 39.8% · guest 60.2%1:27:00 · Lenny 39.8% · guest 60.2%1:30:00 · Lenny 11.6% · guest 88.4%1:30:00 · Lenny 11.6% · guest 88.4%1:33:00 · Lenny 65% · guest 35%1:33:00 · Lenny 65% · guest 35%
Sharpest disagreement ▶ 20:37 Reframing Garry Tan's token spend thesis

Dianne gently challenges the premise of spending $100k on tokens as a benchmark, arguing that token spend is merely an input while collaborative experimentation is the actual output.

Hardest push from Lenny ▶ 48:02 Pushing back on whether PRDs are dead

Lenny presses Dianne on whether PRDs are obsolete in favor of evals, prompting her to clarify where comprehensive written specs remain indispensable.

Biggest teaching moment ▶ 28:40 Deconstructing vague feedback into researcher-actionable evals

Dianne gives Lenny an operational masterclass on why telling researchers Claude hallucinated is useless, demonstrating how PMs must break failures down into tool-use, search synthesis, or alignment issues.

Lenny holds their own ▶ 46:42 Comparing eval-driven development to TDD

Lenny demonstrates sharp domain expertise by synthesizing Dianne's eval framework into the classic software paradigm of test-driven development.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Episode Preview: Key Themes and High-Velocity AI Insights 1000 Lenny introduces Dianne Penn and previews the episode themes alongside brief promotional snippets from the conversation.
Anthropic's Humble Beginnings and the Golden Gate Claude Inflection 3401 Lenny recalls Anthropic's early days competing against OpenAI, while Dianne explains the culture and the Golden Gate Claude experiment.
Sponsor Message: WorkOS Enterprise Solutions 4500 Following an ad read, Lenny asks about inflection points, and Dianne breaks down how Opus 3 and long-form coding differentiation established Anthropic's position.
The Synergy of Opus 4.5 and Claude Code 5400 Lenny notes Dario Amodei's accurate coding predictions and the exponential curve, prompting Dianne to discuss adaptability and first-principles reasoning.
Discontinuous Capabilities, Scaling Laws, and Product Overhang 5511 Lenny summarizes model adaptability, and Dianne refines the explanation using scaling law graphs and discontinuous emergent capabilities.
Token Spending, Communal Experimentation, and Working in Public 4410 Lenny brings up Garry Tan's token-spending thesis, which Dianne gently reframes from raw spend input to communal experimentation output.
Anthropic Labs: Incubating Discontinuous Bets Outside the Core Roadmap 4400 Lenny inquires about Anthropic Labs' operational structure, and Dianne outlines how small pods incubate high-risk discontinuous bets.
Inside the AI Research Loop: Translating User Feedback into Evals 4600 Lenny asks what AI researchers do all day, and Dianne provides a detailed operational breakdown of translating vague user complaints into rigorous evals.
Traits of World-Class Researchers and Designing for Forward Compatibility 4400 Lenny asks about researcher traits, and Dianne describes how forward compatibility and first-principles thinking define top AI talent.
Frontier Model Safety, Safeguard Packages, and Responsible Deployment 4311 Lenny highlights increasing safety scrutiny and restricted releases; Dianne sets boundaries regarding policy but explains fallback UX systems.
Sponsor Message: Mercury Banking with Command AI Operator 5600 After an ad break, Dianne explains how PM hiring criteria remained steady while the core PM artifact shifted from PRDs to evals.
Case Study: Resolving JSON Output Reliability Through Custom Evals 5500 Dianne illustrates an eval case study concerning JSON output compliance, which Lenny compares to test-driven development.
Balancing Evals and PRDs for Large-Scale Alignment and Vision 4401 Lenny checks if PRDs are truly dead; Dianne confirms they remain crucial for broad multi-team alignment and exploratory, ambiguous capabilities.
Hands-On AI Leadership: Why Managers Must Build and Ship 4400 Dianne emphasizes that engineering and product managers must actively build and ship to evaluate AI output effectively.
Collaborative Discovery and Deep Immersion in Specific AI Use Cases 5300 Lenny and Dianne discuss community survey insights showing that going deep on narrow use cases yields more satisfaction than superficial tinkering.
Augmenting EQ: Tactical Claude Workflows for Coaching and Thinking 4401 Dianne shares how she uses Claude skills for Crucial Conversations coaching, and addresses Lenny's query about cognitive reliance on AI.
The Power of Claude's Constitution: Why Pushback Drives Quality 5400 Lenny notes Claude's distinctive conversational personality, and Dianne explains how constitutional alignment creates a genuine thinking partner through pushback.
The Jagged Frontier of AI Writing and Future Improvements 5401 Lenny asks why LLMs struggle with natural prose writing; Dianne explains the jagged capability frontier and prioritizing agentic milestones.
Human Judgment, Lifelong Learning, and Developing an Inner Voice 4400 Dianne highlights enduring human strengths like judgment and persistence, relating them to parenting strategies in an AI-dominated era.
Sustaining High Performance: The Hive Mind and Team Resilience 4400 Dianne describes how the Anthropic team maintains resilience through collective ownership, shared coverage, and an ego-free hive mind.
Closing Reflections: The Enduring Need for User-Centric PMs 5400 Dianne summarizes why deeply user-centric PMs will remain vital, which Lenny enthusiastically validates.
Lightning Round: Books, Culture, and Trading Floor Lessons 3300 During the lightning round, Dianne shares recommendations and reflects on how her bond trading desk experience taught her to champion conviction and truth over hierarchy.

Statements from this episode (26)

Assertion Not checkable as stated
Penn: Anthropic had five product engineers and one API engineer in 2023
“So I joined in 20, 23. Like you said, we had five product engineers. There was one engineer for the entirety of our API business, if you believe.”
Dianne Penn Jul 26, 2026 ▶ 3:53
Opinion
Penn: Golden Gate Claude was Anthropic's hidden product inflection point
“That to me was like one of those, like maybe hidden inflection points of we were starting to find our identity, that we could build products, build experiences that were different for what our competitors had seen, what was already out there.”
Dianne Penn Jul 26, 2026 ▶ 6:56
Assertion Contradicted
Penn: Anthropic had fewer than 200 employees while building Claude 3 Opus
“Definitely when we were training and testing Opus three, I think that was the moment when the company, I think we were less than 200 people still at that point.”
Dianne Penn Jul 26, 2026 ▶ 9:11
Insight
Penn: Optimizing Claude 3 Opus for long-form coding drove Anthropic's early differentiation
“People are starting to use code these models, not just for code, not just like code autocomplete, but actually writing long form code. And it's that an opportunity for us to train, you know, Opus three to be better at. And it ended up being a relatively smalle…”
Dianne Penn Jul 26, 2026 ▶ 11:33
Insight
Penn: Users need frontier product harnesses to feel frontier model magic
“One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic. Of frontier models.”
Dianne Penn Jul 26, 2026 ▶ 12:57
What-if
Penn: Opus 4.5 and Claude Code depended on each other for breakout success
“I think Opus Four Five wouldn't have had that moment without a product like Cloud Code. And Cloud Code, I think, wouldn't have had that type of adoption accelerated without Opus Four Five.”
Dianne Penn Jul 26, 2026 ▶ 13:38
Insight
Penn: Scaling loss is smooth, but AI capabilities emerge discontinuously
“As you add in more compute and data, what's called loss, aka the loss from next token prediction goes down. And so it's a very smooth linear curve of, like, the models get more intelligent as you scale them up. What's actually also interesting in that paper is…”
Dianne Penn Jul 26, 2026 ▶ 17:53
Opinion
Penn: Anthropic's Opus and Fable models still have product overhang
“I think there's like product overhang and user overhang, like to maybe put it in our PM language, even on today's models. And I think there's like a lot that we could be exploring on like our current opuses and definitely with like Fable, for example”
Dianne Penn Jul 26, 2026 ▶ 19:28
Assertion Supported
Penn: Anthropic Labs incubated Claude Code, Skills, Claude Design, and MCP
“Things like cloud code things like skills, and most recently cloud design, MCP.”
Dianne Penn Jul 26, 2026 ▶ 24:28
Assertion Not checkable as stated
Penn: Anthropic discussed computer use for Claude at founding
“I think even at the founding of the company, researchers were talking about how do we get clawed to you? Use a computer. How do we get AI to, like, navigate a screen, right?”
Dianne Penn Jul 26, 2026 ▶ 28:27
Insight
Penn: Telling AI researchers Claude hallucinated is unactionable without failure-mode breakdowns
“If you bring that to a researcher and you say, please fix Claude from being hallucinated, it's not very actionable. And so part of the time of the team is understanding, okay, what's the trajectory of why that user gave that feedback?”
Dianne Penn Jul 26, 2026 ▶ 30:23
Assertion Not checkable as stated
Anthropic Research Leaders Personally Inspect Training Runs and Evals
“Our like leadership, our chief scientists, our heads of like fine tuning and like RL folks are actually really close to the training runs and actually look at things like how the training run is going, evals, looking at the underlying data.”
Dianne Penn Jul 26, 2026 ▶ 32:50
Insight
Penn: Design AI products for forward compatibility with future Claude models
“One thing I asked the team frequently or how I think about when we're building a product is let's say Claude eight comes around. What do you, what changes in what users do? And then what should, what does that mean for how you're building today? Is it going to…”
Dianne Penn Jul 26, 2026 ▶ 34:43
Disclosure
Penn: Anthropic Kept Research PM Hiring Process Unchanged for Three Years
“We actually on my team have not changed our hiring loop for three years now. So what we actually look for and the traits and how we evaluate generalists like PNs, generalists like research product managers have actually been the same.”
Dianne Penn Jul 26, 2026 ▶ 40:04
Insight
Penn: Evals Have Replaced Traditional PRDs in AI Product Management
“We actually have a saying on the team of evals are the new PRDs. Cause in order to deliver that user value it's not that exact artifact that people used to write in the last like one to two decades. It's a new way of working.”
Dianne Penn Jul 26, 2026 ▶ 41:36
Insight
Penn: AI PMs Must Sweat Tokens as Much as Pixels
“Here, you have to sweat the tokens as much as you sweat the pixels. And so one activity we have on the team is reading the transcripts. And understanding, ah, what was the trajectories that failed?”
Dianne Penn Jul 26, 2026 ▶ 42:52
Assertion Not checkable as stated
Penn: 80% of early Claude instruction failures were invalid JSON outputs
“And what I saw was something like 80% of what people meant in the early days for this failure was Claude would not write the right JSON.”
Dianne Penn Jul 26, 2026 ▶ 45:25
Assertion Not checkable as stated
Penn: Claude scores 99.9% to 100% on JSON schema formatting evals
“When we have versions of Claude, we actually run that eval and just check. I think at this point it's always a hundred percent or like 99.9. And so it's no longer a pain point.”
Dianne Penn Jul 26, 2026 ▶ 46:14
Disclosure
Anthropic creates a PRD for every AI model it develops
“So when we do have a model, we actually, for every model, we do have a PRD. Less necessarily for our researchers, but more for our growing product surfaces, for our engineering teams, for our stakeholders like legal and safety and others as just a source of tr…”
Dianne Penn Jul 26, 2026 ▶ 48:26
Disclosure
Penn: Built Claude skill using Crucial Conversations to coach management conversations
“I don't think it's necessarily just about raising the IQ of like experiences we build, but also I use it a lot and actually like prepping for how to have better conversations in the moment during like crucial conversations. So I love that book. And so I actual…”
Dianne Penn Jul 26, 2026 ▶ 59:06
Insight
Penn: Useful AI models push back on unformed ideas rather than agree
“What you don't want is like a AI that just agrees with you, right? What you want is this technology to actually augment and grow and like get to a better outcome. And so sometimes it's Having Claude push back makes me better, and so that's great. Like a cowork…”
Dianne Penn Jul 26, 2026 ▶ 1:03:23
Disclosure
Penn: Anthropic uses research Opus models to determine future Claude pricing
“So I've used Claude to help with things like, are we making the right pricing decision on the next version of Claude? It's a little bit meta, but using a research version of Opus, asking it to figure out how is your price and being able to come out with better…”
Dianne Penn Jul 26, 2026 ▶ 1:05:09
Insight
Penn: Human judgment will stay critical as AI lacks human experience
“I think judgment is one, is an area where it's accumulation of so much nuance and so much experience, and these systems haven't experienced as much as humans have. And so I think that hard earned, like, judgment Is a area for product leaders and just generally…”
Dianne Penn Jul 26, 2026 ▶ 1:12:20
Opinion
Penn: AI in biology and life sciences is only at foot of exponential
“I think, you know, software engineering has been really transformed by AI. I think there's areas like biology, life sciences. These are all things that we're just kind of at like the foot of the exponential on. Like, maybe software engineering were on the expo…”
Dianne Penn Jul 26, 2026 ▶ 1:13:30
Opinion
Penn: Highly capable AI models increase the need for user-centric PMs
“I think fundamentally, you didn't ask me this, but there is this question in the community of, do we still need PMs when the models are still capable, when engineers are leaning in I think the role of people who are user centric, who go into the details of und…”
Dianne Penn Jul 26, 2026 ▶ 1:22:59
Assertion Not checkable as stated
Penn: Anthropic UI ratings and sales feedback route directly to research PMs
“If you thumbs up or thumbs down on any of our product surfaces, if you contact your salesperson with feedback about the model, it will make its way to me. We actually, with every, like, research model, I actually get pretty close into understanding favorabilit…”
Dianne Penn Jul 26, 2026 ▶ 1:31:54
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.