Jan 11, 2026 · 1h 26m · lennys-podcast

Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon

Aishwarya Reganti (Ash) · 40m spoken Kiriti Badam · 20m spoken Lenny Rachitsky · 18m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Lenny Rachitsky interviews AI practitioners Aishwarya Reganti and Kiriti Badam to unpack proven frameworks for building production AI products, emphasizing graduated autonomy, continuous evaluation loops, and problem-first workflow design.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 24.4% of the talking time here. How this is scored →

Lenny as informed peer 3.8 Guest teaching 6.2 Guest disagreement 1.4 Lenny pushing back 1.1
05100:0020:0040:001:00:001:20:005:07–7:37 · Lenny as informed peer 4/10 The Evolution of AI Product Development in 2025 Lenny sets the stage by framing their collaborative research and asks what is working versus failing in enterprise AI. Ash educates on the shift from 2024 skepticism to 2025 execution friction, breaking down broken handoffs between PMs and engineers.7:37–11:38 · Lenny as informed peer 5/10 Two Fundamental Shifts: Non-Determinism and Agency Control Lenny cites the core takeaways from their joint post. Ash expands in technical detail on non-deterministic APIs and the agency versus control trade-off.11:39–16:01 · Lenny as informed peer 4/10 Graduating Agency: The Half Dome Analogy Kiriti illustrates the Half Dome analogy to argue against day-one complex agent architectures. Lenny synthesizes the ladder concept of starting with high control and low agency.16:02–18:36 · Lenny as informed peer 4/10 Behavior Calibration in Healthcare, Coding, and Marketing Ash explains behavior calibration and risk categorization in healthcare pre-authorizations. She demonstrates how logging human actions creates the necessary training flywheel.18:38–21:45 · Lenny as informed peer 6/10 Prioritizing the Problem Over Solution Complexity Lenny demonstrates deep grasp of the concept by reciting tiered examples across coding and marketing assistants. Kiriti reinforces the problem-first mindset.21:45–25:17 · Lenny as informed peer 5/10 Enterprise Reliability Concerns and Emerging Security Risks Ash cites Databricks research on enterprise reliability barriers. Lenny pushes on unresolved security threats like jailbreaking and prompt injection, while Kiriti offers a pragmatic counterpoint urging early adoption despite risks.25:18–29:15 · Lenny as informed peer 4/10 Leadership Vulnerability and Culture in AI Transformation Ash breaks down the success triangle of leadership, culture, and technical depth. She shares an anecdote of a CEO setting aside 4 to 6 AM daily to relearn intuitions.29:15–33:21 · Lenny as informed peer 4/10 Technical Depth and Debunking One-Click Agents Ash forcefully denounces vendors selling one-click agents as pure marketing, explaining the reality of enterprise technical debt and messy taxonomy trees.33:21–38:27 · Lenny as informed peer 4/10 The False Dichotomy: Evals Versus Production Monitoring Lenny asks about the evals versus vibes debate. Kiriti rejects the false dichotomy and thoroughly articulates how offline evals combine with granular implicit production monitoring.38:27–41:26 · Lenny as informed peer 3/10 Semantic Diffusion and Context-Dependent Evaluation Strategies Ash invokes Martin Fowler's semantic diffusion concept to critique how the industry conflates LLM judges, data labeling, and leaderboards under the blanket label of evals.41:27–44:57 · Lenny as informed peer 4/10 Evaluation Practices on OpenAI Codex Lenny asks how Codex handles evals given public skepticism from competitors. Kiriti explains Codex's balance of regression evals, A/B testing, and social feedback monitoring.44:58–57:17 · Lenny as informed peer 3/10 Sponsor Message: Brex Intelligent Finance Following a sponsor message, Ash delivers an extensive masterclass on the Continuous Calibration and Continuous Development framework, detailing customer support failure modes.57:17–1:01:24 · Lenny as informed peer 4/10 Knowing When to Recalibrate: Model Updates and User Evolution Ash explains how to detect calibration readiness by minimizing surprise and illustrates how model deprecations or evolving user expectations force full recalibration.1:01:24–1:05:15 · Lenny as informed peer 3/10 Overrated vs Underrated Trends in AI Kiriti critiques naive peer-to-peer multi-agent gossip protocols while defending coding agents as underrated. Ash argues that building is cheap while product design is underrated.1:05:16–1:08:39 · Lenny as informed peer 3/10 The Next Horizon: Proactive Background Agents and Multimodality Kiriti forecasts proactive background agents resolving tickets overnight, while Ash discusses rich multimodal context processing for messy enterprise documents.1:08:40–1:12:05 · Lenny as informed peer 4/10 Essential Skills: Taste, Judgment, and Proactive Ownership Ash emphasizes taste and judgment as implementation becomes commoditized, sharing an example of a junior hire vibe-coding internal tooling.1:12:06–1:15:14 · Lenny as informed peer 4/10 Pain as the Ultimate Competitive Moat Kiriti coins the phrase pain is the new moat, explaining how enduring trial-and-error experimentation forms defensible organizational knowledge.1:15:15–1:24:32 · Lenny as informed peer 4/10 Lightning Round: Books, Media, Productivity Hacks, and Personal Reflections The conversation shifts to a friendly lightning round covering book recommendations, sci-fi media, productivity utilities like Raycast and Caffeinate, and personal marital reflections.1:24:33–1:25:55 · Lenny as informed peer 1/10 Guest Plugs, Course Resources, and Farewell Lenny invites the guests to share their LinkedIn profiles, GitHub open source repositories, and Maven course details before closing the episode.5:07–7:37 · Guest teaching 6/10 The Evolution of AI Product Development in 2025 Lenny sets the stage by framing their collaborative research and asks what is working versus failing in enterprise AI. Ash educates on the shift from 2024 skepticism to 2025 execution friction, breaking down broken handoffs between PMs and engineers.7:37–11:38 · Guest teaching 7/10 Two Fundamental Shifts: Non-Determinism and Agency Control Lenny cites the core takeaways from their joint post. Ash expands in technical detail on non-deterministic APIs and the agency versus control trade-off.11:39–16:01 · Guest teaching 7/10 Graduating Agency: The Half Dome Analogy Kiriti illustrates the Half Dome analogy to argue against day-one complex agent architectures. Lenny synthesizes the ladder concept of starting with high control and low agency.16:02–18:36 · Guest teaching 7/10 Behavior Calibration in Healthcare, Coding, and Marketing Ash explains behavior calibration and risk categorization in healthcare pre-authorizations. She demonstrates how logging human actions creates the necessary training flywheel.18:38–21:45 · Guest teaching 5/10 Prioritizing the Problem Over Solution Complexity Lenny demonstrates deep grasp of the concept by reciting tiered examples across coding and marketing assistants. Kiriti reinforces the problem-first mindset.21:45–25:17 · Guest teaching 5/10 Enterprise Reliability Concerns and Emerging Security Risks Ash cites Databricks research on enterprise reliability barriers. Lenny pushes on unresolved security threats like jailbreaking and prompt injection, while Kiriti offers a pragmatic counterpoint urging early adoption despite risks.25:18–29:15 · Guest teaching 7/10 Leadership Vulnerability and Culture in AI Transformation Ash breaks down the success triangle of leadership, culture, and technical depth. She shares an anecdote of a CEO setting aside 4 to 6 AM daily to relearn intuitions.29:15–33:21 · Guest teaching 7/10 Technical Depth and Debunking One-Click Agents Ash forcefully denounces vendors selling one-click agents as pure marketing, explaining the reality of enterprise technical debt and messy taxonomy trees.33:21–38:27 · Guest teaching 8/10 The False Dichotomy: Evals Versus Production Monitoring Lenny asks about the evals versus vibes debate. Kiriti rejects the false dichotomy and thoroughly articulates how offline evals combine with granular implicit production monitoring.38:27–41:26 · Guest teaching 8/10 Semantic Diffusion and Context-Dependent Evaluation Strategies Ash invokes Martin Fowler's semantic diffusion concept to critique how the industry conflates LLM judges, data labeling, and leaderboards under the blanket label of evals.41:27–44:57 · Guest teaching 7/10 Evaluation Practices on OpenAI Codex Lenny asks how Codex handles evals given public skepticism from competitors. Kiriti explains Codex's balance of regression evals, A/B testing, and social feedback monitoring.44:58–57:17 · Guest teaching 8/10 Sponsor Message: Brex Intelligent Finance Following a sponsor message, Ash delivers an extensive masterclass on the Continuous Calibration and Continuous Development framework, detailing customer support failure modes.57:17–1:01:24 · Guest teaching 7/10 Knowing When to Recalibrate: Model Updates and User Evolution Ash explains how to detect calibration readiness by minimizing surprise and illustrates how model deprecations or evolving user expectations force full recalibration.1:01:24–1:05:15 · Guest teaching 6/10 Overrated vs Underrated Trends in AI Kiriti critiques naive peer-to-peer multi-agent gossip protocols while defending coding agents as underrated. Ash argues that building is cheap while product design is underrated.1:05:16–1:08:39 · Guest teaching 6/10 The Next Horizon: Proactive Background Agents and Multimodality Kiriti forecasts proactive background agents resolving tickets overnight, while Ash discusses rich multimodal context processing for messy enterprise documents.1:08:40–1:12:05 · Guest teaching 6/10 Essential Skills: Taste, Judgment, and Proactive Ownership Ash emphasizes taste and judgment as implementation becomes commoditized, sharing an example of a junior hire vibe-coding internal tooling.1:12:06–1:15:14 · Guest teaching 7/10 Pain as the Ultimate Competitive Moat Kiriti coins the phrase pain is the new moat, explaining how enduring trial-and-error experimentation forms defensible organizational knowledge.1:15:15–1:24:32 · Guest teaching 3/10 Lightning Round: Books, Media, Productivity Hacks, and Personal Reflections The conversation shifts to a friendly lightning round covering book recommendations, sci-fi media, productivity utilities like Raycast and Caffeinate, and personal marital reflections.1:24:33–1:25:55 · Guest teaching 1/10 Guest Plugs, Course Resources, and Farewell Lenny invites the guests to share their LinkedIn profiles, GitHub open source repositories, and Maven course details before closing the episode.5:07–7:37 · Guest disagreement 1/10 The Evolution of AI Product Development in 2025 Lenny sets the stage by framing their collaborative research and asks what is working versus failing in enterprise AI. Ash educates on the shift from 2024 skepticism to 2025 execution friction, breaking down broken handoffs between PMs and engineers.7:37–11:38 · Guest disagreement 1/10 Two Fundamental Shifts: Non-Determinism and Agency Control Lenny cites the core takeaways from their joint post. Ash expands in technical detail on non-deterministic APIs and the agency versus control trade-off.11:39–16:01 · Guest disagreement 1/10 Graduating Agency: The Half Dome Analogy Kiriti illustrates the Half Dome analogy to argue against day-one complex agent architectures. Lenny synthesizes the ladder concept of starting with high control and low agency.16:02–18:36 · Guest disagreement 1/10 Behavior Calibration in Healthcare, Coding, and Marketing Ash explains behavior calibration and risk categorization in healthcare pre-authorizations. She demonstrates how logging human actions creates the necessary training flywheel.18:38–21:45 · Guest disagreement 1/10 Prioritizing the Problem Over Solution Complexity Lenny demonstrates deep grasp of the concept by reciting tiered examples across coding and marketing assistants. Kiriti reinforces the problem-first mindset.21:45–25:17 · Guest disagreement 3/10 Enterprise Reliability Concerns and Emerging Security Risks Ash cites Databricks research on enterprise reliability barriers. Lenny pushes on unresolved security threats like jailbreaking and prompt injection, while Kiriti offers a pragmatic counterpoint urging early adoption despite risks.25:18–29:15 · Guest disagreement 1/10 Leadership Vulnerability and Culture in AI Transformation Ash breaks down the success triangle of leadership, culture, and technical depth. She shares an anecdote of a CEO setting aside 4 to 6 AM daily to relearn intuitions.29:15–33:21 · Guest disagreement 4/10 Technical Depth and Debunking One-Click Agents Ash forcefully denounces vendors selling one-click agents as pure marketing, explaining the reality of enterprise technical debt and messy taxonomy trees.33:21–38:27 · Guest disagreement 2/10 The False Dichotomy: Evals Versus Production Monitoring Lenny asks about the evals versus vibes debate. Kiriti rejects the false dichotomy and thoroughly articulates how offline evals combine with granular implicit production monitoring.38:27–41:26 · Guest disagreement 2/10 Semantic Diffusion and Context-Dependent Evaluation Strategies Ash invokes Martin Fowler's semantic diffusion concept to critique how the industry conflates LLM judges, data labeling, and leaderboards under the blanket label of evals.41:27–44:57 · Guest disagreement 1/10 Evaluation Practices on OpenAI Codex Lenny asks how Codex handles evals given public skepticism from competitors. Kiriti explains Codex's balance of regression evals, A/B testing, and social feedback monitoring.44:58–57:17 · Guest disagreement 1/10 Sponsor Message: Brex Intelligent Finance Following a sponsor message, Ash delivers an extensive masterclass on the Continuous Calibration and Continuous Development framework, detailing customer support failure modes.57:17–1:01:24 · Guest disagreement 1/10 Knowing When to Recalibrate: Model Updates and User Evolution Ash explains how to detect calibration readiness by minimizing surprise and illustrates how model deprecations or evolving user expectations force full recalibration.1:01:24–1:05:15 · Guest disagreement 2/10 Overrated vs Underrated Trends in AI Kiriti critiques naive peer-to-peer multi-agent gossip protocols while defending coding agents as underrated. Ash argues that building is cheap while product design is underrated.1:05:16–1:08:39 · Guest disagreement 1/10 The Next Horizon: Proactive Background Agents and Multimodality Kiriti forecasts proactive background agents resolving tickets overnight, while Ash discusses rich multimodal context processing for messy enterprise documents.1:08:40–1:12:05 · Guest disagreement 1/10 Essential Skills: Taste, Judgment, and Proactive Ownership Ash emphasizes taste and judgment as implementation becomes commoditized, sharing an example of a junior hire vibe-coding internal tooling.1:12:06–1:15:14 · Guest disagreement 1/10 Pain as the Ultimate Competitive Moat Kiriti coins the phrase pain is the new moat, explaining how enduring trial-and-error experimentation forms defensible organizational knowledge.1:15:15–1:24:32 · Guest disagreement 1/10 Lightning Round: Books, Media, Productivity Hacks, and Personal Reflections The conversation shifts to a friendly lightning round covering book recommendations, sci-fi media, productivity utilities like Raycast and Caffeinate, and personal marital reflections.1:24:33–1:25:55 · Guest disagreement 0/10 Guest Plugs, Course Resources, and Farewell Lenny invites the guests to share their LinkedIn profiles, GitHub open source repositories, and Maven course details before closing the episode.5:07–7:37 · Lenny pushing back 1/10 The Evolution of AI Product Development in 2025 Lenny sets the stage by framing their collaborative research and asks what is working versus failing in enterprise AI. Ash educates on the shift from 2024 skepticism to 2025 execution friction, breaking down broken handoffs between PMs and engineers.7:37–11:38 · Lenny pushing back 1/10 Two Fundamental Shifts: Non-Determinism and Agency Control Lenny cites the core takeaways from their joint post. Ash expands in technical detail on non-deterministic APIs and the agency versus control trade-off.11:39–16:01 · Lenny pushing back 1/10 Graduating Agency: The Half Dome Analogy Kiriti illustrates the Half Dome analogy to argue against day-one complex agent architectures. Lenny synthesizes the ladder concept of starting with high control and low agency.16:02–18:36 · Lenny pushing back 1/10 Behavior Calibration in Healthcare, Coding, and Marketing Ash explains behavior calibration and risk categorization in healthcare pre-authorizations. She demonstrates how logging human actions creates the necessary training flywheel.18:38–21:45 · Lenny pushing back 1/10 Prioritizing the Problem Over Solution Complexity Lenny demonstrates deep grasp of the concept by reciting tiered examples across coding and marketing assistants. Kiriti reinforces the problem-first mindset.21:45–25:17 · Lenny pushing back 4/10 Enterprise Reliability Concerns and Emerging Security Risks Ash cites Databricks research on enterprise reliability barriers. Lenny pushes on unresolved security threats like jailbreaking and prompt injection, while Kiriti offers a pragmatic counterpoint urging early adoption despite risks.25:18–29:15 · Lenny pushing back 1/10 Leadership Vulnerability and Culture in AI Transformation Ash breaks down the success triangle of leadership, culture, and technical depth. She shares an anecdote of a CEO setting aside 4 to 6 AM daily to relearn intuitions.29:15–33:21 · Lenny pushing back 1/10 Technical Depth and Debunking One-Click Agents Ash forcefully denounces vendors selling one-click agents as pure marketing, explaining the reality of enterprise technical debt and messy taxonomy trees.33:21–38:27 · Lenny pushing back 1/10 The False Dichotomy: Evals Versus Production Monitoring Lenny asks about the evals versus vibes debate. Kiriti rejects the false dichotomy and thoroughly articulates how offline evals combine with granular implicit production monitoring.38:27–41:26 · Lenny pushing back 1/10 Semantic Diffusion and Context-Dependent Evaluation Strategies Ash invokes Martin Fowler's semantic diffusion concept to critique how the industry conflates LLM judges, data labeling, and leaderboards under the blanket label of evals.41:27–44:57 · Lenny pushing back 1/10 Evaluation Practices on OpenAI Codex Lenny asks how Codex handles evals given public skepticism from competitors. Kiriti explains Codex's balance of regression evals, A/B testing, and social feedback monitoring.44:58–57:17 · Lenny pushing back 1/10 Sponsor Message: Brex Intelligent Finance Following a sponsor message, Ash delivers an extensive masterclass on the Continuous Calibration and Continuous Development framework, detailing customer support failure modes.57:17–1:01:24 · Lenny pushing back 1/10 Knowing When to Recalibrate: Model Updates and User Evolution Ash explains how to detect calibration readiness by minimizing surprise and illustrates how model deprecations or evolving user expectations force full recalibration.1:01:24–1:05:15 · Lenny pushing back 1/10 Overrated vs Underrated Trends in AI Kiriti critiques naive peer-to-peer multi-agent gossip protocols while defending coding agents as underrated. Ash argues that building is cheap while product design is underrated.1:05:16–1:08:39 · Lenny pushing back 1/10 The Next Horizon: Proactive Background Agents and Multimodality Kiriti forecasts proactive background agents resolving tickets overnight, while Ash discusses rich multimodal context processing for messy enterprise documents.1:08:40–1:12:05 · Lenny pushing back 1/10 Essential Skills: Taste, Judgment, and Proactive Ownership Ash emphasizes taste and judgment as implementation becomes commoditized, sharing an example of a junior hire vibe-coding internal tooling.1:12:06–1:15:14 · Lenny pushing back 1/10 Pain as the Ultimate Competitive Moat Kiriti coins the phrase pain is the new moat, explaining how enduring trial-and-error experimentation forms defensible organizational knowledge.1:15:15–1:24:32 · Lenny pushing back 1/10 Lightning Round: Books, Media, Productivity Hacks, and Personal Reflections The conversation shifts to a friendly lightning round covering book recommendations, sci-fi media, productivity utilities like Raycast and Caffeinate, and personal marital reflections.1:24:33–1:25:55 · Lenny pushing back 0/10 Guest Plugs, Course Resources, and Farewell Lenny invites the guests to share their LinkedIn profiles, GitHub open source repositories, and Maven course details before closing the episode.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 60.8% · guest 39.2%0:00 · Lenny 60.8% · guest 39.2%3:00 · Lenny 95.6% · guest 4.4%3:00 · Lenny 95.6% · guest 4.4%6:00 · Lenny 12.8% · guest 87.2%6:00 · Lenny 12.8% · guest 87.2%9:00 · Lenny 27.3% · guest 72.7%9:00 · Lenny 27.3% · guest 72.7%12:00 · Lenny 10% · guest 90%12:00 · Lenny 10% · guest 90%15:00 · Lenny 31.5% · guest 68.5%15:00 · Lenny 31.5% · guest 68.5%18:00 · Lenny 50.5% · guest 49.5%18:00 · Lenny 50.5% · guest 49.5%21:00 · Lenny 36.2% · guest 63.8%21:00 · Lenny 36.2% · guest 63.8%24:00 · Lenny 20.2% · guest 79.8%24:00 · Lenny 20.2% · guest 79.8%27:00 · Lenny 0% · guest 100%27:00 · Lenny 0% · guest 100%30:00 · Lenny 20.4% · guest 79.6%30:00 · Lenny 20.4% · guest 79.6%33:00 · Lenny 14.3% · guest 85.7%33:00 · Lenny 14.3% · guest 85.7%36:00 · Lenny 7.2% · guest 92.8%36:00 · Lenny 7.2% · guest 92.8%39:00 · Lenny 21.4% · guest 78.6%39:00 · Lenny 21.4% · guest 78.6%42:00 · Lenny 9.8% · guest 90.2%42:00 · Lenny 9.8% · guest 90.2%45:00 · Lenny 43.2% · guest 56.8%45:00 · Lenny 43.2% · guest 56.8%48:00 · Lenny 0% · guest 100%48:00 · Lenny 0% · guest 100%51:00 · Lenny 0% · guest 100%51:00 · Lenny 0% · guest 100%54:00 · Lenny 0% · guest 100%54:00 · Lenny 0% · guest 100%57:00 · Lenny 34% · guest 66%57:00 · Lenny 34% · guest 66%1:00:00 · Lenny 4.2% · guest 95.8%1:00:00 · Lenny 4.2% · guest 95.8%1:03:00 · Lenny 21.1% · guest 78.9%1:03:00 · Lenny 21.1% · guest 78.9%1:06:00 · Lenny 28.5% · guest 71.5%1:06:00 · Lenny 28.5% · guest 71.5%1:09:00 · Lenny 17.7% · guest 82.3%1:09:00 · Lenny 17.7% · guest 82.3%1:12:00 · Lenny 19.5% · guest 80.5%1:12:00 · Lenny 19.5% · guest 80.5%1:15:00 · Lenny 29.5% · guest 70.5%1:15:00 · Lenny 29.5% · guest 70.5%1:18:00 · Lenny 41.1% · guest 58.9%1:18:00 · Lenny 41.1% · guest 58.9%1:21:00 · Lenny 23.7% · guest 76.3%1:21:00 · Lenny 23.7% · guest 76.3%1:24:00 · Lenny 31.3% · guest 68.7%1:24:00 · Lenny 31.3% · guest 68.7%
Sharpest disagreement ▶ 31:20 Ash dismisses one-click enterprise agents as pure marketing

Ash forcefully calls out vendor hype, stating directly that one-click autonomous deployment in messy enterprise infrastructure is impossible and pure marketing.

Hardest push from Lenny ▶ 23:28 Lenny challenges the guest's optimism by raising prompt injection risks

Lenny refuses to let the conversation gloss over major security risks, pressing Kiriti on how easily guardrails are bypassed and whether prompt injection remains fundamentally unsolved.

Biggest teaching moment ▶ 39:20 Ash breaks down semantic diffusion around AI evals

Ash reframes the entire industry debate by explaining semantic diffusion, showing how teams mistakenly treat leaderboards and labeling notes as complete evaluation systems.

Lenny holds their own ▶ 17:41 Lenny lays out multi-stage agent development progressions

Lenny demonstrates deep command of product architecture by walking through concrete V1-to-V3 development ladders across coding assistants and automated marketing systems.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
The Evolution of AI Product Development in 2025 4611 Lenny sets the stage by framing their collaborative research and asks what is working versus failing in enterprise AI. Ash educates on the shift from 2024 skepticism to 2025 execution friction, breaking down broken handoffs between PMs and engineers.
Two Fundamental Shifts: Non-Determinism and Agency Control 5711 Lenny cites the core takeaways from their joint post. Ash expands in technical detail on non-deterministic APIs and the agency versus control trade-off.
Graduating Agency: The Half Dome Analogy 4711 Kiriti illustrates the Half Dome analogy to argue against day-one complex agent architectures. Lenny synthesizes the ladder concept of starting with high control and low agency.
Behavior Calibration in Healthcare, Coding, and Marketing 4711 Ash explains behavior calibration and risk categorization in healthcare pre-authorizations. She demonstrates how logging human actions creates the necessary training flywheel.
Prioritizing the Problem Over Solution Complexity 6511 Lenny demonstrates deep grasp of the concept by reciting tiered examples across coding and marketing assistants. Kiriti reinforces the problem-first mindset.
Enterprise Reliability Concerns and Emerging Security Risks 5534 Ash cites Databricks research on enterprise reliability barriers. Lenny pushes on unresolved security threats like jailbreaking and prompt injection, while Kiriti offers a pragmatic counterpoint urging early adoption despite risks.
Leadership Vulnerability and Culture in AI Transformation 4711 Ash breaks down the success triangle of leadership, culture, and technical depth. She shares an anecdote of a CEO setting aside 4 to 6 AM daily to relearn intuitions.
Technical Depth and Debunking One-Click Agents 4741 Ash forcefully denounces vendors selling one-click agents as pure marketing, explaining the reality of enterprise technical debt and messy taxonomy trees.
The False Dichotomy: Evals Versus Production Monitoring 4821 Lenny asks about the evals versus vibes debate. Kiriti rejects the false dichotomy and thoroughly articulates how offline evals combine with granular implicit production monitoring.
Semantic Diffusion and Context-Dependent Evaluation Strategies 3821 Ash invokes Martin Fowler's semantic diffusion concept to critique how the industry conflates LLM judges, data labeling, and leaderboards under the blanket label of evals.
Evaluation Practices on OpenAI Codex 4711 Lenny asks how Codex handles evals given public skepticism from competitors. Kiriti explains Codex's balance of regression evals, A/B testing, and social feedback monitoring.
Sponsor Message: Brex Intelligent Finance 3811 Following a sponsor message, Ash delivers an extensive masterclass on the Continuous Calibration and Continuous Development framework, detailing customer support failure modes.
Knowing When to Recalibrate: Model Updates and User Evolution 4711 Ash explains how to detect calibration readiness by minimizing surprise and illustrates how model deprecations or evolving user expectations force full recalibration.
Overrated vs Underrated Trends in AI 3621 Kiriti critiques naive peer-to-peer multi-agent gossip protocols while defending coding agents as underrated. Ash argues that building is cheap while product design is underrated.
The Next Horizon: Proactive Background Agents and Multimodality 3611 Kiriti forecasts proactive background agents resolving tickets overnight, while Ash discusses rich multimodal context processing for messy enterprise documents.
Essential Skills: Taste, Judgment, and Proactive Ownership 4611 Ash emphasizes taste and judgment as implementation becomes commoditized, sharing an example of a junior hire vibe-coding internal tooling.
Pain as the Ultimate Competitive Moat 4711 Kiriti coins the phrase pain is the new moat, explaining how enduring trial-and-error experimentation forms defensible organizational knowledge.
Lightning Round: Books, Media, Productivity Hacks, and Personal Reflections 4311 The conversation shifts to a friendly lightning round covering book recommendations, sci-fi media, productivity utilities like Raycast and Caffeinate, and personal marital reflections.
Guest Plugs, Course Resources, and Farewell 1100 Lenny invites the guests to share their LinkedIn profiles, GitHub open source repositories, and Maven course details before closing the episode.

Statements from this episode (37)

Insight
Reganti: Successful AI products require rethinking workflows, not just adding chat
“A lot of the use cases that I saw last year were more of slap chat on your data, right? And that was, you know calling themselves an AI product. And this year, a ton of companies are really rethinking their user experiences. And their workflows and all of that…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 6:10
Insight
Reganti: AI development breaks traditional handoffs between PMs and engineers
“The AI lifecycle both Pre-deployment and post-deployment is very different as compared to a traditional software life cycle. And so, so a lot of old contracts and handoffs between traditional roles, like say PMs and engineers and data folks has now been broken…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 6:48
Insight
Reganti: AI product development involves non-determinism across input, process, and output
“You don't know how the user might behave with your product, and you also don't know how the LLM might respond to that, so you're now working with an input, output, and a process, and you don't understand all the three very well.”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 9:37
Insight
Reganti: Granting AI agent autonomy inherently demands sacrificing operational control
“Every time you hand over decision-making capabilities or autonomy to agentic systems, you're kind of relinquishing some amount of control on your end trade, and when you do that, you want to make sure that your agent has Gained your trust, or it is reliable en…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 10:07
Insight
Badam: AI products should start with high human control before granting agency
“When you don't start with like agents with all the tools and all the context that you have in the company in day one and expect it to work or like, you know, even tinker at that level, you need to be deliberately starting in places where there is minimal impac…”
Kiriti Badam Jan 11, 2026 ▶ 12:05
Assertion Not checkable as stated
Badam: OpenAI Faced Major Support Ticket Spikes During New Product Launches
“Open AI face to say exact same thing when we were launching products and there was like a huge spike of support volume as like, you know, we launched successful products like image and you know, like GPT five and things like that.”
Kiriti Badam Jan 11, 2026 ▶ 13:46
Opinion
Reganti: Insurance pre-authorization is a prime use case for AI
“Insurance pre-authorization is a very ripe use case for AI because clinicians spend a lot of time pre-authorizing things like blood tests, MRIs, and things like that, right?”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 16:42
Insight
Reganti: AI challenges stem from achieving deterministic outcomes with non-deterministic tech
“You want to make sure that that intent is rightly communicated and the right actions are taken because most of your systems are deterministic, and you want to achieve a deterministic outcome, but with non-deterministic technology, and that's where it gets a li…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 19:31
Insight
Badam: Starting AI with low agency forces problem-first focus
“When you start small, and when you start with a building, like a, Very minimalistic version with high human control and low agency. It also forces you to think about what is the problem that I'm going to solve. We use this term called problem first, and to me,…”
Kiriti Badam Jan 11, 2026 ▶ 20:36
Assertion Partly supported
Reganti: 75% of enterprises cite reliability as main barrier to customer AI
“It said about 74 or 75% of the enterprises that they had spoken to their biggest problem was reliability, and that's also why they weren't comfortable deploying products to their end users and building customer-facing products”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 21:51
Prediction Not checkable as stated
Reganti: Prompt injection will become a major crisis as AI goes mainstream
“I think that will be a huge problem once systems go mainstream. We're still so busy building AI products that we're not worried about security, but it will be such a huge problem to kind of especially with this non-deterministic API again, right? So you're kin…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 23:03
Assertion Not checkable as stated
Badam: Every enterprise client at OpenAI found viable AI use cases
“No company at OpenAI we talk to is, has never had been the case that, oh, AI cannot help me in this case. It has always been that, oh, there is this, like, set of things that it can optimize for me, and then let me see how I can adopt it.”
Kiriti Badam Jan 11, 2026 ▶ 24:52
Assertion Not checkable as stated
Reganti: Rackspace CEO blocked 4-6 AM daily for AI learning
“I used to work with the CEO of now Rackspace, Gajen, so he would have this block every day in the morning, which would say catching up with AI four to six a.m., And he would not have any meetings or anything like that, and that was just his time to pick up on …”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 26:28
Insight
Reganti: Enterprise AI transformation cannot succeed as a bottom-up initiative
“It's almost always impossible for it to be bottom-up. You can't have a bunch of engineers go and get buy-in from the leader if they just don't trust in the technology, or if they have misaligned expectations about the technology, right?”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 27:22
Insight
Reganti: Workflow Automation Always Requires Combining ML Models and Deterministic Code
“Whenever you're trying to automate some part of a workflow, it's never the case that you could use an AI agent and that will kind of solve your problems, right? It's always, you probably have a machine learning model that's going to do some part of the job. Yo…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 29:36
Opinion
Reganti: Vendors Selling 'One-Click AI Agents' Are Pushing Pure Marketing
“I probably will go as far to say that if someone's selling you one click agents, it's pure marketing. You don't want to buy into that.”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 31:36
Insight
Reganti: Enterprise AI Workflows Require 4 to 6 Months for Real ROI
“To replace any critical workflow or to build something that can give you significant ROI easily takes four to six months of work, even if you have the best data layer and infrastructure layer.”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 31:53
Insight
Badam: AI agent monitoring requires tracking granular implicit feedback
“And this production monitoring has existed for products like for a long time, just that now with AI agents, you need to be monitoring like a lot more granularity. It's not just the customer always giving you explicit feedback, but there is many implicit feedba…”
Kiriti Badam Jan 11, 2026 ▶ 34:57
Insight
Badam: Relying solely on either evals or production monitoring is inadequate
“So I feel devals are important. Production monitoring is important, but this notion of only one of them is going to solve things for you. That is completely dismissible in my opinion.”
Kiriti Badam Jan 11, 2026 ▶ 37:49
Insight
Reganti: Terms Like Evals and Agents Suffer From Semantic Diffusion
“I think Martin Fowler at some point had this term called semantic diffusion back in The 2000 which kind of means that someone comes up with a term, everybody starts butchering it with their own definitions, and then you kind of lose the actual definition of it…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 39:31
Insight
Reganti: LLM judges fail in complex AI cases due to emerging patterns
“When you go to complex use cases, it's incredibly hard to build LLM judges because you see a lot of emerging patterns. If you build a judge that would you know, test for verbosity or something like that, it turns out that you're seeing newer patterns that your…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 40:07
Insight
Badam: Coding agents cannot rely solely on offline evaluation datasets
“Coding agents are extremely unique compared to agents for other domains in the sense that These are actually built for customizability and these are built for engineers. So coding agent is not a product which is going to solve like this top five workflows or l…”
Kiriti Badam Jan 11, 2026 ▶ 42:03
Insight
Badam: Relying entirely on fixed evals without team testing fails
“I don't think like if anybody's coming and seeing that, like my, I have this Concrete set of evals that I can, like, bet my life on, and then I don't need to think about anything else. Like, it's not going to work, and every new model that we're going to launc…”
Kiriti Badam Jan 11, 2026 ▶ 44:25
Disclosure
Reganti: Team shut down their AI support product due to endless hotfixes
“At a time we were building for a customer support. Use case, which is what, which is the example that we give in the newsletter as well, and we had to shut down the product because we were doing so many hot fixes, and there was no way we could count all the em…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 47:00
Insight
Reganti: Predefined AI evals only catch anticipated failure patterns
“The issue with just building a bunch of evaluation metrics and then having them in production is evaluation metrics catch only the errors that you're already aware of, but there can be a lot of emerging patterns that you understand. Only after you put things i…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 56:01
Insight
Reganti: Advance AI autonomy when monitoring yields minimal new surprise
“There's not really a rule book you can follow, but it's all about minimizing surprise, which means let's say you're calibrating every one or two days and you figured out that you're not seeing new data distribution patterns. Your users have been pretty consist…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 58:28
Insight
Reganti: AI systems need recalibration as evolving user behavior increases complexity
“Something that might seem very natural to an end user might be very hard to build as a product builder, and you see that user behavior also evolves over time, and that's when you know that you want to go back and recalibrate.”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 1:01:11
Insight
Badam: Peer-to-peer multi-agent gossip protocols exceed current AI model capabilities
“If you're building a supervisor agent and there are like sub agents that actually do the work for the super agent, supervisor agent, That is a very successful pattern, but coming with this notion of I'm going to divide the responsibilities based on functionali…”
Kiriti Badam Jan 11, 2026 ▶ 1:02:14
Prediction Not checkable as stated
Badam: 2025 and 2026 will unlock massive value from coding agents
“Talking to an engineer in like any random company especially outside of Bay Area, you can see like the amount of impact this coding agents can create and the penetration is very low. So I feel like 20, 25 and 2026 is going to be like an incredible year for opt…”
Kiriti Badam Jan 11, 2026 ▶ 1:02:55
Insight
Reganti: Cheap AI building makes product design and problem focus far more valuable
“Building is really cheap today. Design is more expensive. Really thinking about your product, what you're going to build. Is it going to really solve a pain point? Is, is what is way more valuable today? And it will only become more true in the near future, ri…”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 1:04:54
Insight
Badam: AI fails to create value because it lacks workflow context
“Where is AI failing to create value today, it's mainly about not understanding the context. And the reason that it's not understanding the context is it's not plugged into the right places where actual work is happening.”
Kiriti Badam Jan 11, 2026 ▶ 1:05:42
Prediction Not checkable as stated
Badam: Future AI products will be proactive background agents
“And now when you extend this to more complex tasks, like a coding agent, which says that like, okay, I have fixed five of your linear tickets and here are the patches just to review them at the start of your day. So I feel that is going to be like extremely us…”
Kiriti Badam Jan 11, 2026 ▶ 1:06:28
Assertion Not checkable as stated
Reganti: Frontier AI models still cannot parse messy PDFs and handwriting
“There are so many handwritten documents and really messy PDFs that cannot be passed even by the best of the models as of today.”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 1:08:07
Assertion Supported
Rachitsky: Jason Lemkin replaced a 10-person sales team with 20 AI agents
“I just had Jason Lemkin on the podcast. He's very smart on sales, go to market, run Saster, and he replaced his whole sales team with agents. He had 10 salespeople. Now he has 1.2 and 20 agents.”
Lenny Rachitsky Jan 11, 2026 ▶ 1:11:29
Insight
Badam: Enduring iterative implementation pain is the new moat for AI builders
“And as you are going through this, like pain of like developing multiple approaches and then solving the problem, I feel that is like going to be the real boat as an individual. Like I like to call it like pain is the new moat, but I feel that is exactly super…”
Kiriti Badam Jan 11, 2026 ▶ 1:12:36
Insight
Badam: Organizational knowledge from trial-and-error evals is the decisive AI moat
“And that kind of knowledge that you've built across the organization or across like your own experience, lived experiences. I feel that the, that pain is what translates into the mode of the company, right? This could be like a product of evals or like somethi…”
Kiriti Badam Jan 11, 2026 ▶ 1:13:41
Insight
Reganti: 80% of AI engineering is understanding workflows, not complex models
“80% of so-called AI engineers, AI PMs spend their time actually understanding their workflows very well. They're not building the fanciest and the, you know, most cool models or workflows around it. They're actually in the weeds understanding their customers' …”
Aishwarya Reganti (Ash) Jan 11, 2026 ▶ 1:14:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.