Jan 11, 2026 · 1h 26m · lennys-podcast
Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Lenny Rachitsky interviews AI practitioners Aishwarya Reganti and Kiriti Badam to unpack proven frameworks for building production AI products, emphasizing graduated autonomy, continuous evaluation loops, and problem-first workflow design.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 24.4% of the talking time here. How this is scored →
speaking balance: gold is Lenny, purple is the guest (3 minute bins)
Ash forcefully calls out vendor hype, stating directly that one-click autonomous deployment in messy enterprise infrastructure is impossible and pure marketing.
Hardest push from Lenny ▶ 23:28 Lenny challenges the guest's optimism by raising prompt injection risksLenny refuses to let the conversation gloss over major security risks, pressing Kiriti on how easily guardrails are bypassed and whether prompt injection remains fundamentally unsolved.
Biggest teaching moment ▶ 39:20 Ash breaks down semantic diffusion around AI evalsAsh reframes the entire industry debate by explaining semantic diffusion, showing how teams mistakenly treat leaderboards and labeling notes as complete evaluation systems.
Lenny holds their own ▶ 17:41 Lenny lays out multi-stage agent development progressionsLenny demonstrates deep command of product architecture by walking through concrete V1-to-V3 development ladders across coding assistants and automated marketing systems.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Lenny as informed peer | Guest teaching | Guest disagreement | Lenny pushing back | Why |
|---|---|---|---|---|---|---|
| The Evolution of AI Product Development in 2025 | 4 | 6 | 1 | 1 | Lenny sets the stage by framing their collaborative research and asks what is working versus failing in enterprise AI. Ash educates on the shift from 2024 skepticism to 2025 execution friction, breaking down broken handoffs between PMs and engineers. | |
| Two Fundamental Shifts: Non-Determinism and Agency Control | 5 | 7 | 1 | 1 | Lenny cites the core takeaways from their joint post. Ash expands in technical detail on non-deterministic APIs and the agency versus control trade-off. | |
| Graduating Agency: The Half Dome Analogy | 4 | 7 | 1 | 1 | Kiriti illustrates the Half Dome analogy to argue against day-one complex agent architectures. Lenny synthesizes the ladder concept of starting with high control and low agency. | |
| Behavior Calibration in Healthcare, Coding, and Marketing | 4 | 7 | 1 | 1 | Ash explains behavior calibration and risk categorization in healthcare pre-authorizations. She demonstrates how logging human actions creates the necessary training flywheel. | |
| Prioritizing the Problem Over Solution Complexity | 6 | 5 | 1 | 1 | Lenny demonstrates deep grasp of the concept by reciting tiered examples across coding and marketing assistants. Kiriti reinforces the problem-first mindset. | |
| Enterprise Reliability Concerns and Emerging Security Risks | 5 | 5 | 3 | 4 | Ash cites Databricks research on enterprise reliability barriers. Lenny pushes on unresolved security threats like jailbreaking and prompt injection, while Kiriti offers a pragmatic counterpoint urging early adoption despite risks. | |
| Leadership Vulnerability and Culture in AI Transformation | 4 | 7 | 1 | 1 | Ash breaks down the success triangle of leadership, culture, and technical depth. She shares an anecdote of a CEO setting aside 4 to 6 AM daily to relearn intuitions. | |
| Technical Depth and Debunking One-Click Agents | 4 | 7 | 4 | 1 | Ash forcefully denounces vendors selling one-click agents as pure marketing, explaining the reality of enterprise technical debt and messy taxonomy trees. | |
| The False Dichotomy: Evals Versus Production Monitoring | 4 | 8 | 2 | 1 | Lenny asks about the evals versus vibes debate. Kiriti rejects the false dichotomy and thoroughly articulates how offline evals combine with granular implicit production monitoring. | |
| Semantic Diffusion and Context-Dependent Evaluation Strategies | 3 | 8 | 2 | 1 | Ash invokes Martin Fowler's semantic diffusion concept to critique how the industry conflates LLM judges, data labeling, and leaderboards under the blanket label of evals. | |
| Evaluation Practices on OpenAI Codex | 4 | 7 | 1 | 1 | Lenny asks how Codex handles evals given public skepticism from competitors. Kiriti explains Codex's balance of regression evals, A/B testing, and social feedback monitoring. | |
| Sponsor Message: Brex Intelligent Finance | 3 | 8 | 1 | 1 | Following a sponsor message, Ash delivers an extensive masterclass on the Continuous Calibration and Continuous Development framework, detailing customer support failure modes. | |
| Knowing When to Recalibrate: Model Updates and User Evolution | 4 | 7 | 1 | 1 | Ash explains how to detect calibration readiness by minimizing surprise and illustrates how model deprecations or evolving user expectations force full recalibration. | |
| Overrated vs Underrated Trends in AI | 3 | 6 | 2 | 1 | Kiriti critiques naive peer-to-peer multi-agent gossip protocols while defending coding agents as underrated. Ash argues that building is cheap while product design is underrated. | |
| The Next Horizon: Proactive Background Agents and Multimodality | 3 | 6 | 1 | 1 | Kiriti forecasts proactive background agents resolving tickets overnight, while Ash discusses rich multimodal context processing for messy enterprise documents. | |
| Essential Skills: Taste, Judgment, and Proactive Ownership | 4 | 6 | 1 | 1 | Ash emphasizes taste and judgment as implementation becomes commoditized, sharing an example of a junior hire vibe-coding internal tooling. | |
| Pain as the Ultimate Competitive Moat | 4 | 7 | 1 | 1 | Kiriti coins the phrase pain is the new moat, explaining how enduring trial-and-error experimentation forms defensible organizational knowledge. | |
| Lightning Round: Books, Media, Productivity Hacks, and Personal Reflections | 4 | 3 | 1 | 1 | The conversation shifts to a friendly lightning round covering book recommendations, sci-fi media, productivity utilities like Raycast and Caffeinate, and personal marital reflections. | |
| Guest Plugs, Course Resources, and Farewell | 1 | 1 | 0 | 0 | Lenny invites the guests to share their LinkedIn profiles, GitHub open source repositories, and Maven course details before closing the episode. |