Nov 6, 2024 · 40m · lennys-podcast

A conversation with OpenAI's CPO Kevin Weil, Anthropic's CPO Mike Krieger, and Sarah Guo

Kevin Weil · 16m spoken Mike Krieger · 15m spoken Sarah Guo · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI Chief Product Officer Kevin Weil and Anthropic Chief Product Officer Mike Krieger join moderator Sarah Guo to discuss the evolving paradigm of AI product leadership, enterprise deployment strategies, and frontier model capabilities. The conversation highlights the transition toward non-deterministic product development, the centrality of rigorous evaluations, and the emergence of proactive, multimodal reasoning systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

Lenny as informed peer 5.8 Guest teaching 2.8 Guest disagreement 0.3 Lenny pushing back 0.5
05100:0015:0030:000:18–3:01 · Lenny as informed peer 5/10 Welcome and Stepping into AI Chief Product Officer Roles Sarah Guo opens the panel with playful, informed framing about both CPOs' backgrounds leading Instagram before asking about their transitions into AI leadership. Both Kevin Weil and Mike Krieger share warm, collaborative reflections on the unique pace of AI development.3:03–8:01 · Lenny as informed peer 6/10 Navigating Enterprise Realities and Fast-Moving Tech Shifts Guo prompts the guests with specific observations about enterprise feedback loops and research-driven cultures. Krieger and Weil elaborate on the contrast between consumer product iteration and long enterprise procurement cycles alongside unpredictable emergent model capabilities.8:03–10:13 · Lenny as informed peer 5/10 The Product and Research Collaboration Cycle Guo asks how product planning functions when underlying capabilities emerge unpredictably from research. Krieger and Weil detail the collaborative co-design and fine-tuning cycles between product managers and research scientists.10:14–15:48 · Lenny as informed peer 7/10 Designing for 60% Accuracy and the Centrality of Evaluations Guo poses a sharp product-investing question regarding how to design around 60% model accuracy versus 99%, accurately noting GitHub Copilot was built on early GPT models. Weil and Krieger explain how writing evaluations has become the critical gating competency for modern AI product managers.15:48–22:46 · Lenny as informed peer 5/10 Developing Intuition and Essential Skills for AI Product Teams Guo asks how builders without access to proprietary internal bootcamps can develop eval intuition. Krieger and Weil emphasize inspecting raw failure data directly, prototyping with models, and getting comfortable managing non-deterministic systems.22:47–27:12 · Lenny as informed peer 6/10 Educating End Users and Enterprise Change Management Guo raises the challenge of enterprise change management for unintuitive AI tools across non-technical workforces. Weil uses a Waymo analogy to illustrate how rapidly user expectations adjust, while Krieger details training Claude to guide users through its own documentation.27:12–33:12 · Lenny as informed peer 6/10 Breakthrough Capabilities: Anthropic Computer Use and OpenAI o1 Guo asks for technical clarity on Anthropic's Computer Use and OpenAI's o1 model, explicitly prompting Weil to define reasoning. Weil delivers a detailed breakdown contrasting pre-training scaling with inference-time reasoning and multi-model orchestration.33:14–36:55 · Lenny as informed peer 6/10 Future Horizons: Proactivity, Multimodal Voice, and Ambient AI Guo prompts both CPOs to project forward 6 to 12 months on upcoming UX paradigms and surprising user behaviors. Krieger emphasizes proactivity and asynchronous reasoning, while Weil highlights real-time speech translation and anthropomorphic model interaction.0:18–3:01 · Guest teaching 1/10 Welcome and Stepping into AI Chief Product Officer Roles Sarah Guo opens the panel with playful, informed framing about both CPOs' backgrounds leading Instagram before asking about their transitions into AI leadership. Both Kevin Weil and Mike Krieger share warm, collaborative reflections on the unique pace of AI development.3:03–8:01 · Guest teaching 3/10 Navigating Enterprise Realities and Fast-Moving Tech Shifts Guo prompts the guests with specific observations about enterprise feedback loops and research-driven cultures. Krieger and Weil elaborate on the contrast between consumer product iteration and long enterprise procurement cycles alongside unpredictable emergent model capabilities.8:03–10:13 · Guest teaching 2/10 The Product and Research Collaboration Cycle Guo asks how product planning functions when underlying capabilities emerge unpredictably from research. Krieger and Weil detail the collaborative co-design and fine-tuning cycles between product managers and research scientists.10:14–15:48 · Guest teaching 4/10 Designing for 60% Accuracy and the Centrality of Evaluations Guo poses a sharp product-investing question regarding how to design around 60% model accuracy versus 99%, accurately noting GitHub Copilot was built on early GPT models. Weil and Krieger explain how writing evaluations has become the critical gating competency for modern AI product managers.15:48–22:46 · Guest teaching 3/10 Developing Intuition and Essential Skills for AI Product Teams Guo asks how builders without access to proprietary internal bootcamps can develop eval intuition. Krieger and Weil emphasize inspecting raw failure data directly, prototyping with models, and getting comfortable managing non-deterministic systems.22:47–27:12 · Guest teaching 2/10 Educating End Users and Enterprise Change Management Guo raises the challenge of enterprise change management for unintuitive AI tools across non-technical workforces. Weil uses a Waymo analogy to illustrate how rapidly user expectations adjust, while Krieger details training Claude to guide users through its own documentation.27:12–33:12 · Guest teaching 5/10 Breakthrough Capabilities: Anthropic Computer Use and OpenAI o1 Guo asks for technical clarity on Anthropic's Computer Use and OpenAI's o1 model, explicitly prompting Weil to define reasoning. Weil delivers a detailed breakdown contrasting pre-training scaling with inference-time reasoning and multi-model orchestration.33:14–36:55 · Guest teaching 2/10 Future Horizons: Proactivity, Multimodal Voice, and Ambient AI Guo prompts both CPOs to project forward 6 to 12 months on upcoming UX paradigms and surprising user behaviors. Krieger emphasizes proactivity and asynchronous reasoning, while Weil highlights real-time speech translation and anthropomorphic model interaction.0:18–3:01 · Guest disagreement 0/10 Welcome and Stepping into AI Chief Product Officer Roles Sarah Guo opens the panel with playful, informed framing about both CPOs' backgrounds leading Instagram before asking about their transitions into AI leadership. Both Kevin Weil and Mike Krieger share warm, collaborative reflections on the unique pace of AI development.3:03–8:01 · Guest disagreement 1/10 Navigating Enterprise Realities and Fast-Moving Tech Shifts Guo prompts the guests with specific observations about enterprise feedback loops and research-driven cultures. Krieger and Weil elaborate on the contrast between consumer product iteration and long enterprise procurement cycles alongside unpredictable emergent model capabilities.8:03–10:13 · Guest disagreement 0/10 The Product and Research Collaboration Cycle Guo asks how product planning functions when underlying capabilities emerge unpredictably from research. Krieger and Weil detail the collaborative co-design and fine-tuning cycles between product managers and research scientists.10:14–15:48 · Guest disagreement 1/10 Designing for 60% Accuracy and the Centrality of Evaluations Guo poses a sharp product-investing question regarding how to design around 60% model accuracy versus 99%, accurately noting GitHub Copilot was built on early GPT models. Weil and Krieger explain how writing evaluations has become the critical gating competency for modern AI product managers.15:48–22:46 · Guest disagreement 0/10 Developing Intuition and Essential Skills for AI Product Teams Guo asks how builders without access to proprietary internal bootcamps can develop eval intuition. Krieger and Weil emphasize inspecting raw failure data directly, prototyping with models, and getting comfortable managing non-deterministic systems.22:47–27:12 · Guest disagreement 0/10 Educating End Users and Enterprise Change Management Guo raises the challenge of enterprise change management for unintuitive AI tools across non-technical workforces. Weil uses a Waymo analogy to illustrate how rapidly user expectations adjust, while Krieger details training Claude to guide users through its own documentation.27:12–33:12 · Guest disagreement 0/10 Breakthrough Capabilities: Anthropic Computer Use and OpenAI o1 Guo asks for technical clarity on Anthropic's Computer Use and OpenAI's o1 model, explicitly prompting Weil to define reasoning. Weil delivers a detailed breakdown contrasting pre-training scaling with inference-time reasoning and multi-model orchestration.33:14–36:55 · Guest disagreement 0/10 Future Horizons: Proactivity, Multimodal Voice, and Ambient AI Guo prompts both CPOs to project forward 6 to 12 months on upcoming UX paradigms and surprising user behaviors. Krieger emphasizes proactivity and asynchronous reasoning, while Weil highlights real-time speech translation and anthropomorphic model interaction.0:18–3:01 · Lenny pushing back 0/10 Welcome and Stepping into AI Chief Product Officer Roles Sarah Guo opens the panel with playful, informed framing about both CPOs' backgrounds leading Instagram before asking about their transitions into AI leadership. Both Kevin Weil and Mike Krieger share warm, collaborative reflections on the unique pace of AI development.3:03–8:01 · Lenny pushing back 1/10 Navigating Enterprise Realities and Fast-Moving Tech Shifts Guo prompts the guests with specific observations about enterprise feedback loops and research-driven cultures. Krieger and Weil elaborate on the contrast between consumer product iteration and long enterprise procurement cycles alongside unpredictable emergent model capabilities.8:03–10:13 · Lenny pushing back 0/10 The Product and Research Collaboration Cycle Guo asks how product planning functions when underlying capabilities emerge unpredictably from research. Krieger and Weil detail the collaborative co-design and fine-tuning cycles between product managers and research scientists.10:14–15:48 · Lenny pushing back 1/10 Designing for 60% Accuracy and the Centrality of Evaluations Guo poses a sharp product-investing question regarding how to design around 60% model accuracy versus 99%, accurately noting GitHub Copilot was built on early GPT models. Weil and Krieger explain how writing evaluations has become the critical gating competency for modern AI product managers.15:48–22:46 · Lenny pushing back 0/10 Developing Intuition and Essential Skills for AI Product Teams Guo asks how builders without access to proprietary internal bootcamps can develop eval intuition. Krieger and Weil emphasize inspecting raw failure data directly, prototyping with models, and getting comfortable managing non-deterministic systems.22:47–27:12 · Lenny pushing back 1/10 Educating End Users and Enterprise Change Management Guo raises the challenge of enterprise change management for unintuitive AI tools across non-technical workforces. Weil uses a Waymo analogy to illustrate how rapidly user expectations adjust, while Krieger details training Claude to guide users through its own documentation.27:12–33:12 · Lenny pushing back 1/10 Breakthrough Capabilities: Anthropic Computer Use and OpenAI o1 Guo asks for technical clarity on Anthropic's Computer Use and OpenAI's o1 model, explicitly prompting Weil to define reasoning. Weil delivers a detailed breakdown contrasting pre-training scaling with inference-time reasoning and multi-model orchestration.33:14–36:55 · Lenny pushing back 0/10 Future Horizons: Proactivity, Multimodal Voice, and Ambient AI Guo prompts both CPOs to project forward 6 to 12 months on upcoming UX paradigms and surprising user behaviors. Krieger emphasizes proactivity and asynchronous reasoning, while Weil highlights real-time speech translation and anthropomorphic model interaction.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 0% · guest 100%0:00 · Lenny 0% · guest 100%3:00 · Lenny 0% · guest 100%3:00 · Lenny 0% · guest 100%6:00 · Lenny 0% · guest 100%6:00 · Lenny 0% · guest 100%9:00 · Lenny 0% · guest 100%9:00 · Lenny 0% · guest 100%12:00 · Lenny 0% · guest 100%12:00 · Lenny 0% · guest 100%15:00 · Lenny 0% · guest 100%15:00 · Lenny 0% · guest 100%18:00 · Lenny 0% · guest 100%18:00 · Lenny 0% · guest 100%21:00 · Lenny 0% · guest 100%21:00 · Lenny 0% · guest 100%24:00 · Lenny 0% · guest 100%24:00 · Lenny 0% · guest 100%27:00 · Lenny 0% · guest 100%27:00 · Lenny 0% · guest 100%30:00 · Lenny 0% · guest 100%30:00 · Lenny 0% · guest 100%33:00 · Lenny 0% · guest 100%33:00 · Lenny 0% · guest 100%36:00 · Lenny 0% · guest 100%36:00 · Lenny 0% · guest 100%39:00 · Lenny 0% · guest 100%39:00 · Lenny 0% · guest 100%
Sharpest disagreement ▶ 10:49 Weil challenges the premise of 60% accuracy being a blocker

Weil playfully pushes back on Guo's question about accuracy thresholds, arguing that 60% reliability is already immensely valuable when paired with human-in-the-loop product design.

Hardest push from Lenny ▶ 25:44 Guo presses on enterprise change management reality

Guo redirects the conversation from consumer adoption speed to the rigid status quo and organizational friction inherent in enterprise workflow changes.

Biggest teaching moment ▶ 29:56 Weil explains the mechanics of inference-time reasoning

Weil educates the audience and host on the foundational distinction between pre-training System 1 token prediction and o1's query-time hypothesis formation and validation.

Lenny holds their own ▶ 11:24 Guo cites GPT-2 base architecture for early Copilot

Guo demonstrates deep technical domain knowledge by interjecting that the original GitHub Copilot ran on a relatively small GPT-2 model to validate Weil's point.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Welcome and Stepping into AI Chief Product Officer Roles 5100 Sarah Guo opens the panel with playful, informed framing about both CPOs' backgrounds leading Instagram before asking about their transitions into AI leadership. Both Kevin Weil and Mike Krieger share warm, collaborative reflections on the unique pace of AI development.
Navigating Enterprise Realities and Fast-Moving Tech Shifts 6311 Guo prompts the guests with specific observations about enterprise feedback loops and research-driven cultures. Krieger and Weil elaborate on the contrast between consumer product iteration and long enterprise procurement cycles alongside unpredictable emergent model capabilities.
The Product and Research Collaboration Cycle 5200 Guo asks how product planning functions when underlying capabilities emerge unpredictably from research. Krieger and Weil detail the collaborative co-design and fine-tuning cycles between product managers and research scientists.
Designing for 60% Accuracy and the Centrality of Evaluations 7411 Guo poses a sharp product-investing question regarding how to design around 60% model accuracy versus 99%, accurately noting GitHub Copilot was built on early GPT models. Weil and Krieger explain how writing evaluations has become the critical gating competency for modern AI product managers.
Developing Intuition and Essential Skills for AI Product Teams 5300 Guo asks how builders without access to proprietary internal bootcamps can develop eval intuition. Krieger and Weil emphasize inspecting raw failure data directly, prototyping with models, and getting comfortable managing non-deterministic systems.
Educating End Users and Enterprise Change Management 6201 Guo raises the challenge of enterprise change management for unintuitive AI tools across non-technical workforces. Weil uses a Waymo analogy to illustrate how rapidly user expectations adjust, while Krieger details training Claude to guide users through its own documentation.
Breakthrough Capabilities: Anthropic Computer Use and OpenAI o1 6501 Guo asks for technical clarity on Anthropic's Computer Use and OpenAI's o1 model, explicitly prompting Weil to define reasoning. Weil delivers a detailed breakdown contrasting pre-training scaling with inference-time reasoning and multi-model orchestration.
Future Horizons: Proactivity, Multimodal Voice, and Ambient AI 6200 Guo prompts both CPOs to project forward 6 to 12 months on upcoming UX paradigms and surprising user behaviors. Krieger emphasizes proactivity and asynchronous reasoning, while Weil highlights real-time speech translation and anthropomorphic model interaction.

Statements from this episode (17)

Insight
Weil: AI product development fundamentally differs because capabilities shift every two months
“Normally when you're building product, you're building off of kind of a fixed technology base, right? You know what you have to work with, and you're trying to build the best product you can. Here it's like every two months computers can do something computers…”
Kevin Weil Nov 6, 2024 ▶ 1:46
Opinion
Krieger: Only About Three Companies Could Have Recruited Me Post-Instagram
“I mean, not many people could, but like, this is like probably a list of three companies that would have been interesting.”
Mike Krieger Nov 6, 2024 ▶ 2:46
Insight
Weil: AI product design fundamentally differs based on model reliability tiers
“You don't know whether it's going to be like, 60% good, or 90% good, or 99% good, and the product that you would build that would make The sense with something that works 60% of the time is super different than 90 or 99% of the time, right?”
Kevin Weil Nov 6, 2024 ▶ 6:55
Insight
Krieger compares internal AI research disruptions to Apple WWDC announcements
“It's the thing it most reminds me of like from the Instagram days where like Apple, like WWC announcements, you're like, this could either be awesome for us or could like absolutely like cause chaos for it. It's like that, but your own company is the one kind …”
Mike Krieger Nov 6, 2024 ▶ 7:42
Insight
Weil: Current AI models are eval-limited rather than intelligence-limited
“I think there's a very real sense in which models today are not intelligence limited. They're eval limited. They can actually do much more and be much more correct on a wider range of things than they are today, and it's really about sort of teaching them.”
Kevin Weil Nov 6, 2024 ▶ 13:12
Prediction Not checkable as stated
Weil: Writing evals will become a core skill for PMs
“Yeah, writing evals, I mean, it's, I actually think it's going to become a core skill for PMs.”
Kevin Weil Nov 6, 2024 ▶ 14:37
Disclosure
Weil: OpenAI runs internal boot camps teaching product managers to write evals
“We set up a boot camp and, like, took every PM through Writing evals, and like, what it was like, difference between good and bad evals, and you know, we're definitely not done there.”
Kevin Weil Nov 6, 2024 ▶ 15:30
Insight
Krieger: AI benchmark evals often contain flawed ground-truth answers
“And some of these model these evals we've seen, like, even the golden answer, I'm like, I'm not sure a human would say it, or, like, I think that math is actually a little wrong, like, getting a hundred percent is gonna be really hard, because even just gradin…”
Mike Krieger Nov 6, 2024 ▶ 16:50
Assertion Supported
Weil: Frontier AI increasingly generates answers that evaluators prefer over humans
“The models are getting to the point where they can often beat humans at certain tasks. Like, people prefer the model's answers to a human's answers, and so if you're humans writing your evals, like, you know, so what does that mean?”
Kevin Weil Nov 6, 2024 ▶ 18:35
Insight
Krieger: Best product managers rapidly prototype UI designs directly using AI models
“I think prototyping with these models is a thing that is underused. Like, our best PMs do this, where we'll get in some long conversation about, like, should the UI be this or that, and before our designers have even, like, picked up their Figma, like, our, of…”
Mike Krieger Nov 6, 2024 ▶ 19:04
Prediction Not checkable as stated
Weil predicts today's cutting-edge AI will feel like garbage within 12 months
“And, you know, the stuff that's happening today that we're working on, that you guys are working on, it all feels like magic. 12 months from now, we're going to be like, can you believe we use that garbage? Because it's going to, I mean, that's how fast this t…”
Kevin Weil Nov 6, 2024 ▶ 24:17
Disclosure
Krieger: Anthropic is training Claude to teach users its own features
“One thing we're trying to get better at, and that's also letting the product be like educational in a very literal way, which is like the thing we did not do early and now we're changing is just tell Claude more about itself, which was like, you know, it's in …”
Mike Krieger Nov 6, 2024 ▶ 24:53
Assertion Not checkable as stated
Weil: Enterprise organizations often create thousands of internal custom GPTs
“With OpenAI, we have these custom GPTs that you can make, and organizations make thousands of them often, and it's a way for the power users to make something that makes AI easier and, like, immediately valuable for the people that might not know how to use it…”
Kevin Weil Nov 6, 2024 ▶ 26:43
Opinion
Krieger: Anthropic's Computer Use feature shows early promise for brittle UI testing
“Some early things that we're seeing that we think are really interesting, one is UI testing, which is like, I was, like, at Instagram, we had basically no UI tests, because they're hard to write, they're, like, they're brittle and they're, like, often, like, a…”
Mike Krieger Nov 6, 2024 ▶ 28:12
Insight
Weil: Top AI customers orchestrate workflows across multiple models, not just one
“People maybe don't realize that actually a lot of the most sophisticated customers of ours are doing, and that we're certainly doing internally, is it's not really about one model for any particular thing. You end up putting together sort of workflows and orch…”
Kevin Weil Nov 6, 2024 ▶ 29:30
Opinion
Weil: AI reasoning models are currently only in their GPT-1 phase
“So, it's basically a new way to scale intelligence, and we feel like we're just at the very beginning, you know, we're at the, like, GPT-I phase of this new form of reasoning.”
Kevin Weil Nov 6, 2024 ▶ 31:49
Insight
Weil: AI model behavior and personality are core product responsibilities
“Yeah, model behavior is absolutely a product role. Like, the personality of the model is, is key, and there are interesting questions around how much should it customize versus how much should, you know, OpenAI have one personality, and Claude has some distinc…”
Kevin Weil Nov 6, 2024 ▶ 39:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.