Feb 9, 2025 · 1h 14m · lennys-podcast

OpenAI researcher on why soft skills are the future of work | Karina Nguyen

Karina Nguyen · 48m spoken Lenny Rachitsky · 18m spoken Christina Cacioppo · 47s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI researcher Karina Nguyen explains how post-training, synthetic data, and evaluation benchmarks are reshaping AI product development while highlighting why human creativity, aesthetic taste, and soft skills represent the future of knowledge work.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 27.2% of the talking time here. How this is scored →

Lenny as informed peer 3.7 Guest teaching 5.1 Guest disagreement 1.1 Lenny pushing back 1.3
05100:0015:0030:0045:001:00:004:45–8:21 · Lenny as informed peer 3/10 Model Training Misconceptions and Philosophical Edge Cases Lenny asks about common misconceptions in model training. Karina educates him on the subtle art of post-training debugging, sharing how Claude 3 became confused when taught it lacked a physical body while also being taught tool-use functions like setting an alarm.8:22–12:49 · Lenny as informed peer 4/10 Debunking the Data Wall with Synthetic Data and RL Lenny brings up the popular theory that models will hit a data wall after exhausting internet data. Karina reframes the problem, explaining that post-training via reinforcement learning creates an infinite task space, shifting the true bottleneck to evaluations.12:50–18:33 · Lenny as informed peer 3/10 Designing OpenAI Canvas with Synthetic Model Training Karina walks through how OpenAI Canvas was built from scratch using synthetic data generation to train core decision boundaries like triggering, targeted in-place edits, and document critique. Lenny listens and mirrors back the synthetic training paradigm.18:33–23:22 · Lenny as informed peer 3/10 Day-to-Day Research Workflows and Deterministic Evals Karina describes her day-to-day workflow transitioning from IC research to management and explains how product managers use deterministic evaluation spreadsheets and win-rate benchmarks.23:23–26:56 · Lenny as informed peer 4/10 The Evolution of Product Development Through Evals and Prompting Lenny synthesizes how product development is shifting from PRDs to evals and rapid prototyping. Karina gives examples from Anthropic on how prompt-based prototyping unlocked features like 100k context document uploads and title generation.26:57–32:33 · Lenny as informed peer 3/10 Building Tasks at OpenAI and Iterative Deployment Karina breaks down the design and staffing lifecycle of OpenAI Tasks, detailing how tool specs with JSON schemas are crafted and iterated through synthetic data rather than slow human data collection.32:35–35:34 · Lenny as informed peer 2/10 Sponsor Segment: Loom Following a sponsor read, Lenny asks for clarification on what researchers do versus model designers. Karina explains the difference between product research and long-term capability research focused on synthetic data diversity.35:35–42:16 · Lenny as informed peer 5/10 Societal and Economic Shifts from Falling Intelligence Costs Karina discusses how falling intelligence costs and model distillation enable smarter small models. Lenny demonstrates subject familiarity by citing a New York Times study evaluating diagnostic accuracy between doctors and ChatGPT.42:16–47:50 · Lenny as informed peer 5/10 Soft Skills, Creative Thinking, and Research Prioritization Karina argues soft skills like creative thinking and research prioritization under compute constraints will remain defensible human skills. Lenny validates this by referencing his own newsletter analysis on how AI alters product management.47:50–53:34 · Lenny as informed peer 5/10 Strategic Reasoning, AI Idea Generation, and Research Management Lenny pushes a contrarian thesis that high-level business strategy can be completely taken over by AI through multi-source data synthesis. Karina agrees, pointing toward self-improving product feedback loops.53:35–57:10 · Lenny as informed peer 4/10 Comparing Company Cultures: OpenAI vs Anthropic Karina compares Anthropic's tight prioritization and craft-oriented model personality against OpenAI's bottom-up risk-taking and research freedom. Lenny summarizes the organizational contrast.57:11–1:02:00 · Lenny as informed peer 3/10 Prototyping Early Product Form Factors and Collaborative Agents Karina reflects on early prototypes like Claude in Slack and discusses moving from synchronous chat to asynchronous collaborative agents. Lenny presses on why workplace Slack bots were deprioritized.1:02:00–1:05:23 · Lenny as informed peer 5/10 Breakthrough Milestones: Context Windows to Simulated Personas Karina highlights major step-function moments in LLMs, such as large context windows and simulated personas. Lenny shares his firsthand workflow of using large context windows to process guest books and building his own voice-enabled Lennybot persona.1:05:24–1:11:35 · Lenny as informed peer 5/10 Chat Interfaces, Desktop Control, and OpenAI Operator Lenny references OpenAI CPO Kevin Weil's observations on why chat interfaces endure and mentions Tuple for pair programming. Karina unpacks the technical hurdles behind computer-use agents like OpenAI Operator, specifically pixel perception and user intent modeling.1:11:36–1:14:05 · Lenny as informed peer 2/10 Post-AGI Aspirations: Fiction Writing and Art Conservation In a closing lighthearted exchange, Karina shares her post-AGI dreams of writing sci-fi novels and restoring museum art, followed by hiring calls for her Frontier Product Research team.4:45–8:21 · Guest teaching 6/10 Model Training Misconceptions and Philosophical Edge Cases Lenny asks about common misconceptions in model training. Karina educates him on the subtle art of post-training debugging, sharing how Claude 3 became confused when taught it lacked a physical body while also being taught tool-use functions like setting an alarm.8:22–12:49 · Guest teaching 7/10 Debunking the Data Wall with Synthetic Data and RL Lenny brings up the popular theory that models will hit a data wall after exhausting internet data. Karina reframes the problem, explaining that post-training via reinforcement learning creates an infinite task space, shifting the true bottleneck to evaluations.12:50–18:33 · Guest teaching 6/10 Designing OpenAI Canvas with Synthetic Model Training Karina walks through how OpenAI Canvas was built from scratch using synthetic data generation to train core decision boundaries like triggering, targeted in-place edits, and document critique. Lenny listens and mirrors back the synthetic training paradigm.18:33–23:22 · Guest teaching 6/10 Day-to-Day Research Workflows and Deterministic Evals Karina describes her day-to-day workflow transitioning from IC research to management and explains how product managers use deterministic evaluation spreadsheets and win-rate benchmarks.23:23–26:56 · Guest teaching 5/10 The Evolution of Product Development Through Evals and Prompting Lenny synthesizes how product development is shifting from PRDs to evals and rapid prototyping. Karina gives examples from Anthropic on how prompt-based prototyping unlocked features like 100k context document uploads and title generation.26:57–32:33 · Guest teaching 6/10 Building Tasks at OpenAI and Iterative Deployment Karina breaks down the design and staffing lifecycle of OpenAI Tasks, detailing how tool specs with JSON schemas are crafted and iterated through synthetic data rather than slow human data collection.32:35–35:34 · Guest teaching 5/10 Sponsor Segment: Loom Following a sponsor read, Lenny asks for clarification on what researchers do versus model designers. Karina explains the difference between product research and long-term capability research focused on synthetic data diversity.35:35–42:16 · Guest teaching 5/10 Societal and Economic Shifts from Falling Intelligence Costs Karina discusses how falling intelligence costs and model distillation enable smarter small models. Lenny demonstrates subject familiarity by citing a New York Times study evaluating diagnostic accuracy between doctors and ChatGPT.42:16–47:50 · Guest teaching 4/10 Soft Skills, Creative Thinking, and Research Prioritization Karina argues soft skills like creative thinking and research prioritization under compute constraints will remain defensible human skills. Lenny validates this by referencing his own newsletter analysis on how AI alters product management.47:50–53:34 · Guest teaching 4/10 Strategic Reasoning, AI Idea Generation, and Research Management Lenny pushes a contrarian thesis that high-level business strategy can be completely taken over by AI through multi-source data synthesis. Karina agrees, pointing toward self-improving product feedback loops.53:35–57:10 · Guest teaching 5/10 Comparing Company Cultures: OpenAI vs Anthropic Karina compares Anthropic's tight prioritization and craft-oriented model personality against OpenAI's bottom-up risk-taking and research freedom. Lenny summarizes the organizational contrast.57:11–1:02:00 · Guest teaching 5/10 Prototyping Early Product Form Factors and Collaborative Agents Karina reflects on early prototypes like Claude in Slack and discusses moving from synchronous chat to asynchronous collaborative agents. Lenny presses on why workplace Slack bots were deprioritized.1:02:00–1:05:23 · Guest teaching 4/10 Breakthrough Milestones: Context Windows to Simulated Personas Karina highlights major step-function moments in LLMs, such as large context windows and simulated personas. Lenny shares his firsthand workflow of using large context windows to process guest books and building his own voice-enabled Lennybot persona.1:05:24–1:11:35 · Guest teaching 6/10 Chat Interfaces, Desktop Control, and OpenAI Operator Lenny references OpenAI CPO Kevin Weil's observations on why chat interfaces endure and mentions Tuple for pair programming. Karina unpacks the technical hurdles behind computer-use agents like OpenAI Operator, specifically pixel perception and user intent modeling.1:11:36–1:14:05 · Guest teaching 2/10 Post-AGI Aspirations: Fiction Writing and Art Conservation In a closing lighthearted exchange, Karina shares her post-AGI dreams of writing sci-fi novels and restoring museum art, followed by hiring calls for her Frontier Product Research team.4:45–8:21 · Guest disagreement 1/10 Model Training Misconceptions and Philosophical Edge Cases Lenny asks about common misconceptions in model training. Karina educates him on the subtle art of post-training debugging, sharing how Claude 3 became confused when taught it lacked a physical body while also being taught tool-use functions like setting an alarm.8:22–12:49 · Guest disagreement 2/10 Debunking the Data Wall with Synthetic Data and RL Lenny brings up the popular theory that models will hit a data wall after exhausting internet data. Karina reframes the problem, explaining that post-training via reinforcement learning creates an infinite task space, shifting the true bottleneck to evaluations.12:50–18:33 · Guest disagreement 1/10 Designing OpenAI Canvas with Synthetic Model Training Karina walks through how OpenAI Canvas was built from scratch using synthetic data generation to train core decision boundaries like triggering, targeted in-place edits, and document critique. Lenny listens and mirrors back the synthetic training paradigm.18:33–23:22 · Guest disagreement 1/10 Day-to-Day Research Workflows and Deterministic Evals Karina describes her day-to-day workflow transitioning from IC research to management and explains how product managers use deterministic evaluation spreadsheets and win-rate benchmarks.23:23–26:56 · Guest disagreement 1/10 The Evolution of Product Development Through Evals and Prompting Lenny synthesizes how product development is shifting from PRDs to evals and rapid prototyping. Karina gives examples from Anthropic on how prompt-based prototyping unlocked features like 100k context document uploads and title generation.26:57–32:33 · Guest disagreement 1/10 Building Tasks at OpenAI and Iterative Deployment Karina breaks down the design and staffing lifecycle of OpenAI Tasks, detailing how tool specs with JSON schemas are crafted and iterated through synthetic data rather than slow human data collection.32:35–35:34 · Guest disagreement 1/10 Sponsor Segment: Loom Following a sponsor read, Lenny asks for clarification on what researchers do versus model designers. Karina explains the difference between product research and long-term capability research focused on synthetic data diversity.35:35–42:16 · Guest disagreement 1/10 Societal and Economic Shifts from Falling Intelligence Costs Karina discusses how falling intelligence costs and model distillation enable smarter small models. Lenny demonstrates subject familiarity by citing a New York Times study evaluating diagnostic accuracy between doctors and ChatGPT.42:16–47:50 · Guest disagreement 1/10 Soft Skills, Creative Thinking, and Research Prioritization Karina argues soft skills like creative thinking and research prioritization under compute constraints will remain defensible human skills. Lenny validates this by referencing his own newsletter analysis on how AI alters product management.47:50–53:34 · Guest disagreement 1/10 Strategic Reasoning, AI Idea Generation, and Research Management Lenny pushes a contrarian thesis that high-level business strategy can be completely taken over by AI through multi-source data synthesis. Karina agrees, pointing toward self-improving product feedback loops.53:35–57:10 · Guest disagreement 1/10 Comparing Company Cultures: OpenAI vs Anthropic Karina compares Anthropic's tight prioritization and craft-oriented model personality against OpenAI's bottom-up risk-taking and research freedom. Lenny summarizes the organizational contrast.57:11–1:02:00 · Guest disagreement 1/10 Prototyping Early Product Form Factors and Collaborative Agents Karina reflects on early prototypes like Claude in Slack and discusses moving from synchronous chat to asynchronous collaborative agents. Lenny presses on why workplace Slack bots were deprioritized.1:02:00–1:05:23 · Guest disagreement 1/10 Breakthrough Milestones: Context Windows to Simulated Personas Karina highlights major step-function moments in LLMs, such as large context windows and simulated personas. Lenny shares his firsthand workflow of using large context windows to process guest books and building his own voice-enabled Lennybot persona.1:05:24–1:11:35 · Guest disagreement 1/10 Chat Interfaces, Desktop Control, and OpenAI Operator Lenny references OpenAI CPO Kevin Weil's observations on why chat interfaces endure and mentions Tuple for pair programming. Karina unpacks the technical hurdles behind computer-use agents like OpenAI Operator, specifically pixel perception and user intent modeling.1:11:36–1:14:05 · Guest disagreement 1/10 Post-AGI Aspirations: Fiction Writing and Art Conservation In a closing lighthearted exchange, Karina shares her post-AGI dreams of writing sci-fi novels and restoring museum art, followed by hiring calls for her Frontier Product Research team.4:45–8:21 · Lenny pushing back 1/10 Model Training Misconceptions and Philosophical Edge Cases Lenny asks about common misconceptions in model training. Karina educates him on the subtle art of post-training debugging, sharing how Claude 3 became confused when taught it lacked a physical body while also being taught tool-use functions like setting an alarm.8:22–12:49 · Lenny pushing back 2/10 Debunking the Data Wall with Synthetic Data and RL Lenny brings up the popular theory that models will hit a data wall after exhausting internet data. Karina reframes the problem, explaining that post-training via reinforcement learning creates an infinite task space, shifting the true bottleneck to evaluations.12:50–18:33 · Lenny pushing back 1/10 Designing OpenAI Canvas with Synthetic Model Training Karina walks through how OpenAI Canvas was built from scratch using synthetic data generation to train core decision boundaries like triggering, targeted in-place edits, and document critique. Lenny listens and mirrors back the synthetic training paradigm.18:33–23:22 · Lenny pushing back 1/10 Day-to-Day Research Workflows and Deterministic Evals Karina describes her day-to-day workflow transitioning from IC research to management and explains how product managers use deterministic evaluation spreadsheets and win-rate benchmarks.23:23–26:56 · Lenny pushing back 2/10 The Evolution of Product Development Through Evals and Prompting Lenny synthesizes how product development is shifting from PRDs to evals and rapid prototyping. Karina gives examples from Anthropic on how prompt-based prototyping unlocked features like 100k context document uploads and title generation.26:57–32:33 · Lenny pushing back 1/10 Building Tasks at OpenAI and Iterative Deployment Karina breaks down the design and staffing lifecycle of OpenAI Tasks, detailing how tool specs with JSON schemas are crafted and iterated through synthetic data rather than slow human data collection.32:35–35:34 · Lenny pushing back 1/10 Sponsor Segment: Loom Following a sponsor read, Lenny asks for clarification on what researchers do versus model designers. Karina explains the difference between product research and long-term capability research focused on synthetic data diversity.35:35–42:16 · Lenny pushing back 1/10 Societal and Economic Shifts from Falling Intelligence Costs Karina discusses how falling intelligence costs and model distillation enable smarter small models. Lenny demonstrates subject familiarity by citing a New York Times study evaluating diagnostic accuracy between doctors and ChatGPT.42:16–47:50 · Lenny pushing back 1/10 Soft Skills, Creative Thinking, and Research Prioritization Karina argues soft skills like creative thinking and research prioritization under compute constraints will remain defensible human skills. Lenny validates this by referencing his own newsletter analysis on how AI alters product management.47:50–53:34 · Lenny pushing back 2/10 Strategic Reasoning, AI Idea Generation, and Research Management Lenny pushes a contrarian thesis that high-level business strategy can be completely taken over by AI through multi-source data synthesis. Karina agrees, pointing toward self-improving product feedback loops.53:35–57:10 · Lenny pushing back 1/10 Comparing Company Cultures: OpenAI vs Anthropic Karina compares Anthropic's tight prioritization and craft-oriented model personality against OpenAI's bottom-up risk-taking and research freedom. Lenny summarizes the organizational contrast.57:11–1:02:00 · Lenny pushing back 2/10 Prototyping Early Product Form Factors and Collaborative Agents Karina reflects on early prototypes like Claude in Slack and discusses moving from synchronous chat to asynchronous collaborative agents. Lenny presses on why workplace Slack bots were deprioritized.1:02:00–1:05:23 · Lenny pushing back 1/10 Breakthrough Milestones: Context Windows to Simulated Personas Karina highlights major step-function moments in LLMs, such as large context windows and simulated personas. Lenny shares his firsthand workflow of using large context windows to process guest books and building his own voice-enabled Lennybot persona.1:05:24–1:11:35 · Lenny pushing back 1/10 Chat Interfaces, Desktop Control, and OpenAI Operator Lenny references OpenAI CPO Kevin Weil's observations on why chat interfaces endure and mentions Tuple for pair programming. Karina unpacks the technical hurdles behind computer-use agents like OpenAI Operator, specifically pixel perception and user intent modeling.1:11:36–1:14:05 · Lenny pushing back 1/10 Post-AGI Aspirations: Fiction Writing and Art Conservation In a closing lighthearted exchange, Karina shares her post-AGI dreams of writing sci-fi novels and restoring museum art, followed by hiring calls for her Frontier Product Research team.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 78.5% · guest 21.5%0:00 · Lenny 78.5% · guest 21.5%3:00 · Lenny 67% · guest 33%3:00 · Lenny 67% · guest 33%6:00 · Lenny 33.6% · guest 66.4%6:00 · Lenny 33.6% · guest 66.4%9:00 · Lenny 10.4% · guest 89.6%9:00 · Lenny 10.4% · guest 89.6%12:00 · Lenny 3.7% · guest 96.3%12:00 · Lenny 3.7% · guest 96.3%15:00 · Lenny 0% · guest 100%15:00 · Lenny 0% · guest 100%18:00 · Lenny 41.4% · guest 58.6%18:00 · Lenny 41.4% · guest 58.6%21:00 · Lenny 14.7% · guest 85.3%21:00 · Lenny 14.7% · guest 85.3%24:00 · Lenny 23% · guest 77%24:00 · Lenny 23% · guest 77%27:00 · Lenny 6.4% · guest 93.6%27:00 · Lenny 6.4% · guest 93.6%30:00 · Lenny 13.8% · guest 86.2%30:00 · Lenny 13.8% · guest 86.2%33:00 · Lenny 50.5% · guest 49.5%33:00 · Lenny 50.5% · guest 49.5%36:00 · Lenny 4.3% · guest 95.7%36:00 · Lenny 4.3% · guest 95.7%39:00 · Lenny 17.5% · guest 82.5%39:00 · Lenny 17.5% · guest 82.5%42:00 · Lenny 14.5% · guest 85.5%42:00 · Lenny 14.5% · guest 85.5%45:00 · Lenny 35% · guest 65%45:00 · Lenny 35% · guest 65%48:00 · Lenny 39% · guest 61%48:00 · Lenny 39% · guest 61%51:00 · Lenny 34.9% · guest 65.1%51:00 · Lenny 34.9% · guest 65.1%54:00 · Lenny 7% · guest 93%54:00 · Lenny 7% · guest 93%57:00 · Lenny 16.4% · guest 83.6%57:00 · Lenny 16.4% · guest 83.6%1:00:00 · Lenny 30.4% · guest 69.6%1:00:00 · Lenny 30.4% · guest 69.6%1:03:00 · Lenny 45.6% · guest 54.4%1:03:00 · Lenny 45.6% · guest 54.4%1:06:00 · Lenny 30.5% · guest 69.5%1:06:00 · Lenny 30.5% · guest 69.5%1:09:00 · Lenny 24.5% · guest 75.5%1:09:00 · Lenny 24.5% · guest 75.5%1:12:00 · Lenny 44.9% · guest 55.1%1:12:00 · Lenny 44.9% · guest 55.1%
Sharpest disagreement ▶ 8:46 Dismantling the internet data wall narrative

Karina directly pushes back against the premise that AI progress is stopping due to exhausted internet data, explaining that post-training RL on infinite tasks has replaced raw pre-training token collection.

Hardest push from Lenny ▶ 49:40 Challenging the idea that strategy is immune to AI automation

Lenny challenges the popular consensus that business strategy requires uniquely human judgment, asserting that strategy is simply multi-source data synthesis and plan creation that LLMs will excel at.

Biggest teaching moment ▶ 6:45 Revealing paradoxical model training edge cases

Karina educates Lenny on how conflicting post-training data—teaching a model it has no physical body while teaching it tool execution like setting an alarm—causes severe confusion and task refusals.

Lenny holds their own ▶ 40:05 Citing clinical study on LLM diagnostic superiority

Lenny validates and expands upon Karina's point about domain intelligence by citing a specific New York Times comparative study where ChatGPT outperformed human doctors.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Model Training Misconceptions and Philosophical Edge Cases 3611 Lenny asks about common misconceptions in model training. Karina educates him on the subtle art of post-training debugging, sharing how Claude 3 became confused when taught it lacked a physical body while also being taught tool-use functions like setting an alarm.
Debunking the Data Wall with Synthetic Data and RL 4722 Lenny brings up the popular theory that models will hit a data wall after exhausting internet data. Karina reframes the problem, explaining that post-training via reinforcement learning creates an infinite task space, shifting the true bottleneck to evaluations.
Designing OpenAI Canvas with Synthetic Model Training 3611 Karina walks through how OpenAI Canvas was built from scratch using synthetic data generation to train core decision boundaries like triggering, targeted in-place edits, and document critique. Lenny listens and mirrors back the synthetic training paradigm.
Day-to-Day Research Workflows and Deterministic Evals 3611 Karina describes her day-to-day workflow transitioning from IC research to management and explains how product managers use deterministic evaluation spreadsheets and win-rate benchmarks.
The Evolution of Product Development Through Evals and Prompting 4512 Lenny synthesizes how product development is shifting from PRDs to evals and rapid prototyping. Karina gives examples from Anthropic on how prompt-based prototyping unlocked features like 100k context document uploads and title generation.
Building Tasks at OpenAI and Iterative Deployment 3611 Karina breaks down the design and staffing lifecycle of OpenAI Tasks, detailing how tool specs with JSON schemas are crafted and iterated through synthetic data rather than slow human data collection.
Sponsor Segment: Loom 2511 Following a sponsor read, Lenny asks for clarification on what researchers do versus model designers. Karina explains the difference between product research and long-term capability research focused on synthetic data diversity.
Societal and Economic Shifts from Falling Intelligence Costs 5511 Karina discusses how falling intelligence costs and model distillation enable smarter small models. Lenny demonstrates subject familiarity by citing a New York Times study evaluating diagnostic accuracy between doctors and ChatGPT.
Soft Skills, Creative Thinking, and Research Prioritization 5411 Karina argues soft skills like creative thinking and research prioritization under compute constraints will remain defensible human skills. Lenny validates this by referencing his own newsletter analysis on how AI alters product management.
Strategic Reasoning, AI Idea Generation, and Research Management 5412 Lenny pushes a contrarian thesis that high-level business strategy can be completely taken over by AI through multi-source data synthesis. Karina agrees, pointing toward self-improving product feedback loops.
Comparing Company Cultures: OpenAI vs Anthropic 4511 Karina compares Anthropic's tight prioritization and craft-oriented model personality against OpenAI's bottom-up risk-taking and research freedom. Lenny summarizes the organizational contrast.
Prototyping Early Product Form Factors and Collaborative Agents 3512 Karina reflects on early prototypes like Claude in Slack and discusses moving from synchronous chat to asynchronous collaborative agents. Lenny presses on why workplace Slack bots were deprioritized.
Breakthrough Milestones: Context Windows to Simulated Personas 5411 Karina highlights major step-function moments in LLMs, such as large context windows and simulated personas. Lenny shares his firsthand workflow of using large context windows to process guest books and building his own voice-enabled Lennybot persona.
Chat Interfaces, Desktop Control, and OpenAI Operator 5611 Lenny references OpenAI CPO Kevin Weil's observations on why chat interfaces endure and mentions Tuple for pair programming. Karina unpacks the technical hurdles behind computer-use agents like OpenAI Operator, specifically pixel perception and user intent modeling.
Post-AGI Aspirations: Fiction Writing and Art Conservation 2211 In a closing lighthearted exchange, Karina shares her post-AGI dreams of writing sci-fi novels and restoring museum art, followed by hiring calls for her Frontier Product Research team.

Statements from this episode (23)

Assertion Supported
Rachitsky: Reader Survey Ranks ChatGPT Ahead of Gmail and Slack
“I did a survey of my readers and asked them what tools to use every day in your work and most use. And ChatGPT was Number one above Gmail, above Slack, above anything else. 90% of people said they use ChatGPT regularly.”
Lenny Rachitsky Feb 9, 2025 ▶ 5:09
Insight
Nguyen: Model Training Is More Art Than Science, Debugged Like Software
“Model training is more an art than a science, and in a lot of ways, like, we as, like, model trainers think a lot about, like, data quality. So, like, it's one of the most important things in model training is, like how do you ensure the highest quality data f…”
Karina Nguyen Feb 9, 2025 ▶ 6:36
Assertion Not checkable as stated
Nguyen: Claude Refused Setting Alarms After Realizing It Lacked a Body
“One of the things that I've learned early days at Anthropic was, like, we've discovered, especially with, like, cloud three training, when you taught the model some of the self-knowledge of, like, hey, like, you actually don't have a physical body to operate, …”
Karina Nguyen Feb 9, 2025 ▶ 7:05
Insight
Nguyen: Post-training scaling avoids data walls through infinite learnable tasks
“The scaling in post-chaining itself is not hitting the wall, and that's because Basically, we went from, like, raw data sets from pre-trained models to infinite amount of tasks that you can teach the model in the post-training world via reinforcement learning.…”
Karina Nguyen Feb 9, 2025 ▶ 9:56
Opinion
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Karina Nguyen Feb 9, 2025 ▶ 10:49
Disclosure
Nguyen: OpenAI built Canvas and Tasks features mostly via synthetic data
“The way we made Canvas and tasks and, like, new, like, product features for HTTP was mostly done by synthetic training.”
Karina Nguyen Feb 9, 2025 ▶ 12:27
Assertion Not checkable as stated
Canvas was OpenAI's first project uniting research and applied engineering early
“Actually, like, Canvas is, like, one of the, I would say, like, the first project at OpenAI where researchers and applied engineers started working together from the very beginning of the product development cycle.”
Karina Nguyen Feb 9, 2025 ▶ 13:46
Disclosure
OpenAI used o1 synthetic data to train Canvas commenting behaviors
“The way we used it is, like, we would use a one model to produce, to, like, simulate, like, use a conversation. Let's say, like, write me a document about XYZ, but then we used a one to, like, produce the document, and then we kind of injected, like, user prom…”
Karina Nguyen Feb 9, 2025 ▶ 17:11
Prediction Not checkable as stated
Rachitsky: Writing model evaluations will become central to AI product management
“Writing these evaluations is going to become increasingly an important part of the job. Of product teams, especially when they're building AI features and working with them.”
Lenny Rachitsky Feb 9, 2025 ▶ 20:40
Disclosure
Nguyen: OpenAI introduced 'model designers' to define behavioral evals
“There's also like new role that we have like model designers. To kind of, like, go through some of the user feedback, maybe, or, like, think of, like, various, like, user conversations that should have triggered, like, under this circumstances, it should trigg…”
Karina Nguyen Feb 9, 2025 ▶ 21:09
Assertion Not checkable as stated
Nguyen: Optimizing AI models constantly causes capability regressions across all labs
“If you optimize the model for this behavior, like, you kind of don't want to, like, brain damage in, like, other areas of intelligence, or, and this is happening, like, all the time in every lab and every, like, research team.”
Karina Nguyen Feb 9, 2025 ▶ 24:20
Assertion Not checkable as stated
OpenAI Built Tasks in Two Months and Canvas in Four to Five
“For tasks, it took, like, I don't know, like, two months or so to go from, like, zero to one, basically. Oh, wow. For canvases was, like, four or five months, I guess. To go from zero to one.”
Karina Nguyen Feb 9, 2025 ▶ 30:48
Insight
Nguyen: Synthetic Data Outperforms Human Data for AI Product Development
“And the reason why I really love, like, synthetic, like, relying purely on synthetic data instead of, like, collecting Data from humans is because it's, like, much more scalable. It's cheap, less than how, like, you literally sample from the model, and you tea…”
Karina Nguyen Feb 9, 2025 ▶ 31:55
Assertion Supported
Nguyen: Small distilled models like Claude 3 Haiku outperform larger predecessors
“Smart, small models are becoming even smarter than, like, large models. And that's because of, like, the distillation research. This happened with, like, Cloud Tree Haiku. I was like working on like post-chaining of like Cloudy Haiku, and I realized it was muc…”
Karina Nguyen Feb 9, 2025 ▶ 38:46
Opinion
Nguyen: ChatGPT still struggles with writing due to creative reasoning limits
“I think it's actually really, really hard to teach the model how to be aesthetic or, like, do, like, visual, really good, like, visual design or, like, how to be extremely creative in the way they write. I think, like, I still think, like, Chai GP kind of suck…”
Karina Nguyen Feb 9, 2025 ▶ 46:02
Insight
Nguyen: AI research progress is bottlenecked by research management
“I actually, like, AI research progress is bottlenecked by, like, management. Like, research management is because you have, like, constrained set of compute, and you need to, like, allocate the compute to the research path that you feel the most Commenced abou…”
Karina Nguyen Feb 9, 2025 ▶ 46:28
Prediction Not checkable as stated
Nguyen: Extended model reasoning will lead to better AI writing quality
“And actually, like, I'm thinking, like, this new paradigm of, like, the models think more should actually lead to, like, better writing in itself.”
Karina Nguyen Feb 9, 2025 ▶ 48:36
Prediction Not checkable as stated
Nguyen: AI is not far from autonomous self-improving product development
“And I don't think, like, we are far away from that kind of, like, self-improvement, models becoming, like, self-improved via, like, then, like, the product development is basically kind of, like, self-improving, like, it's kind of, like, its own, like, organis…”
Karina Nguyen Feb 9, 2025 ▶ 50:47
Opinion
Nguyen: AI will excel at strategy by synthesizing disparate data sources
“Strategy is, like, it's more, like, data analysis and, like coming up with, like, I think what models are really good at is, like, connecting the dots, I think. It's like, okay, if you have user feedback from this source, but you also have an internal, like, d…”
Karina Nguyen Feb 9, 2025 ▶ 51:03
Insight
Nguyen: Claude's librarian personality reflects the craft of its Anthropic creators
“And it's like the reason why cloud has so much more personality and like is more like a librarian. I don't know. Like, I don't know. I am like visualizing Cloud being, like, a librarian. Like very, like, nerdy or something. It's because I feel like it's a refl…”
Karina Nguyen Feb 9, 2025 ▶ 54:47
Opinion
Nguyen: Anthropic excels at prioritization, while OpenAI takes more product risks
“I would say, like, Antarctic, I learned from Antarctic that, like, They're much better at, like, focusing and, like, prioritization or, like, very, very hard, like, very hardcore prioritization, I guess, and they need to do it. Like, but I think, like, OpenAI …”
Karina Nguyen Feb 9, 2025 ▶ 55:53
Prediction Not checkable as stated
Nguyen: Computer-use AI agents will learn and replicate individual browsing behaviors
“The computer use agents, like the model operating the desktop, and you can essentially think of like, you know, new kind of like Experience where the model can learn the way you browse. And from that preference, it can just like browse as just like you.”
Karina Nguyen Feb 9, 2025 ▶ 1:02:40
Insight
Nguyen: Pixel-based perception is much harder to scale than language in AI
“Much of it is, like because right now the models operating on, like, pixels instead of, like, language or whatnot, like, pixels is actually really, really hard for the models because, like, perception or visual perception. I think there's still, like, a lot of…”
Karina Nguyen Feb 9, 2025 ▶ 1:10:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.