Feb 1, 2025 · 1h 6m · latent-space

The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI

Karina Nguyen · 44m spoken Shawn Wang · 11m spoken Alessio Fanelli · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, OpenAI research lead Karina Nguyen explains how post-training, behavioral design, and synthetic data power interactive interfaces like ChatGPT Canvas and Tasks. She shares firsthand engineering and cultural insights from both Anthropic and OpenAI while outlining the future of AI agents and generative operating systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 27.1% of the talking time here. How this is scored →

The hosts as informed peer 4.8 Guest teaching 4.8 Guest disagreement 1.4 The hosts pushing back 1.9
05100:0015:0030:0045:001:00:002:22–5:31 · The hosts as informed peer 3/10 Early Career: Computer Vision, Journalism, and Entering AI Hosts ask introductory questions about Karina's background transitioning from computer vision in journalism at Berkeley to AI labs. Karina gently clarifies that her work was reporting for publications rather than taught by Pulitzer-winning professors.5:31–9:16 · The hosts as informed peer 4/10 Pioneering Products at Anthropic: Claude in Slack and Claude.ai Karina details building Claude in Slack and creating Claude.ai from scratch under tight deadlines. Swyx shares his experience interviewing at Anthropic, and Karina explains why safety and hallucination concerns delayed early web UI releases.9:16–11:30 · The hosts as informed peer 4/10 The Conceptual Evolution of Collaborative Workspaces and Canvas Precursors Karina discusses her early 2023 conceptual sketches for shared human-AI workspaces inspired by Tom Riddle's diary. She corrects Swyx's assumption that this was simply Claude Projects, noting product research was rarely prioritized at that time.11:30–17:36 · The hosts as informed peer 6/10 Claude 3 Post-Training, Compute Allocation, and Benchmark Evals Swyx presses Karina on how labs square meticulous dataset curation and eval tracking with chaotic YOLO training runs. Karina reframes the dynamic around compute allocation and ruthless experimental prioritization.17:37–21:56 · The hosts as informed peer 5/10 Prompting Reasoning Models and the Verification Challenge Alessio and Swyx explore prompting strategies for reasoning models like o1. Karina explains that labs rely heavily on external user discovery because emergent behaviors are hard to verify even internally without specialized domain knowledge.21:57–27:37 · The hosts as informed peer 5/10 Behavioral Design: Crafting Model Personas and Balancing Values Karina introduces the concept of behavioral design, comparing persona engineering to crafting video game characters. She breaks down the technical art of balancing contradictory core values like honesty and harmlessness during synthetic data generation.27:37–41:45 · The hosts as informed peer 6/10 Engineering ChatGPT Canvas: Post-Training, Code Diffs, and Product Integration Alessio challenges why Canvas model improvements are kept separate from the base GPT-4o API model, citing transcript tests. Karina breaks down the difficulties of code diffs, behavioral routing, and rapid iteration via dedicated post-trained checkpoints.41:45–49:02 · The hosts as informed peer 5/10 ChatGPT Tasks: Proactive Agent Capabilities and Organizational Workflows Karina describes supervising the ChatGPT Tasks project and formalizing an operational framework connecting product engineers and research scientists. Swyx probes the exact PRD-to-eval development workflow.49:02–56:01 · The hosts as informed peer 6/10 Defining Agents: Trust Building, Collaboration, and Computer Use Swyx challenges hype surrounding computer use agents, citing high latency, high cost, and low accuracy. Karina argues that direct collaboration in UI workspaces is a prerequisite to establishing the trust necessary for full computer delegation.56:01–1:00:20 · The hosts as informed peer 5/10 The Shift to Generative Operating Systems and Dynamic User Interfaces Alessio and Karina discuss future generative operating systems where dynamic UIs and on-the-fly React components replace direct website navigation, using expense reporting as an agentic benchmark.1:00:21–1:03:46 · The hosts as informed peer 4/10 Culture and Leadership: Comparing OpenAI and Anthropic Karina contrasts the cultures of Anthropic and OpenAI, highlighting Anthropic's structured enterprise focus against OpenAI's rapid risk-taking and bottom-up resource reallocation.2:22–5:31 · Guest teaching 3/10 Early Career: Computer Vision, Journalism, and Entering AI Hosts ask introductory questions about Karina's background transitioning from computer vision in journalism at Berkeley to AI labs. Karina gently clarifies that her work was reporting for publications rather than taught by Pulitzer-winning professors.5:31–9:16 · Guest teaching 4/10 Pioneering Products at Anthropic: Claude in Slack and Claude.ai Karina details building Claude in Slack and creating Claude.ai from scratch under tight deadlines. Swyx shares his experience interviewing at Anthropic, and Karina explains why safety and hallucination concerns delayed early web UI releases.9:16–11:30 · Guest teaching 4/10 The Conceptual Evolution of Collaborative Workspaces and Canvas Precursors Karina discusses her early 2023 conceptual sketches for shared human-AI workspaces inspired by Tom Riddle's diary. She corrects Swyx's assumption that this was simply Claude Projects, noting product research was rarely prioritized at that time.11:30–17:36 · Guest teaching 5/10 Claude 3 Post-Training, Compute Allocation, and Benchmark Evals Swyx presses Karina on how labs square meticulous dataset curation and eval tracking with chaotic YOLO training runs. Karina reframes the dynamic around compute allocation and ruthless experimental prioritization.17:37–21:56 · Guest teaching 5/10 Prompting Reasoning Models and the Verification Challenge Alessio and Swyx explore prompting strategies for reasoning models like o1. Karina explains that labs rely heavily on external user discovery because emergent behaviors are hard to verify even internally without specialized domain knowledge.21:57–27:37 · Guest teaching 5/10 Behavioral Design: Crafting Model Personas and Balancing Values Karina introduces the concept of behavioral design, comparing persona engineering to crafting video game characters. She breaks down the technical art of balancing contradictory core values like honesty and harmlessness during synthetic data generation.27:37–41:45 · Guest teaching 6/10 Engineering ChatGPT Canvas: Post-Training, Code Diffs, and Product Integration Alessio challenges why Canvas model improvements are kept separate from the base GPT-4o API model, citing transcript tests. Karina breaks down the difficulties of code diffs, behavioral routing, and rapid iteration via dedicated post-trained checkpoints.41:45–49:02 · Guest teaching 5/10 ChatGPT Tasks: Proactive Agent Capabilities and Organizational Workflows Karina describes supervising the ChatGPT Tasks project and formalizing an operational framework connecting product engineers and research scientists. Swyx probes the exact PRD-to-eval development workflow.49:02–56:01 · Guest teaching 6/10 Defining Agents: Trust Building, Collaboration, and Computer Use Swyx challenges hype surrounding computer use agents, citing high latency, high cost, and low accuracy. Karina argues that direct collaboration in UI workspaces is a prerequisite to establishing the trust necessary for full computer delegation.56:01–1:00:20 · Guest teaching 5/10 The Shift to Generative Operating Systems and Dynamic User Interfaces Alessio and Karina discuss future generative operating systems where dynamic UIs and on-the-fly React components replace direct website navigation, using expense reporting as an agentic benchmark.1:00:21–1:03:46 · Guest teaching 5/10 Culture and Leadership: Comparing OpenAI and Anthropic Karina contrasts the cultures of Anthropic and OpenAI, highlighting Anthropic's structured enterprise focus against OpenAI's rapid risk-taking and bottom-up resource reallocation.2:22–5:31 · Guest disagreement 1/10 Early Career: Computer Vision, Journalism, and Entering AI Hosts ask introductory questions about Karina's background transitioning from computer vision in journalism at Berkeley to AI labs. Karina gently clarifies that her work was reporting for publications rather than taught by Pulitzer-winning professors.5:31–9:16 · Guest disagreement 1/10 Pioneering Products at Anthropic: Claude in Slack and Claude.ai Karina details building Claude in Slack and creating Claude.ai from scratch under tight deadlines. Swyx shares his experience interviewing at Anthropic, and Karina explains why safety and hallucination concerns delayed early web UI releases.9:16–11:30 · Guest disagreement 2/10 The Conceptual Evolution of Collaborative Workspaces and Canvas Precursors Karina discusses her early 2023 conceptual sketches for shared human-AI workspaces inspired by Tom Riddle's diary. She corrects Swyx's assumption that this was simply Claude Projects, noting product research was rarely prioritized at that time.11:30–17:36 · Guest disagreement 2/10 Claude 3 Post-Training, Compute Allocation, and Benchmark Evals Swyx presses Karina on how labs square meticulous dataset curation and eval tracking with chaotic YOLO training runs. Karina reframes the dynamic around compute allocation and ruthless experimental prioritization.17:37–21:56 · Guest disagreement 1/10 Prompting Reasoning Models and the Verification Challenge Alessio and Swyx explore prompting strategies for reasoning models like o1. Karina explains that labs rely heavily on external user discovery because emergent behaviors are hard to verify even internally without specialized domain knowledge.21:57–27:37 · Guest disagreement 1/10 Behavioral Design: Crafting Model Personas and Balancing Values Karina introduces the concept of behavioral design, comparing persona engineering to crafting video game characters. She breaks down the technical art of balancing contradictory core values like honesty and harmlessness during synthetic data generation.27:37–41:45 · Guest disagreement 2/10 Engineering ChatGPT Canvas: Post-Training, Code Diffs, and Product Integration Alessio challenges why Canvas model improvements are kept separate from the base GPT-4o API model, citing transcript tests. Karina breaks down the difficulties of code diffs, behavioral routing, and rapid iteration via dedicated post-trained checkpoints.41:45–49:02 · Guest disagreement 1/10 ChatGPT Tasks: Proactive Agent Capabilities and Organizational Workflows Karina describes supervising the ChatGPT Tasks project and formalizing an operational framework connecting product engineers and research scientists. Swyx probes the exact PRD-to-eval development workflow.49:02–56:01 · Guest disagreement 2/10 Defining Agents: Trust Building, Collaboration, and Computer Use Swyx challenges hype surrounding computer use agents, citing high latency, high cost, and low accuracy. Karina argues that direct collaboration in UI workspaces is a prerequisite to establishing the trust necessary for full computer delegation.56:01–1:00:20 · Guest disagreement 1/10 The Shift to Generative Operating Systems and Dynamic User Interfaces Alessio and Karina discuss future generative operating systems where dynamic UIs and on-the-fly React components replace direct website navigation, using expense reporting as an agentic benchmark.1:00:21–1:03:46 · Guest disagreement 1/10 Culture and Leadership: Comparing OpenAI and Anthropic Karina contrasts the cultures of Anthropic and OpenAI, highlighting Anthropic's structured enterprise focus against OpenAI's rapid risk-taking and bottom-up resource reallocation.2:22–5:31 · The hosts pushing back 1/10 Early Career: Computer Vision, Journalism, and Entering AI Hosts ask introductory questions about Karina's background transitioning from computer vision in journalism at Berkeley to AI labs. Karina gently clarifies that her work was reporting for publications rather than taught by Pulitzer-winning professors.5:31–9:16 · The hosts pushing back 1/10 Pioneering Products at Anthropic: Claude in Slack and Claude.ai Karina details building Claude in Slack and creating Claude.ai from scratch under tight deadlines. Swyx shares his experience interviewing at Anthropic, and Karina explains why safety and hallucination concerns delayed early web UI releases.9:16–11:30 · The hosts pushing back 1/10 The Conceptual Evolution of Collaborative Workspaces and Canvas Precursors Karina discusses her early 2023 conceptual sketches for shared human-AI workspaces inspired by Tom Riddle's diary. She corrects Swyx's assumption that this was simply Claude Projects, noting product research was rarely prioritized at that time.11:30–17:36 · The hosts pushing back 4/10 Claude 3 Post-Training, Compute Allocation, and Benchmark Evals Swyx presses Karina on how labs square meticulous dataset curation and eval tracking with chaotic YOLO training runs. Karina reframes the dynamic around compute allocation and ruthless experimental prioritization.17:37–21:56 · The hosts pushing back 2/10 Prompting Reasoning Models and the Verification Challenge Alessio and Swyx explore prompting strategies for reasoning models like o1. Karina explains that labs rely heavily on external user discovery because emergent behaviors are hard to verify even internally without specialized domain knowledge.21:57–27:37 · The hosts pushing back 1/10 Behavioral Design: Crafting Model Personas and Balancing Values Karina introduces the concept of behavioral design, comparing persona engineering to crafting video game characters. She breaks down the technical art of balancing contradictory core values like honesty and harmlessness during synthetic data generation.27:37–41:45 · The hosts pushing back 3/10 Engineering ChatGPT Canvas: Post-Training, Code Diffs, and Product Integration Alessio challenges why Canvas model improvements are kept separate from the base GPT-4o API model, citing transcript tests. Karina breaks down the difficulties of code diffs, behavioral routing, and rapid iteration via dedicated post-trained checkpoints.41:45–49:02 · The hosts pushing back 2/10 ChatGPT Tasks: Proactive Agent Capabilities and Organizational Workflows Karina describes supervising the ChatGPT Tasks project and formalizing an operational framework connecting product engineers and research scientists. Swyx probes the exact PRD-to-eval development workflow.49:02–56:01 · The hosts pushing back 4/10 Defining Agents: Trust Building, Collaboration, and Computer Use Swyx challenges hype surrounding computer use agents, citing high latency, high cost, and low accuracy. Karina argues that direct collaboration in UI workspaces is a prerequisite to establishing the trust necessary for full computer delegation.56:01–1:00:20 · The hosts pushing back 1/10 The Shift to Generative Operating Systems and Dynamic User Interfaces Alessio and Karina discuss future generative operating systems where dynamic UIs and on-the-fly React components replace direct website navigation, using expense reporting as an agentic benchmark.1:00:21–1:03:46 · The hosts pushing back 1/10 Culture and Leadership: Comparing OpenAI and Anthropic Karina contrasts the cultures of Anthropic and OpenAI, highlighting Anthropic's structured enterprise focus against OpenAI's rapid risk-taking and bottom-up resource reallocation.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35.1% · guest 64.9%0:00 · the hosts 35.1% · guest 64.9%3:00 · the hosts 11.6% · guest 88.4%3:00 · the hosts 11.6% · guest 88.4%6:00 · the hosts 21.9% · guest 78.1%6:00 · the hosts 21.9% · guest 78.1%9:00 · the hosts 23.1% · guest 76.9%9:00 · the hosts 23.1% · guest 76.9%12:00 · the hosts 26.3% · guest 73.7%12:00 · the hosts 26.3% · guest 73.7%15:00 · the hosts 28.3% · guest 71.7%15:00 · the hosts 28.3% · guest 71.7%18:00 · the hosts 44% · guest 56%18:00 · the hosts 44% · guest 56%21:00 · the hosts 21.5% · guest 78.5%21:00 · the hosts 21.5% · guest 78.5%24:00 · the hosts 18.7% · guest 81.3%24:00 · the hosts 18.7% · guest 81.3%27:00 · the hosts 22% · guest 78%27:00 · the hosts 22% · guest 78%30:00 · the hosts 2.5% · guest 97.5%30:00 · the hosts 2.5% · guest 97.5%33:00 · the hosts 53% · guest 47%33:00 · the hosts 53% · guest 47%36:00 · the hosts 66% · guest 34%36:00 · the hosts 66% · guest 34%39:00 · the hosts 10.8% · guest 89.2%39:00 · the hosts 10.8% · guest 89.2%42:00 · the hosts 13.4% · guest 86.6%42:00 · the hosts 13.4% · guest 86.6%45:00 · the hosts 28.7% · guest 71.3%45:00 · the hosts 28.7% · guest 71.3%48:00 · the hosts 29.8% · guest 70.2%48:00 · the hosts 29.8% · guest 70.2%51:00 · the hosts 17% · guest 83%51:00 · the hosts 17% · guest 83%54:00 · the hosts 44.5% · guest 55.5%54:00 · the hosts 44.5% · guest 55.5%57:00 · the hosts 36.5% · guest 63.5%57:00 · the hosts 36.5% · guest 63.5%1:00:00 · the hosts 19.1% · guest 80.9%1:00:00 · the hosts 19.1% · guest 80.9%1:03:00 · the hosts 22% · guest 78%1:03:00 · the hosts 22% · guest 78%1:06:00 · the hosts 54.2% · guest 45.8%1:06:00 · the hosts 54.2% · guest 45.8%
Sharpest disagreement ▶ 9:44 Clarifying workspace origin vs Claude Projects

Karina rejects Swyx's attempt to equate her early collaborative document prototypes with Anthropic's later Claude Projects feature, clarifying the distinct evolutionary timeline.

Hardest push from the hosts ▶ 55:02 Host skepticism regarding computer use viability

Swyx directly expresses heavy skepticism regarding computer use agents, highlighting serious issues with slowness, cost, and imprecise pixel-level accuracy.

Biggest teaching moment ▶ 23:46 The craft of behavioral design and trade-offs

Karina educates the hosts on how behavioral design functions as a rigorous discipline of decomposing conflicting values like helpfulness and harmlessness via synthetic data generation.

The host holds their own ▶ 34:06 Empirical comparison of API vs Canvas output

Alessio demonstrates technical domain expertise by comparing outputs from his custom podcast application on GPT-4o against Canvas, challenging the host lab's model integration strategy.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Early Career: Computer Vision, Journalism, and Entering AI 3311 Hosts ask introductory questions about Karina's background transitioning from computer vision in journalism at Berkeley to AI labs. Karina gently clarifies that her work was reporting for publications rather than taught by Pulitzer-winning professors.
Pioneering Products at Anthropic: Claude in Slack and Claude.ai 4411 Karina details building Claude in Slack and creating Claude.ai from scratch under tight deadlines. Swyx shares his experience interviewing at Anthropic, and Karina explains why safety and hallucination concerns delayed early web UI releases.
The Conceptual Evolution of Collaborative Workspaces and Canvas Precursors 4421 Karina discusses her early 2023 conceptual sketches for shared human-AI workspaces inspired by Tom Riddle's diary. She corrects Swyx's assumption that this was simply Claude Projects, noting product research was rarely prioritized at that time.
Claude 3 Post-Training, Compute Allocation, and Benchmark Evals 6524 Swyx presses Karina on how labs square meticulous dataset curation and eval tracking with chaotic YOLO training runs. Karina reframes the dynamic around compute allocation and ruthless experimental prioritization.
Prompting Reasoning Models and the Verification Challenge 5512 Alessio and Swyx explore prompting strategies for reasoning models like o1. Karina explains that labs rely heavily on external user discovery because emergent behaviors are hard to verify even internally without specialized domain knowledge.
Behavioral Design: Crafting Model Personas and Balancing Values 5511 Karina introduces the concept of behavioral design, comparing persona engineering to crafting video game characters. She breaks down the technical art of balancing contradictory core values like honesty and harmlessness during synthetic data generation.
Engineering ChatGPT Canvas: Post-Training, Code Diffs, and Product Integration 6623 Alessio challenges why Canvas model improvements are kept separate from the base GPT-4o API model, citing transcript tests. Karina breaks down the difficulties of code diffs, behavioral routing, and rapid iteration via dedicated post-trained checkpoints.
ChatGPT Tasks: Proactive Agent Capabilities and Organizational Workflows 5512 Karina describes supervising the ChatGPT Tasks project and formalizing an operational framework connecting product engineers and research scientists. Swyx probes the exact PRD-to-eval development workflow.
Defining Agents: Trust Building, Collaboration, and Computer Use 6624 Swyx challenges hype surrounding computer use agents, citing high latency, high cost, and low accuracy. Karina argues that direct collaboration in UI workspaces is a prerequisite to establishing the trust necessary for full computer delegation.
The Shift to Generative Operating Systems and Dynamic User Interfaces 5511 Alessio and Karina discuss future generative operating systems where dynamic UIs and on-the-fly React components replace direct website navigation, using expense reporting as an agentic benchmark.
Culture and Leadership: Comparing OpenAI and Anthropic 4511 Karina contrasts the cultures of Anthropic and OpenAI, highlighting Anthropic's structured enterprise focus against OpenAI's rapid risk-taking and bottom-up resource reallocation.

Statements from this episode (30)

Assertion Not checkable as stated
Nguyen: Writing and Coding Are Most Common Use Cases for Canvas
“So for Canvas, for example, one of the most common use cases is basically writing and coding”
Karina Nguyen Feb 1, 2025 ▶ 1:21
Prediction Not checkable as stated
Nguyen: Canvas and Tasks Will Evolve ChatGPT into Something Completely New
“There are different types of like. Features like Canvas, tasks, but all those components that go, they compose together to evolve ChatGPT into something completely new, I think, in the new year.”
Karina Nguyen Feb 1, 2025 ▶ 1:58
Assertion Not checkable as stated
Anthropic built commercial products to self-fund AI safety research
“And that was a time when Antarctic Anthropik really decided to, like, do more product-y related things, and the vision was like, we need to, like, fund research, and, like, building product is, like, the best way to, like, fund safety research”
Karina Nguyen Feb 1, 2025 ▶ 5:45
Assertion Not checkable as stated
Nguyen wrote Claude.ai's first 50,000 lines of code unreviewed
“Yeah, like I think like the first like 50,000 code of lines without any reviews at that time because there's no one. Yeah, it was like very small team. It was like six, seven team who we were called a deployment team.”
Karina Nguyen Feb 1, 2025 ▶ 7:43
Assertion Not checkable as stated
Anthropic delayed web UI due to Claude 1.3 hallucinations
“And I think, like, at that time, Cloud 1.3 I.E. Had a lot of hallucinations, actually. So I think there was, like, one of the concerns is, like, I don't think, like, the leadership was convinced, had a conviction that this is the model that you need to, like, …”
Karina Nguyen Feb 1, 2025 ▶ 8:42
Disclosure
Nguyen: Explored a Claude collaborative workspace concept at Anthropic in 2023
“I was working on something similar to, like, Canvas-y, but for Claude at that time, in, like, twenty-twenty-three, it was the same similar idea of, like, Claude workspace where a human and a Claude could have, like, a shared workspace which is like a document.”
Karina Nguyen Feb 1, 2025 ▶ 9:27
What-if
Nguyen: AI canvas interfaces could have happened two years earlier
“I think like those ideas could have happened like two years ago. Just like maybe, I don't think it was like a priority at that time. It was like very unclear. I think like AI landscape at that time was very nascent, if that makes sense.”
Karina Nguyen Feb 1, 2025 ▶ 10:49
Assertion Not checkable as stated
Nguyen: Anthropic Claude 3 post-training team had only 10-12 people
“I was a part of the post-training fine-tuning team. We only had, like, what, like, 10, 12 people involved”
Karina Nguyen Feb 1, 2025 ▶ 11:49
Assertion Supported
Nguyen: Anthropic was first AI lab to publish GPQA benchmark numbers
“I think it was like the first, I think we were the first lab, like, Antarctica was the first lab to, like, run. Publish GPQA, like, numbers”
Karina Nguyen Feb 1, 2025 ▶ 14:37
Insight
Nguyen: AI model card benchmark numbers are never apples-to-apples across labs
“None of the numbers are, like, apples to apples. So you actually need to, like, go back to, like, I don't know, like, GPT-E for model card and, like, read the appendix just to, like, make sure that, like, The settings were the same as you're running the settin…”
Karina Nguyen Feb 1, 2025 ▶ 15:11
Assertion Not checkable as stated
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Karina Nguyen Feb 1, 2025 ▶ 16:39
Insight
Karina Nguyen: OpenAI o1 excels when given explicit hard constraints
“If you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and match like the candidate that is most like, fulfill the criteria that you…”
Karina Nguyen Feb 1, 2025 ▶ 18:13
Insight
Swix: AI engineering exists because labs crowdsource emergent capability discovery
“The reason that AI engineering can exist outside of the model labs is because the model labs release Models with capabilities that they don't even fully know because you never train specifically for it. It's emergent. And you can rely on basically crowdsourcin…”
Shawn Wang Feb 1, 2025 ▶ 20:57
Opinion
Karina Nguyen: Verification difficulty makes alignment crucial for reasoning models
“The question of like alignment is actually more important for this like complex reasoning models to like, how do we help humans to like verify the outputs of these models is quite important.”
Karina Nguyen Feb 1, 2025 ▶ 21:41
Disclosure
Nguyen: Claude 2's distinct personality was unintentional until Claude 3
“People said, like, Cloud II is, like, so much better at, like, writing and, like, has a certain personality, even though it was, like, unintentional at all. And we did not pay that much attention and didn't know even how to, like, productionize this property o…”
Karina Nguyen Feb 1, 2025 ▶ 26:16
Disclosure
Karina Nguyen: OpenAI Retrained GPT-4o to Handle Canvas Edge Cases
“The only way to like fix some of the edge cases is actually through post training. So we actually, what we did was actually retrain the entire full O plus our canvas stuff.”
Karina Nguyen Feb 1, 2025 ▶ 30:17
Opinion
Nguyen: Original ChatGPT Canvas Beta Model Was More Creative Than GPT-4o Canvas
“I would say like the original better model that we released this canvas was actually much more creative than even right now when I use like for, oh, this canvas”
Karina Nguyen Feb 1, 2025 ▶ 35:05
Insight
Swix: ChatGPT Canvas Inverts Google Docs and Gemini's Interface Architecture
“It's basically an inversion of what Google Docs is, wants to do with Gemini. It's like Google Docs on the main screen and then Gemini on the side. And right, whatnot, what ChatGPT has done is Do the chat thing first, and then the docs on the side. But it's kin…”
Shawn Wang Feb 1, 2025 ▶ 36:56
Prediction Open · timeframe Feb 2028
Nguyen: ChatGPT Will Evolve Into an Interface That Morphs Based on User Intent
“Chat CPT evolves into this Blank interface, which can morph itself in whatever you trying, like the model should try to like derive your true intent and then modify the interface based on your intent. And then if you like writing, it should become like the mos…”
Karina Nguyen Feb 1, 2025 ▶ 37:53
Insight
Nguyen: Full Document Rewrites Yield Higher Model Accuracy Than Code Diffs
“We didn't know that, like, code diffs was very difficult for a model, for example. Again, it's like, do we go back to, like, fundamentally improve, like, code diffs as a model capability? Or do you, like, do a workaround where the model will just, like, rewrit…”
Karina Nguyen Feb 1, 2025 ▶ 39:19
Assertion Not checkable as stated
Nguyen: OpenAI developed ChatGPT Tasks in under two months
“And actually, tasks was developed less than, like, two months. So if Canvas took, like, I don't know, four months, then tasks took, like, two months.”
Karina Nguyen Feb 1, 2025 ▶ 42:35
Insight
Nguyen: AI model training requires evals where prompted baselines fail
“Prototype was prompted baseline. It's all, all, everything starts with, like, prompted baseline, and then, like, we craft, like, certain, like, evaluations that we want to, like, capture, that we want to, like, measure progress, at least, for the model, and th…”
Karina Nguyen Feb 1, 2025 ▶ 44:37
Prediction Not checkable as stated
Nguyen: AI models will evolve to proactively suggest recurring user workflows
“I think that ideally we learn from like the user behavior and ideally the model will just be more proactive in suggesting of like Oh, I can either do this for you every day because I've observed that you do that every day or something. So it's like more become…”
Karina Nguyen Feb 1, 2025 ▶ 46:46
Insight
Nguyen: User collaboration is the key milestone before full AI delegation
“Sometimes I feel like a lot of researchers or, like, people in the AI community are, like, so into, like, yeah, agents, delegate everything, like, blah, blah. But, like, on the way towards that, I think, like, collaboration is actually one of the main roadbloc…”
Karina Nguyen Feb 1, 2025 ▶ 51:30
Assertion Not checkable as stated
Nguyen: Computer-use agent execution lagged two to three years behind ideation
“Computer using, oh, agents using desktop or like your computer is like the delegation part. Like when you might want to like delegate an agent to like order a book for me or like order a flight or like search for a flight and then order things for me. And I fe…”
Karina Nguyen Feb 1, 2025 ▶ 52:28
Opinion
Swix: Bearish on computer-use AI agents due to cost, speed, and accuracy
“I have been very bearish in computer use because they're slow. They're expensive. They're imprecise. Like the accuracy is horrible. Still, even with Anthropix new stuff, I'm really waiting to see what opening I might do to change my opinions.”
Shawn Wang Feb 1, 2025 ▶ 55:02
Prediction Not checkable as stated
Nguyen: Website clicks will drop as internet access shifts to AI models
“In my opinion, like, people in, like, few years will click On, like, websites way less. I want to see the plot of, like, website clicks over time, but then my prediction is, like, it will go down and, like, people's access to the internet will be through the m…”
Karina Nguyen Feb 1, 2025 ▶ 56:35
Prediction Held up
Fanelli: Computer use agents will likely automate expense reports within a year
“It's not, you cannot actually do it today, but it feels like a tractable problem, you know, that probably by the end of the year we should be able to do it.”
Alessio Fanelli Feb 1, 2025 ▶ 57:43
Opinion
Nguyen: OpenAI takes bigger product risks while Anthropic focuses on enterprise
“OpenAI and Anthropik is different in terms of like more like maybe like product mindset. Maybe OpenAI is much more willing to take some of the product risks and explore different bets. And I think Anthropik is much more focused and they have, I think it's fine…”
Karina Nguyen Feb 1, 2025 ▶ 1:01:06
Insight
Nguyen: AI progress is bottlenecked by human interface creativity
“I feel like we are bottlenecked by like human creativity on like completely changing the way we think about the internet or like some of the way we think about software, like AI right now pushes us to like rethink everything that we've done before in my view.”
Karina Nguyen Feb 1, 2025 ▶ 1:04:56
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.