Aug 8, 2025 · 42m · a16z

GPT-5 and Agents Breakdown – w/ OpenAI Researchers Isa Fulford & Christina Kim

Christina Kim · 14m spoken Isa Fulford · 14m spoken Sarah Wang · 6m spoken Erik Torenberg · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI researchers Isa Fulford and Christina Kim join hosts Erik Torenberg and Sarah Wang on the a16z podcast to discuss the milestone launch of GPT-5. They explore advancements in coding capabilities, reinforcement learning, autonomous agents, post-training data curation, and OpenAI's ongoing organizational evolution.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 8.8% of the talking time here. How this is scored →

The host as informed peer 3.1 Guest teaching 4.9 Guest disagreement 0.1 The host pushing back 1.1
05100:0015:0030:001:04–4:10 · The host as informed peer 3/10 Christina Kim’s History at OpenAI and Launch Day Reactions The hosts ask foundational questions regarding the launch of GPT-5 and past history at OpenAI. Christina explains the lineage from WebGPT to ChatGPT and details improvements in front-end web development capabilities.4:10–6:50 · The host as informed peer 3/10 Balancing Model Behavior, Engagement, and Hallucination Reduction Hosts enquire about sycophancy, engagement trade-offs, and hallucination reductions. Christina highlights post-training as an art requiring reward trade-offs and explains how reasoning thinking time reduces blurting out hallucinations.6:50–10:27 · The host as informed peer 4/10 Leveraging Existing Products and Reinforcement Learning Data Efficiency Sarah asks how existing products inform model releases, prompting Isa to explain data efficiency in reinforcement learning. Christina discusses how metric evaluation is shifting toward real-world usage rather than saturated public benchmarks.10:27–12:44 · The host as informed peer 4/10 Crafting Capability Evals and Balancing Specialization Sarah probes how OpenAI handles benchmark saturation and prioritizes capability trade-offs between expert and general domains. Guests elaborate on working backward from desired capabilities and constructing internal evals.12:44–14:54 · The host as informed peer 3/10 RL Reasoning Progress and the Criticality of High-Quality Data Isa describes how observing RL progress in math and physics unlocked agent reasoning capabilities. When asked about architectural versus data contributions, Christina declares herself data pilled.14:54–16:59 · The host as informed peer 3/10 Challenges in RL Environments and Broad Tool Execution Sarah asks about RL environment bottlenecks and labor automation. The guests discuss task curation constraints and the theoretical reach of broad browser/terminal tools.16:59–19:43 · The host as informed peer 2/10 Emotional Resonance in Creative Writing and Everyday Prompts Erik asks about creative writing improvements and public adaptation to rapidly advancing technology. Christina reflects on human propensity to quickly take revolutionary tools like pocket wizards for granted.19:43–23:16 · The host as informed peer 3/10 Boundaries of Model Execution and Human-in-the-Loop Oversight Hosts ask what the model categorically cannot do and query future capabilities like end-to-end DevOps. Guests emphasize safety constraints around irreversible actions and the horizon for multi-hour autonomous tasks.23:16–27:11 · The host as informed peer 4/10 Defining Agentic Workflows and Asynchronous Execution Patience Sarah asks Isa to define agents and pushes on user willingness to accept asynchronous latency over immediate answers. Isa discusses how removing latency constraints enabled deep multi-step research.27:11–33:04 · The host as informed peer 3/10 Calibrating Thinking Time, Output Length, and Reliability Bottlenecks Isa and Christina discuss user psychological expectations around output length and thinking duration. Christina provides a clear technical explanation of mid-training's role between pre-training and post-training.33:04–36:25 · The host as informed peer 2/10 ChatGPT's 50-Person Beta Test Origins and Joining OpenAI Erik labels Christina an AI historian and asks for reflections on OpenAI's early days. Christina shares anecdotes about testing early chatbots on her roommates, while Isa recalls being a power user prior to joining.36:25–39:35 · The host as informed peer 3/10 OpenAI's Growth from 200 to Thousands and Research Integration Sarah asks about organizational changes as OpenAI scaled from 200 employees to thousands. Guests discuss preserving a high-agency startup culture and tight integration between research and applied product teams.39:35–41:43 · The host as informed peer 3/10 Dual Target Audience and Defining Researcher Taste via Simplicity Sarah asks about OpenAI balancing consumer and enterprise identities and queries the definition of taste. Isa defines research taste as simplifying problems to their most straightforward, elegant elements.1:04–4:10 · Guest teaching 5/10 Christina Kim’s History at OpenAI and Launch Day Reactions The hosts ask foundational questions regarding the launch of GPT-5 and past history at OpenAI. Christina explains the lineage from WebGPT to ChatGPT and details improvements in front-end web development capabilities.4:10–6:50 · Guest teaching 5/10 Balancing Model Behavior, Engagement, and Hallucination Reduction Hosts enquire about sycophancy, engagement trade-offs, and hallucination reductions. Christina highlights post-training as an art requiring reward trade-offs and explains how reasoning thinking time reduces blurting out hallucinations.6:50–10:27 · Guest teaching 5/10 Leveraging Existing Products and Reinforcement Learning Data Efficiency Sarah asks how existing products inform model releases, prompting Isa to explain data efficiency in reinforcement learning. Christina discusses how metric evaluation is shifting toward real-world usage rather than saturated public benchmarks.10:27–12:44 · Guest teaching 5/10 Crafting Capability Evals and Balancing Specialization Sarah probes how OpenAI handles benchmark saturation and prioritizes capability trade-offs between expert and general domains. Guests elaborate on working backward from desired capabilities and constructing internal evals.12:44–14:54 · Guest teaching 5/10 RL Reasoning Progress and the Criticality of High-Quality Data Isa describes how observing RL progress in math and physics unlocked agent reasoning capabilities. When asked about architectural versus data contributions, Christina declares herself data pilled.14:54–16:59 · Guest teaching 5/10 Challenges in RL Environments and Broad Tool Execution Sarah asks about RL environment bottlenecks and labor automation. The guests discuss task curation constraints and the theoretical reach of broad browser/terminal tools.16:59–19:43 · Guest teaching 4/10 Emotional Resonance in Creative Writing and Everyday Prompts Erik asks about creative writing improvements and public adaptation to rapidly advancing technology. Christina reflects on human propensity to quickly take revolutionary tools like pocket wizards for granted.19:43–23:16 · Guest teaching 5/10 Boundaries of Model Execution and Human-in-the-Loop Oversight Hosts ask what the model categorically cannot do and query future capabilities like end-to-end DevOps. Guests emphasize safety constraints around irreversible actions and the horizon for multi-hour autonomous tasks.23:16–27:11 · Guest teaching 5/10 Defining Agentic Workflows and Asynchronous Execution Patience Sarah asks Isa to define agents and pushes on user willingness to accept asynchronous latency over immediate answers. Isa discusses how removing latency constraints enabled deep multi-step research.27:11–33:04 · Guest teaching 6/10 Calibrating Thinking Time, Output Length, and Reliability Bottlenecks Isa and Christina discuss user psychological expectations around output length and thinking duration. Christina provides a clear technical explanation of mid-training's role between pre-training and post-training.33:04–36:25 · Guest teaching 5/10 ChatGPT's 50-Person Beta Test Origins and Joining OpenAI Erik labels Christina an AI historian and asks for reflections on OpenAI's early days. Christina shares anecdotes about testing early chatbots on her roommates, while Isa recalls being a power user prior to joining.36:25–39:35 · Guest teaching 4/10 OpenAI's Growth from 200 to Thousands and Research Integration Sarah asks about organizational changes as OpenAI scaled from 200 employees to thousands. Guests discuss preserving a high-agency startup culture and tight integration between research and applied product teams.39:35–41:43 · Guest teaching 5/10 Dual Target Audience and Defining Researcher Taste via Simplicity Sarah asks about OpenAI balancing consumer and enterprise identities and queries the definition of taste. Isa defines research taste as simplifying problems to their most straightforward, elegant elements.1:04–4:10 · Guest disagreement 0/10 Christina Kim’s History at OpenAI and Launch Day Reactions The hosts ask foundational questions regarding the launch of GPT-5 and past history at OpenAI. Christina explains the lineage from WebGPT to ChatGPT and details improvements in front-end web development capabilities.4:10–6:50 · Guest disagreement 0/10 Balancing Model Behavior, Engagement, and Hallucination Reduction Hosts enquire about sycophancy, engagement trade-offs, and hallucination reductions. Christina highlights post-training as an art requiring reward trade-offs and explains how reasoning thinking time reduces blurting out hallucinations.6:50–10:27 · Guest disagreement 0/10 Leveraging Existing Products and Reinforcement Learning Data Efficiency Sarah asks how existing products inform model releases, prompting Isa to explain data efficiency in reinforcement learning. Christina discusses how metric evaluation is shifting toward real-world usage rather than saturated public benchmarks.10:27–12:44 · Guest disagreement 0/10 Crafting Capability Evals and Balancing Specialization Sarah probes how OpenAI handles benchmark saturation and prioritizes capability trade-offs between expert and general domains. Guests elaborate on working backward from desired capabilities and constructing internal evals.12:44–14:54 · Guest disagreement 0/10 RL Reasoning Progress and the Criticality of High-Quality Data Isa describes how observing RL progress in math and physics unlocked agent reasoning capabilities. When asked about architectural versus data contributions, Christina declares herself data pilled.14:54–16:59 · Guest disagreement 0/10 Challenges in RL Environments and Broad Tool Execution Sarah asks about RL environment bottlenecks and labor automation. The guests discuss task curation constraints and the theoretical reach of broad browser/terminal tools.16:59–19:43 · Guest disagreement 0/10 Emotional Resonance in Creative Writing and Everyday Prompts Erik asks about creative writing improvements and public adaptation to rapidly advancing technology. Christina reflects on human propensity to quickly take revolutionary tools like pocket wizards for granted.19:43–23:16 · Guest disagreement 0/10 Boundaries of Model Execution and Human-in-the-Loop Oversight Hosts ask what the model categorically cannot do and query future capabilities like end-to-end DevOps. Guests emphasize safety constraints around irreversible actions and the horizon for multi-hour autonomous tasks.23:16–27:11 · Guest disagreement 0/10 Defining Agentic Workflows and Asynchronous Execution Patience Sarah asks Isa to define agents and pushes on user willingness to accept asynchronous latency over immediate answers. Isa discusses how removing latency constraints enabled deep multi-step research.27:11–33:04 · Guest disagreement 1/10 Calibrating Thinking Time, Output Length, and Reliability Bottlenecks Isa and Christina discuss user psychological expectations around output length and thinking duration. Christina provides a clear technical explanation of mid-training's role between pre-training and post-training.33:04–36:25 · Guest disagreement 0/10 ChatGPT's 50-Person Beta Test Origins and Joining OpenAI Erik labels Christina an AI historian and asks for reflections on OpenAI's early days. Christina shares anecdotes about testing early chatbots on her roommates, while Isa recalls being a power user prior to joining.36:25–39:35 · Guest disagreement 0/10 OpenAI's Growth from 200 to Thousands and Research Integration Sarah asks about organizational changes as OpenAI scaled from 200 employees to thousands. Guests discuss preserving a high-agency startup culture and tight integration between research and applied product teams.39:35–41:43 · Guest disagreement 0/10 Dual Target Audience and Defining Researcher Taste via Simplicity Sarah asks about OpenAI balancing consumer and enterprise identities and queries the definition of taste. Isa defines research taste as simplifying problems to their most straightforward, elegant elements.1:04–4:10 · The host pushing back 1/10 Christina Kim’s History at OpenAI and Launch Day Reactions The hosts ask foundational questions regarding the launch of GPT-5 and past history at OpenAI. Christina explains the lineage from WebGPT to ChatGPT and details improvements in front-end web development capabilities.4:10–6:50 · The host pushing back 1/10 Balancing Model Behavior, Engagement, and Hallucination Reduction Hosts enquire about sycophancy, engagement trade-offs, and hallucination reductions. Christina highlights post-training as an art requiring reward trade-offs and explains how reasoning thinking time reduces blurting out hallucinations.6:50–10:27 · The host pushing back 1/10 Leveraging Existing Products and Reinforcement Learning Data Efficiency Sarah asks how existing products inform model releases, prompting Isa to explain data efficiency in reinforcement learning. Christina discusses how metric evaluation is shifting toward real-world usage rather than saturated public benchmarks.10:27–12:44 · The host pushing back 2/10 Crafting Capability Evals and Balancing Specialization Sarah probes how OpenAI handles benchmark saturation and prioritizes capability trade-offs between expert and general domains. Guests elaborate on working backward from desired capabilities and constructing internal evals.12:44–14:54 · The host pushing back 1/10 RL Reasoning Progress and the Criticality of High-Quality Data Isa describes how observing RL progress in math and physics unlocked agent reasoning capabilities. When asked about architectural versus data contributions, Christina declares herself data pilled.14:54–16:59 · The host pushing back 1/10 Challenges in RL Environments and Broad Tool Execution Sarah asks about RL environment bottlenecks and labor automation. The guests discuss task curation constraints and the theoretical reach of broad browser/terminal tools.16:59–19:43 · The host pushing back 1/10 Emotional Resonance in Creative Writing and Everyday Prompts Erik asks about creative writing improvements and public adaptation to rapidly advancing technology. Christina reflects on human propensity to quickly take revolutionary tools like pocket wizards for granted.19:43–23:16 · The host pushing back 1/10 Boundaries of Model Execution and Human-in-the-Loop Oversight Hosts ask what the model categorically cannot do and query future capabilities like end-to-end DevOps. Guests emphasize safety constraints around irreversible actions and the horizon for multi-hour autonomous tasks.23:16–27:11 · The host pushing back 2/10 Defining Agentic Workflows and Asynchronous Execution Patience Sarah asks Isa to define agents and pushes on user willingness to accept asynchronous latency over immediate answers. Isa discusses how removing latency constraints enabled deep multi-step research.27:11–33:04 · The host pushing back 1/10 Calibrating Thinking Time, Output Length, and Reliability Bottlenecks Isa and Christina discuss user psychological expectations around output length and thinking duration. Christina provides a clear technical explanation of mid-training's role between pre-training and post-training.33:04–36:25 · The host pushing back 1/10 ChatGPT's 50-Person Beta Test Origins and Joining OpenAI Erik labels Christina an AI historian and asks for reflections on OpenAI's early days. Christina shares anecdotes about testing early chatbots on her roommates, while Isa recalls being a power user prior to joining.36:25–39:35 · The host pushing back 1/10 OpenAI's Growth from 200 to Thousands and Research Integration Sarah asks about organizational changes as OpenAI scaled from 200 employees to thousands. Guests discuss preserving a high-agency startup culture and tight integration between research and applied product teams.39:35–41:43 · The host pushing back 1/10 Dual Target Audience and Defining Researcher Taste via Simplicity Sarah asks about OpenAI balancing consumer and enterprise identities and queries the definition of taste. Isa defines research taste as simplifying problems to their most straightforward, elegant elements.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 25.7% · guest 74.3%0:00 · the host 25.7% · guest 74.3%3:00 · the host 3.7% · guest 96.3%3:00 · the host 3.7% · guest 96.3%6:00 · the host 19% · guest 81%6:00 · the host 19% · guest 81%9:00 · the host 1% · guest 99%9:00 · the host 1% · guest 99%12:00 · the host 4.3% · guest 95.7%12:00 · the host 4.3% · guest 95.7%15:00 · the host 2.8% · guest 97.2%15:00 · the host 2.8% · guest 97.2%18:00 · the host 25.1% · guest 74.9%18:00 · the host 25.1% · guest 74.9%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 7.2% · guest 92.8%27:00 · the host 7.2% · guest 92.8%30:00 · the host 9.5% · guest 90.5%30:00 · the host 9.5% · guest 90.5%33:00 · the host 8.9% · guest 91.1%33:00 · the host 8.9% · guest 91.1%36:00 · the host 9.6% · guest 90.4%36:00 · the host 9.6% · guest 90.4%39:00 · the host 6.8% · guest 93.2%39:00 · the host 6.8% · guest 93.2%42:00 · the host 13% · guest 87%42:00 · the host 13% · guest 87%
Sharpest disagreement ▶ 27:42 Correcting user assumptions on report length and thinking time

Isa gently pushes back against user assumptions, noting that longer reports or extended thinking times are not inherently superior indicators of effort or quality.

Hardest push from the host ▶ 25:04 Sarah challenges the async speed vs value tradeoff

Sarah directly questions the industry premise that speed is paramount, pointing out the major paradigm shift where users are suddenly willing to wait minutes for high-value asynchronous results.

Biggest teaching moment ▶ 31:55 Christina explains the function of mid-training

Christina clearly educates the hosts on the distinct architectural position of mid-training as an efficient mechanism to expand model knowledge without re-running massive pre-training clusters.

The host holds their own ▶ 9:42 Sarah brings precise benchmark saturation context

Sarah demonstrates sharp industry insight by citing specific internal quote contexts about benchmark saturation from Greg Brockman to frame her question on internal eval methodologies.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Christina Kim’s History at OpenAI and Launch Day Reactions 3501 The hosts ask foundational questions regarding the launch of GPT-5 and past history at OpenAI. Christina explains the lineage from WebGPT to ChatGPT and details improvements in front-end web development capabilities.
Balancing Model Behavior, Engagement, and Hallucination Reduction 3501 Hosts enquire about sycophancy, engagement trade-offs, and hallucination reductions. Christina highlights post-training as an art requiring reward trade-offs and explains how reasoning thinking time reduces blurting out hallucinations.
Leveraging Existing Products and Reinforcement Learning Data Efficiency 4501 Sarah asks how existing products inform model releases, prompting Isa to explain data efficiency in reinforcement learning. Christina discusses how metric evaluation is shifting toward real-world usage rather than saturated public benchmarks.
Crafting Capability Evals and Balancing Specialization 4502 Sarah probes how OpenAI handles benchmark saturation and prioritizes capability trade-offs between expert and general domains. Guests elaborate on working backward from desired capabilities and constructing internal evals.
RL Reasoning Progress and the Criticality of High-Quality Data 3501 Isa describes how observing RL progress in math and physics unlocked agent reasoning capabilities. When asked about architectural versus data contributions, Christina declares herself data pilled.
Challenges in RL Environments and Broad Tool Execution 3501 Sarah asks about RL environment bottlenecks and labor automation. The guests discuss task curation constraints and the theoretical reach of broad browser/terminal tools.
Emotional Resonance in Creative Writing and Everyday Prompts 2401 Erik asks about creative writing improvements and public adaptation to rapidly advancing technology. Christina reflects on human propensity to quickly take revolutionary tools like pocket wizards for granted.
Boundaries of Model Execution and Human-in-the-Loop Oversight 3501 Hosts ask what the model categorically cannot do and query future capabilities like end-to-end DevOps. Guests emphasize safety constraints around irreversible actions and the horizon for multi-hour autonomous tasks.
Defining Agentic Workflows and Asynchronous Execution Patience 4502 Sarah asks Isa to define agents and pushes on user willingness to accept asynchronous latency over immediate answers. Isa discusses how removing latency constraints enabled deep multi-step research.
Calibrating Thinking Time, Output Length, and Reliability Bottlenecks 3611 Isa and Christina discuss user psychological expectations around output length and thinking duration. Christina provides a clear technical explanation of mid-training's role between pre-training and post-training.
ChatGPT's 50-Person Beta Test Origins and Joining OpenAI 2501 Erik labels Christina an AI historian and asks for reflections on OpenAI's early days. Christina shares anecdotes about testing early chatbots on her roommates, while Isa recalls being a power user prior to joining.
OpenAI's Growth from 200 to Thousands and Research Integration 3401 Sarah asks about organizational changes as OpenAI scaled from 200 employees to thousands. Guests discuss preserving a high-agency startup culture and tight integration between research and applied product teams.
Dual Target Audience and Defining Researcher Taste via Simplicity 3501 Sarah asks about OpenAI balancing consumer and enterprise identities and queries the definition of taste. Isa defines research taste as simplifying problems to their most straightforward, elegant elements.

Statements from this episode (36)

Assertion Not checkable as stated
Kim: GPT-5 internal testers felt insulted by instant answers to hard questions
“I think we hear this with GPT-Five internally when people are testing and they're like, oh, I thought I asked like a really hard question. I feel like a little bit insulted that I thought for like two seconds or like when it doesn't even want to think at all.”
Christina Kim Aug 8, 2025 ▶ 0:27
Assertion Contradicted
Kim: WebGPT was the first large language model to use tools
“I originally worked on WebGPT, which was the original first LLM using tool use.”
Christina Kim Aug 8, 2025 ▶ 1:14
Disclosure
Kim: GPT-5 is a step change for personal coding and writing
“I use it for coding and writing all the time, and it's just a huge stuff change.”
Christina Kim Aug 8, 2025 ▶ 2:13
Opinion
Kim: GPT-5 front-end coding is a massive leap over o3
“If you compare it to O three's front end coding capability, this is just totally next level.”
Christina Kim Aug 8, 2025 ▶ 3:43
Insight
Kim: AI post-training functions more like art than traditional research
“For post-training, what's really f- or one of the reasons I really like post-training is it feels more like an art than maybe even, like, other areas of research, because you kind of have to make all these trade-offs, right?”
Christina Kim Aug 8, 2025 ▶ 4:39
Insight
Kim: Step-by-step reasoning reduces hallucinations in AI models
“When the models are able to take step by step, they actually can like pause before blurting out an answer is kind of what I, it feels like with a lot of the previous models or hallucinations.”
Christina Kim Aug 8, 2025 ▶ 6:00
Opinion
Kim: Competitor coding models lacked compelling price points
“Maybe like previous competitor models were, are good at coding, but the price point is not as exciting.”
Christina Kim Aug 8, 2025 ▶ 6:36
Insight
Isa Fulford: Reinforcement learning for specific model capabilities is data-efficient
“Training a model to be good at a specific capability is very data efficient. You don't need that many examples to teach it something new.”
Isa Fulford Aug 8, 2025 ▶ 7:16
Assertion Contradicted
Isa Fulford: Deep Research was the first AI model to do comprehensive browsing
“Deep Research, it was the first model to do, like, very comprehensive browsing.”
Isa Fulford Aug 8, 2025 ▶ 7:31
Disclosure
Fulford: OpenAI recycles agent model datasets to train frontier reasoning models
“We're able to take the data sets that we've created for The, you know, frontier agent models and then contribute it back to the frontier reasoning models.”
Isa Fulford Aug 8, 2025 ▶ 7:37
What-if
Kim: Building OpenAI's launch demo manually would have taken her a week
“I'd literally, I think that would have honestly taken me, like, a week to actually build, like, fully interactive”
Christina Kim Aug 8, 2025 ▶ 8:26
Prediction Not checkable as stated
Kim: AI prompt-based app generation will spur surge in indie businesses
“I think we're just gonna have a lot more, I would expect, like, maybe a lot more, like, indie type of, like, Businesses built around this because of the fact that, like, you just need to have the idea, write a simple prompt, and then you get the full fledged a…”
Christina Kim Aug 8, 2025 ▶ 8:26
Insight
Kim: Real-world usage will replace saturated benchmarks to measure AI progress
“I feel like we've almost saturated a lot of these evals, and the real, like, metric of, like, how good our models are getting is, I think, gonna be, like, usage, right?”
Christina Kim Aug 8, 2025 ▶ 9:15
Insight
Kim: Designing good evaluations is the best way to motivate AI researchers
“If you want to nerdside someone into working on something, you just need to make a good eval, and then people are going to be so happy to try to hill climb that.”
Christina Kim Aug 8, 2025 ▶ 11:07
Insight
Isa Fulford: OpenAI defies startup wisdom by targeting universal users
“I mean, it's like everything they tell you not to do at a startup is just like your user is anyone.”
Isa Fulford Aug 8, 2025 ▶ 11:44
Assertion Not checkable as stated
Kim: OpenAI's Operator required multimodal base model capabilities to launch
“Because we had been working on computer usage, but I think it was hard to finally get the model to actually, without like the multimodal capabilities to really support it, like you couldn't have something like Operator when it launched.”
Christina Kim Aug 8, 2025 ▶ 13:14
Insight
Fulford: RL breakthroughs in math and coding unlocked functional AI agents
“When we saw the reinforcement learning algorithm working really well on math and physics problems and coding problems, It became pretty clear, like, just from reading through the chain of thought, like, okay, this thing's actually, like, thinking and reasoning…”
Isa Fulford Aug 8, 2025 ▶ 13:37
Insight
Fulford: More efficient AI learning increases the necessity of high-quality data
“Now that we have such an efficient way of learning data is even high quality data is even, even more important.”
Isa Fulford Aug 8, 2025 ▶ 14:43
Insight
Kim: Training tasks and RL environments matter more than algorithmic advances
“Tasks matter more at this point, given the fact that we have such a strong algorithm so I think the data, creating data and figuring out, like, the best tasks to train on is, like, the, One of the big questions we have.”
Christina Kim Aug 8, 2025 ▶ 15:53
Insight
Fulford: AI agents must train on target tasks to reach top performance
“There's some generalization from training on, like, one website to another, but if you want to get really, really good at something, the best thing to do is just, like, train on that exact thing.”
Isa Fulford Aug 8, 2025 ▶ 16:13
Assertion Not checkable as stated
Fulford: ChatGPT agent's browser and terminal access enable most human computer tasks
“The ChatGPT agent, for example, has such a general tool. It has a browser and a terminal, and between those two things, you can basically do most of the tasks that A human does on a computer.”
Isa Fulford Aug 8, 2025 ▶ 16:30
Opinion
Kim: GPT-5's creative writing capability is tender and touching
“That's one of my favorite improvements in GBT five. The writing, I honestly find it's very tender and touching, especially for a lot of the creative writing that we want to do.”
Christina Kim Aug 8, 2025 ▶ 17:04
Prediction Not checkable as stated
Kim: AI will remain approachable even as models surpass human intelligence
“I guess people adapt to things rather quickly, in my opinion, with technology, and it is really easy, and I think because the form factor is so easy, even with, like, new tools like Deep Research and ChatGPT Agent, it's, like, presented in such, like, a, like,…”
Christina Kim Aug 8, 2025 ▶ 19:13
Opinion
Kim: The leap from GPT-4 to GPT-5 is OpenAI's most impressive yet
“Maybe I'm biased, recency biased, but I think to jump to four to five is most impressive for me, because I guess with 3.5 when we first released it, the most common use case for me then also was still just for coding. And, but now, like, Even though four was b…”
Christina Kim Aug 8, 2025 ▶ 20:07
Disclosure
Fulford: OpenAI requires user confirmation before agents execute irreversible actions
“We take a conservative approach, especially with like asking the user for confirmation before doing any kind of action that's irreversible. So like sending an email or ordering something, booking something.”
Isa Fulford Aug 8, 2025 ▶ 20:57
Prediction Not checkable as stated
Fulford: Users will eventually grant AI agents autonomy for bulk actions
“So I think I can imagine quite You know, a number of tasks where you'd want to take, like, bulk actions which you might not be able to do right now because it would last you every single time, but I think as people get more comfortable using these things and a…”
Isa Fulford Aug 8, 2025 ▶ 21:09
Assertion Not checkable as stated
Fulford: Current AI models can execute monitoring given proper harnesses
“I'm sure that you could build something that's, like, monitoring, you know, your Humio or, like, Datadog, whatever. Like, with these current models, it's just, like, setting up the harness, like, to make that possible.”
Isa Fulford Aug 8, 2025 ▶ 22:18
Insight
Isa Fulford: AI user patience quickly shifts from minutes to 30 seconds
“Initially people are like, oh, this is amazing. It's doing all this work. That would have taken me so long, and now people are like, ok, but I want it, now I want it in 30 seconds.”
Isa Fulford Aug 8, 2025 ▶ 26:43
Insight
Fulford: Users wrongly associate longer AI answers with thoroughness
“One thing that's interesting is I think sometimes people just bias to thinking that the longer answer is more, like, thorough, or it's done more work for it, which I don't necessarily think is the case.”
Isa Fulford Aug 8, 2025 ▶ 27:19
Insight
Fulford: OpenAI bootstraps browsing models to generate synthetic training data
“For initial deep research, there's not really any data sets that exist for browsing in the same way that you have a math data set that already exists. So we have to create all this data. But once you have good browsing models or good computer use models, you c…”
Isa Fulford Aug 8, 2025 ▶ 31:30
Insight
Kim: Mid-training updates AI model knowledge without requiring full pre-training runs
“We do it before after pre-training, but before post-training you kind of think of a way to like extend the model's like intelligence without having to do a whole new pre-training run. So this is mostly just focused on data and off of the pre-training models. S…”
Christina Kim Aug 8, 2025 ▶ 32:08
Assertion Not checkable as stated
OpenAI tested early ChatGPT with a 50-person beta group
“We gave early access to about 50 people. Most of those people being, like, people I lived with at the time.”
Christina Kim Aug 8, 2025 ▶ 33:59
Disclosure
Kim: OpenAI originally considered narrowing ChatGPT to a meeting or coding bot
“At the time, we were kind of thinking, like, okay, we kind of have this chatbot. Should we make this, like, a really specific, like, meeting bot type of thing? Do we, like, make it a coding helper?”
Christina Kim Aug 8, 2025 ▶ 34:20
Assertion Supported
Kim: OpenAI grew from 200 employees to several thousand
“It was around like two hundred-ish people, and I think we're close to like a few thousand for sure.”
Christina Kim Aug 8, 2025 ▶ 37:29
Assertion Not checkable as stated
Kim: OpenAI's Deep Research project was originally developed by just two people
“Like, when Isa was working on deep research, it was, like, two people.”
Christina Kim Aug 8, 2025 ▶ 38:16
Insight
Fulford: Good researcher taste means simplifying problems to the most basic approach
“I think also I've been surprised by how often the thing that is, is the most simple, like easy to explain is the thing that works the best. And so sometimes it's like sound, seems very obvious, but It, you know, it's quite hard to get the details of something …”
Isa Fulford Aug 8, 2025 ▶ 40:50
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.