Oct 7, 2025 · 44m · latent-space

DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever

Sherwin Wu · 19m spoken Christina Huang · 11m spoken Shawn Wang · 7m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Recorded live at OpenAI DevDay 2025, Latent Space hosts Swix and Alessio Fanelli interview OpenAI Platform leaders Sherwin Wu and Christina Huang about major developer announcements, including the Apps SDK, Agent Kit, and ChatKit. The discussion explores agentic architecture, open ecosystem protocols like MCP, internal Codex development habits, and the production realities of reliability and inference economics.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 26.5% of the talking time here. How this is scored →

The hosts as informed peer 6.0 Guest teaching 3.1 Guest disagreement 1.0 The hosts pushing back 1.9
05100:0015:0030:003:08–8:56 · The hosts as informed peer 6/10 Apps SDK Launch and Model Context Protocol Hosts demonstrate strong knowledge of past DevDay history and MCP's rise, while Sherwin and Christina elaborate on OpenAI's protocol philosophy and UI embedding.8:56–17:27 · The hosts as informed peer 6/10 Agent Kit Architecture and Multi-Model Evaluation Swyx probes into the Agent Kit architecture, prompting Christina and Sherwin to clarify visual versus code-first workflows and multi-model eval integrations.17:27–22:10 · The hosts as informed peer 7/10 Agentic Evaluations and Automated Prompt Optimization Hosts reference zero-gradient optimization concepts while Sherwin discusses long trace evaluations and prompt engineering durability.22:10–28:55 · The hosts as informed peer 6/10 Industry Perspectives on the Tinker API Swyx brings up John Schulman and Tinker API, leading to an agreeable exchange on low-level research fine-tuning abstractions.28:55–32:24 · The hosts as informed peer 6/10 Identity Management and Inference Cost Economics Hosts inquire about ChatGPT identity provider capabilities and BYOK economics, with guests explaining current Apple and Kakao auth paradigms.32:24–39:13 · The hosts as informed peer 6/10 ChatKit UI Components and Enterprise Deployment Alessio and Swyx challenge whether ChatKit will become an open-source standard or internal tool, prompting Christina to explain iframe encapsulation.39:13–42:17 · The hosts as informed peer 5/10 Codex Power User Habits and Development Practices Swyx pushes on the code review etiquette of vibe coding, leading to guests explaining autonomous Codex PR workflows internally.3:08–8:56 · Guest teaching 3/10 Apps SDK Launch and Model Context Protocol Hosts demonstrate strong knowledge of past DevDay history and MCP's rise, while Sherwin and Christina elaborate on OpenAI's protocol philosophy and UI embedding.8:56–17:27 · Guest teaching 4/10 Agent Kit Architecture and Multi-Model Evaluation Swyx probes into the Agent Kit architecture, prompting Christina and Sherwin to clarify visual versus code-first workflows and multi-model eval integrations.17:27–22:10 · Guest teaching 3/10 Agentic Evaluations and Automated Prompt Optimization Hosts reference zero-gradient optimization concepts while Sherwin discusses long trace evaluations and prompt engineering durability.22:10–28:55 · Guest teaching 2/10 Industry Perspectives on the Tinker API Swyx brings up John Schulman and Tinker API, leading to an agreeable exchange on low-level research fine-tuning abstractions.28:55–32:24 · Guest teaching 3/10 Identity Management and Inference Cost Economics Hosts inquire about ChatGPT identity provider capabilities and BYOK economics, with guests explaining current Apple and Kakao auth paradigms.32:24–39:13 · Guest teaching 4/10 ChatKit UI Components and Enterprise Deployment Alessio and Swyx challenge whether ChatKit will become an open-source standard or internal tool, prompting Christina to explain iframe encapsulation.39:13–42:17 · Guest teaching 3/10 Codex Power User Habits and Development Practices Swyx pushes on the code review etiquette of vibe coding, leading to guests explaining autonomous Codex PR workflows internally.3:08–8:56 · Guest disagreement 1/10 Apps SDK Launch and Model Context Protocol Hosts demonstrate strong knowledge of past DevDay history and MCP's rise, while Sherwin and Christina elaborate on OpenAI's protocol philosophy and UI embedding.8:56–17:27 · Guest disagreement 1/10 Agent Kit Architecture and Multi-Model Evaluation Swyx probes into the Agent Kit architecture, prompting Christina and Sherwin to clarify visual versus code-first workflows and multi-model eval integrations.17:27–22:10 · Guest disagreement 1/10 Agentic Evaluations and Automated Prompt Optimization Hosts reference zero-gradient optimization concepts while Sherwin discusses long trace evaluations and prompt engineering durability.22:10–28:55 · Guest disagreement 1/10 Industry Perspectives on the Tinker API Swyx brings up John Schulman and Tinker API, leading to an agreeable exchange on low-level research fine-tuning abstractions.28:55–32:24 · Guest disagreement 1/10 Identity Management and Inference Cost Economics Hosts inquire about ChatGPT identity provider capabilities and BYOK economics, with guests explaining current Apple and Kakao auth paradigms.32:24–39:13 · Guest disagreement 1/10 ChatKit UI Components and Enterprise Deployment Alessio and Swyx challenge whether ChatKit will become an open-source standard or internal tool, prompting Christina to explain iframe encapsulation.39:13–42:17 · Guest disagreement 1/10 Codex Power User Habits and Development Practices Swyx pushes on the code review etiquette of vibe coding, leading to guests explaining autonomous Codex PR workflows internally.3:08–8:56 · The hosts pushing back 1/10 Apps SDK Launch and Model Context Protocol Hosts demonstrate strong knowledge of past DevDay history and MCP's rise, while Sherwin and Christina elaborate on OpenAI's protocol philosophy and UI embedding.8:56–17:27 · The hosts pushing back 2/10 Agent Kit Architecture and Multi-Model Evaluation Swyx probes into the Agent Kit architecture, prompting Christina and Sherwin to clarify visual versus code-first workflows and multi-model eval integrations.17:27–22:10 · The hosts pushing back 2/10 Agentic Evaluations and Automated Prompt Optimization Hosts reference zero-gradient optimization concepts while Sherwin discusses long trace evaluations and prompt engineering durability.22:10–28:55 · The hosts pushing back 2/10 Industry Perspectives on the Tinker API Swyx brings up John Schulman and Tinker API, leading to an agreeable exchange on low-level research fine-tuning abstractions.28:55–32:24 · The hosts pushing back 2/10 Identity Management and Inference Cost Economics Hosts inquire about ChatGPT identity provider capabilities and BYOK economics, with guests explaining current Apple and Kakao auth paradigms.32:24–39:13 · The hosts pushing back 2/10 ChatKit UI Components and Enterprise Deployment Alessio and Swyx challenge whether ChatKit will become an open-source standard or internal tool, prompting Christina to explain iframe encapsulation.39:13–42:17 · The hosts pushing back 2/10 Codex Power User Habits and Development Practices Swyx pushes on the code review etiquette of vibe coding, leading to guests explaining autonomous Codex PR workflows internally.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 54.2% · guest 45.8%0:00 · the hosts 54.2% · guest 45.8%3:00 · the hosts 24% · guest 76%3:00 · the hosts 24% · guest 76%6:00 · the hosts 31% · guest 69%6:00 · the hosts 31% · guest 69%9:00 · the hosts 44.1% · guest 55.9%9:00 · the hosts 44.1% · guest 55.9%12:00 · the hosts 12.4% · guest 87.6%12:00 · the hosts 12.4% · guest 87.6%15:00 · the hosts 19.9% · guest 80.1%15:00 · the hosts 19.9% · guest 80.1%18:00 · the hosts 11.7% · guest 88.3%18:00 · the hosts 11.7% · guest 88.3%21:00 · the hosts 33.9% · guest 66.1%21:00 · the hosts 33.9% · guest 66.1%24:00 · the hosts 39.3% · guest 60.7%24:00 · the hosts 39.3% · guest 60.7%27:00 · the hosts 21% · guest 79%27:00 · the hosts 21% · guest 79%30:00 · the hosts 39% · guest 61%30:00 · the hosts 39% · guest 61%33:00 · the hosts 3.5% · guest 96.5%33:00 · the hosts 3.5% · guest 96.5%36:00 · the hosts 32.2% · guest 67.8%36:00 · the hosts 32.2% · guest 67.8%39:00 · the hosts 13.2% · guest 86.8%39:00 · the hosts 13.2% · guest 86.8%42:00 · the hosts 18.5% · guest 81.5%42:00 · the hosts 18.5% · guest 81.5%
Sharpest disagreement ▶ 40:44 Playful disagreement on PR review etiquette

Guests gently push back on Swyx's concern about vibe coding etiquette by revealing Codex now performs autonomous reviews internally.

Hardest push from the hosts ▶ 35:25 Swyx presses on open sourcing ChatKit

Swyx asks directly whether ChatKit will be open sourced, prompting a structured justification of the closed iframe architecture.

Biggest teaching moment ▶ 29:31 Sherwin explains existing identity routing mechanics

Sherwin explains to the hosts how Siri's Apple integration already functions as a tiered ChatGPT identity provider in production.

The host holds their own ▶ 21:14 Swyx provides the zero-gradient fine-tuning framing

Swyx showcases domain depth by linking prompt optimization techniques directly to Shunyu Yao's research and zero-gradient theory.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Apps SDK Launch and Model Context Protocol 6311 Hosts demonstrate strong knowledge of past DevDay history and MCP's rise, while Sherwin and Christina elaborate on OpenAI's protocol philosophy and UI embedding.
Agent Kit Architecture and Multi-Model Evaluation 6412 Swyx probes into the Agent Kit architecture, prompting Christina and Sherwin to clarify visual versus code-first workflows and multi-model eval integrations.
Agentic Evaluations and Automated Prompt Optimization 7312 Hosts reference zero-gradient optimization concepts while Sherwin discusses long trace evaluations and prompt engineering durability.
Industry Perspectives on the Tinker API 6212 Swyx brings up John Schulman and Tinker API, leading to an agreeable exchange on low-level research fine-tuning abstractions.
Identity Management and Inference Cost Economics 6312 Hosts inquire about ChatGPT identity provider capabilities and BYOK economics, with guests explaining current Apple and Kakao auth paradigms.
ChatKit UI Components and Enterprise Deployment 6412 Alessio and Swyx challenge whether ChatKit will become an open-source standard or internal tool, prompting Christina to explain iframe encapsulation.
Codex Power User Habits and Development Practices 5312 Swyx pushes on the code review etiquette of vibe coding, leading to guests explaining autonomous Codex PR workflows internally.

Statements from this episode (19)

Assertion Supported
OpenAI reports 4 million active developers at DevDay 2025
“Every year in Dev Day, you report the number of developers. This year is four million. I think last year was like three.”
Shawn Wang Oct 7, 2025 ▶ 1:24
Assertion Supported
OpenAI's Nick Cooper sits on Anthropic's MCP steering committee
“We actually have a member of our team, Nick Cooper, who is sitting on kind of like that, that steering committee for MCP as well.”
Sherwin Wu Oct 7, 2025 ▶ 6:45
Assertion Supported
OpenAI launches Agent Kit to build, deploy, and optimize agents
“We launched Agent Kit today. Full set of solutions to build, deploy, and optimize agents.”
Christina Huang Oct 7, 2025 ▶ 9:26
Disclosure
OpenAI plans two-way code sync and code execution in Agent Builder
“Eventually, like, that's definitely what we want to do. Maybe you could start off in code. You could bring it in. We'll also probably have, like, ability to, you know, run code in, in the agent builder as well”
Christina Huang Oct 7, 2025 ▶ 13:42
Assertion Supported
Wu: OpenAI was first to launch stateful responses API
“Obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted, I think Grok has it in their API. I think I saw LMSYS just did something”
Sherwin Wu Oct 7, 2025 ▶ 15:46
Assertion Supported
OpenAI adds third-party model support to its evals product
“One of the things that we launched today with evals too is ability to use, like, third-party models as well and kind of bring that into one place”
Christina Huang Oct 7, 2025 ▶ 16:38
Assertion Supported
OpenAI integrates with OpenRouter for multi-provider evals
“We have a really cool setup with Open Router, where we're working with them, and then you can bring your Open Router setup. And then with that, you can actually, you know, you write your evals using our data sets tool, or use our data set tool to create a bunc…”
Sherwin Wu Oct 7, 2025 ▶ 17:02
Insight
AI industry has completed only 10% of necessary agent evaluation progress
“I actually think agent evals is still a work in progress. So I think we've, like, made maybe 10% of the progress that we need here.”
Sherwin Wu Oct 7, 2025 ▶ 17:44
Insight
Prompt engineering has grown more entrenched despite predictions of its demise
“I feel like two years ago, people were like, oh, at some point, prop, like, prompting's gonna be dead. Like, you know, and it's like, you know... And if anything, it is, like, become more and more entrenched. And I think that, you know, there's this interestin…”
Sherwin Wu Oct 7, 2025 ▶ 20:46
Disclosure
Wu: Managing and serving fine-tuned snapshots is extremely difficult for OpenAI
“We have a fine-tuning API, and, like, it is extremely difficult for us to run, you know, and serve, like, all of these different snapshots... But like, man, it is like pretty difficult for us to like manage all of these different snapshots.”
Sherwin Wu Oct 7, 2025 ▶ 21:38
Assertion Not checkable as stated
John Schulman developed the Tinker API concept across OpenAI and Anthropic
“Right when I joined OpenAI, like, this has actually been, I think, a passion project of John's. Like, he's been talking about doing something in this, like, in this shape for a while, which is, like, a truly, like, low-level research, like, fine-tuning library…”
Sherwin Wu Oct 7, 2025 ▶ 22:35
Disclosure
Wu: OpenAI has no current plans to become a generic IdP
“Direct answer is like no plans right now, of course but I actually think we currently have some version of this, which is our partnership with Apple because with Apple, you can actually sign in to your ChatGPT account, and some of that identity does carry with…”
Sherwin Wu Oct 7, 2025 ▶ 29:32
Assertion Supported
Apple Siri routes requests based on the user's ChatGPT subscription tier
“If you sign into your ChatGPT account the Siri integration will actually use your subscription status to decide what type of model to use when it passes things over to ChatGPT. And so if you're you know just a free user you get, you know, the free model. But i…”
Sherwin Wu Oct 7, 2025 ▶ 29:56
Insight
Wu: Cheaper inference does not cut developer spending due to surging demand
“What we realized is as we make it cheaper, you know, the demand for that goes up even more, and you end up, you know, still spending quite a bit”
Sherwin Wu Oct 7, 2025 ▶ 32:11
Assertion Not publicly verifiable
Huang: OpenAI Customer Support Is Powered by AgentKit
“We use this internally and externally, like our customer support, help.open.au.com already powered on agent kit and then various other like internal use cases as well.”
Christina Huang Oct 7, 2025 ▶ 33:52
Insight
Huang: Chat Interfaces Should Be Standardized Drop-in Components Like Stripe Checkout
“Yeah, so it's very similar philosophically, right? So Stripe, you know, can build elements and check out, and not every business needs to rebuild, right, the pieces that are really common, and I think we see the same with chat. We see chat being built over and…”
Christina Huang Oct 7, 2025 ▶ 36:30
Assertion Not checkable as stated
OpenAI Codex successfully one-shots entire features 30% to 40% of the time
“What a lot of the interns would do is just, like, full YOLO mode, like, trust it to, like, write the whole feature. And it, like, it doesn't work. It, like, doesn't work sometimes. But, like, I don't know, like, 30, 40% of the time it just, like, one-shots it.”
Sherwin Wu Oct 7, 2025 ▶ 40:00
What-if
Huang: Building Visual Agent Builder in under 2 months required Codex
“For the Visual Agents Builder, we only started that Probably less than two months ago, and that, that wouldn't be possible without Codex.”
Christina Huang Oct 7, 2025 ▶ 41:12
Assertion Supported
OpenAI's API throughput has surpassed six billion tokens per minute
“We actually zoomed past that.”
Sherwin Wu Oct 7, 2025 ▶ 44:33
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.