Sherwin Wu

Head of Engineering, OpenAI Platform · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

engineeroperatorexecutive@sherwinwu ↗LinkedIn ↗

Sherwin Wu is the Head of Engineering for OpenAI's developer platform and API products. He previously worked as a software engineer at Quora, Opendoor, and Palantir.

12statements → 7claims → 5claims resolved → 100%fully supported → 3.75/5average certainty → 1.42/5average debate potential →

5 supported 0 partly supported 0 contradicted 2 not checkable as stated how the 7 claims stand · each chip opens the sources

7 assertions · 3 insights · 2 disclosures · every statement was checked. The predictions and assertions are the 7 claims: statements the public record can support or contradict. 5 are resolved, and 2 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Sherwin argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Wu: OpenAI was first to launch stateful responses API
“Obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted, I think Grok has it in their API. I think I saw LMSYS just did something”
Sherwin Wu Oct 7, 2025 ▶ 15:46 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Sherwin Wu on measured tape to publish a rate. This says nothing about how they speak.

Everything Sherwin Wu said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
Wu: OpenAI was first to launch stateful responses API
“Obviously we were the first one to launch responses API, but like a couple of other people have kind of adopted, I think Grok has it in their API. I think I saw LMSYS just did something”
Sherwin Wu Oct 7, 2025 ▶ 15:46 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI integrates with OpenRouter for multi-provider evals
“We have a really cool setup with Open Router, where we're working with them, and then you can bring your Open Router setup. And then with that, you can actually, you know, you write your evals using our data sets tool, or use our data set tool to create a bunc…”
Sherwin Wu Oct 7, 2025 ▶ 17:02 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Insight
AI industry has completed only 10% of necessary agent evaluation progress
“I actually think agent evals is still a work in progress. So I think we've, like, made maybe 10% of the progress that we need here.”
Sherwin Wu Oct 7, 2025 ▶ 17:44 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Insight
Prompt engineering has grown more entrenched despite predictions of its demise
“I feel like two years ago, people were like, oh, at some point, prop, like, prompting's gonna be dead. Like, you know, and it's like, you know... And if anything, it is, like, become more and more entrenched. And I think that, you know, there's this interestin…”
Sherwin Wu Oct 7, 2025 ▶ 20:46 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
Apple Siri routes requests based on the user's ChatGPT subscription tier
“If you sign into your ChatGPT account the Siri integration will actually use your subscription status to decide what type of model to use when it passes things over to ChatGPT. And so if you're you know just a free user you get, you know, the free model. But i…”
Sherwin Wu Oct 7, 2025 ▶ 29:56 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Insight
Wu: Cheaper inference does not cut developer spending due to surging demand
“What we realized is as we make it cheaper, you know, the demand for that goes up even more, and you end up, you know, still spending quite a bit”
Sherwin Wu Oct 7, 2025 ▶ 32:11 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Not checkable as stated
OpenAI Codex successfully one-shots entire features 30% to 40% of the time
“What a lot of the interns would do is just, like, full YOLO mode, like, trust it to, like, write the whole feature. And it, like, it doesn't work. It, like, doesn't work sometimes. But, like, I don't know, like, 30, 40% of the time it just, like, one-shots it.”
Sherwin Wu Oct 7, 2025 ▶ 40:00 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI's API throughput has surpassed six billion tokens per minute
“We actually zoomed past that.”
Sherwin Wu Oct 7, 2025 ▶ 44:33 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Supported
OpenAI's Nick Cooper sits on Anthropic's MCP steering committee
“We actually have a member of our team, Nick Cooper, who is sitting on kind of like that, that steering committee for MCP as well.”
Sherwin Wu Oct 7, 2025 ▶ 6:45 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Disclosure
Wu: Managing and serving fine-tuned snapshots is extremely difficult for OpenAI
“We have a fine-tuning API, and, like, it is extremely difficult for us to run, you know, and serve, like, all of these different snapshots... But like, man, it is like pretty difficult for us to like manage all of these different snapshots.”
Sherwin Wu Oct 7, 2025 ▶ 21:38 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Assertion Not checkable as stated
John Schulman developed the Tinker API concept across OpenAI and Anthropic
“Right when I joined OpenAI, like, this has actually been, I think, a passion project of John's. Like, he's been talking about doing something in this, like, in this shape for a while, which is, like, a truly, like, low-level research, like, fine-tuning library…”
Sherwin Wu Oct 7, 2025 ▶ 22:35 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever
Disclosure
Wu: OpenAI has no current plans to become a generic IdP
“Direct answer is like no plans right now, of course but I actually think we currently have some version of this, which is our partnership with Apple because with Apple, you can actually sign in to your ChatGPT account, and some of that identity does carry with…”
Sherwin Wu Oct 7, 2025 ▶ 29:32 DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever

Appearances (1)

EpisodeDateSpeaking time
DevDay 2025: Apps SDK, Agent Kit, MCP, Codex and why Prompting is More Important than Ever Oct 7, 2025 19m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.