Opinion certainty 4/5 debate potential 3/5

Sherman Wu: Combining language and diffusion models is an anti-pattern

Sherman Wu · How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning · Nov 28, 2025 · at 42:59

Sherwin Wu is Head of Engineering for OpenAI's Developer Platform. He discusses the organizational and technical challenges of developing both LLMs and diffusion models like Sora.

0:00 / 0:04exact quote · 4.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Yeah, I think you're totally right. It's an anti-pattern. It's pretty tough to pull off.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Sherman Wu

Prediction Not checkable as stated
Sherman Wu: AI industry will shift toward specialized models
“It's, like, becoming increasingly clear that there will be room for a bunch of specialized models. There will likely be a proliferation of other types of models.”
Sherman Wu Nov 28, 2025 ▶ 0:16 How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning
Assertion Supported
Sherman Wu: ChatGPT has reached 800 million weekly active users
“The first party app is a really great way to get, you know it was like eight hundred million wows or whatever now.”
Sherman Wu Nov 28, 2025 ▶ 10:04 How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning
Disclosure
Sherman Wu: OpenAI previously believed a single model would replace fine-tuning
“I remember like, even with an open AI, the thinking was that there would be like one model that rules them all. And it's like, why would you, I mean, like this kind of goes to the fine tuning API product. It's like, why would you even have a fine tuning produc…”
Sherman Wu Nov 28, 2025 ▶ 17:58 How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning
Assertion Not checkable as stated
Sherman Wu: OpenAI prices its API services using a cost-plus model
“Internally, one thing we do is, is we always make sure that we actually price our usage-based pricing from a, like, cost-plus perspective. Like, we're actually just, like, trying to make sure that we're being responsible from a margin perspective.”
Sherman Wu Nov 28, 2025 ▶ 32:45 How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning
Insight
Sherman Wu: Usage-based pricing naturally approximates outcome-based pricing in AI
“On outcome-based pricing it sounds very appealing, like, if it can work, but one thing that we've started realizing is it actually ends up correlating quite a bit with usage-based pricing, especially with test time compute. Like, if the thing is just, like, th…”
Sherman Wu Nov 28, 2025 ▶ 35:53 How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning
What-if
Sherman Wu: Replicating OpenAI's flagship model inference performance is extremely hard
“Even if we just like, you know, open source, like if we just literally open sourced GPT-V or something, it would be really, really hard to inference it at the level that we are able to get it to do.”
Sherman Wu Nov 28, 2025 ▶ 39:53 How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.