Nov 28, 2025 · 53m · a16z

How OpenAI Builds for 800 Million Weekly Users: Model Specialization and Fine-Tuning

Sherman Wu · 30m spoken Martin Casado · 18m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The a16z Show, host Martin Casado interviews Sherman Wu, Head of Engineering for OpenAI's Developer Platform, about model specialization, reinforcement fine-tuning (RFT), AI agent architecture, and platform economics.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 5.9 Guest teaching 3.4 Guest disagreement 0.9 The host pushing back 2.4
05100:0015:0030:0045:000:34–4:59 · The host as informed peer 3/10 The a16z Show Title Card Martin asks conversational questions about Sherman's career transitions from OpenDoor ML pricing to OpenAI's developer platform, drawing on his own background at Livermore.4:59–8:35 · The host as informed peer 2/10 Early Career at Quora and MIT Origins Sherman shares stories about his MIT externship and Quora's early engineering talent, while Martin reminisces about Quora's legendary founding team.8:35–12:06 · The host as informed peer 5/10 Internal Strategy: Horizontal APIs vs. Consumer Products Martin asks whether offering both horizontal APIs and vertical consumer apps creates internal product conflict, prompting Sherman to explain OpenAI's 800M user scale and mission framing.12:06–15:10 · The host as informed peer 7/10 Model Retention and Anti-Disintermediation Mechanics Martin proposes that AI models act as anti-disintermediation technology because users form sticky behavioral habits that resist traditional software abstraction.15:10–19:15 · The host as informed peer 6/10 Developer Workflows in a Multi-Model Landscape Martin cites his hands-on developer experience using Cursor across multiple models to challenge the early consensus that one single AGI model would rule everything.19:15–24:50 · The host as informed peer 6/10 Enterprise Customization: Reinforcement Fine-Tuning (RFT) Martin probes whether fine-tuning is capitulation on general intelligence, leading Sherman to explain how reinforcement fine-tuning (RFT) unlocks state-of-the-art domain capabilities.24:50–27:00 · The host as informed peer 8/10 Context Engineering and the Evolution of RAG Martin sharply critiques primitive RAG architectures as insulting to superintelligence, which Sherman validates while explaining how reasoning models transform context retrieval.27:00–30:51 · The host as informed peer 5/10 Architectural Perspective on AI Agents and Interfaces Martin attempts to categorize agents into classic product versus API categories, but Sherman reframes agents as varied user interfaces for underlying model intelligence.30:51–36:46 · The host as informed peer 7/10 AI Platform Economics and Usage-Based Pricing Martin analyzes the industry transition to usage-based billing and challenges the feasibility of outcome-based pricing in complex domains, which Sherman agrees with via test-time compute trends.36:53–40:15 · The host as informed peer 6/10 Open Source Strategy and Community Ecosystem Impact Martin highlights that open-weight models do not cannibalize API revenue because running performant inference infrastructure remains the true operational moat.40:15–42:28 · The host as informed peer 6/10 Exploring Deep Model Verticalization Across Modalities Martin compares text models to image LoRAs, prompting Sherman to explain the heavy pre-training and post-training bottlenecks that limit deep text verticalization.42:28–45:36 · The host as informed peer 7/10 Managing Multimodal Models and World Simulation Infrastructure Martin identifies developing both text and video diffusion models under one roof as an organizational anti-pattern, which Sherman confirms while detailing OpenAI's isolated World Simulation team.45:36–49:06 · The host as informed peer 6/10 The Evolution of AI Agents and SOP Workflows Martin brings up developer skepticism regarding low-code agent builders, and Sherman explains why deterministic nodes are essential for procedural standard operating procedures.49:06–53:07 · The host as informed peer 8/10 Constraining Agent Execution for Games and Regulated Fields Martin details technical implementation patterns like passing Python pseudocode in prompts to constrain NPC behavior in games and regulated fields, which Sherman enthusiastically praises.0:34–4:59 · Guest teaching 2/10 The a16z Show Title Card Martin asks conversational questions about Sherman's career transitions from OpenDoor ML pricing to OpenAI's developer platform, drawing on his own background at Livermore.4:59–8:35 · Guest teaching 1/10 Early Career at Quora and MIT Origins Sherman shares stories about his MIT externship and Quora's early engineering talent, while Martin reminisces about Quora's legendary founding team.8:35–12:06 · Guest teaching 3/10 Internal Strategy: Horizontal APIs vs. Consumer Products Martin asks whether offering both horizontal APIs and vertical consumer apps creates internal product conflict, prompting Sherman to explain OpenAI's 800M user scale and mission framing.12:06–15:10 · Guest teaching 2/10 Model Retention and Anti-Disintermediation Mechanics Martin proposes that AI models act as anti-disintermediation technology because users form sticky behavioral habits that resist traditional software abstraction.15:10–19:15 · Guest teaching 3/10 Developer Workflows in a Multi-Model Landscape Martin cites his hands-on developer experience using Cursor across multiple models to challenge the early consensus that one single AGI model would rule everything.19:15–24:50 · Guest teaching 5/10 Enterprise Customization: Reinforcement Fine-Tuning (RFT) Martin probes whether fine-tuning is capitulation on general intelligence, leading Sherman to explain how reinforcement fine-tuning (RFT) unlocks state-of-the-art domain capabilities.24:50–27:00 · Guest teaching 4/10 Context Engineering and the Evolution of RAG Martin sharply critiques primitive RAG architectures as insulting to superintelligence, which Sherman validates while explaining how reasoning models transform context retrieval.27:00–30:51 · Guest teaching 4/10 Architectural Perspective on AI Agents and Interfaces Martin attempts to categorize agents into classic product versus API categories, but Sherman reframes agents as varied user interfaces for underlying model intelligence.30:51–36:46 · Guest teaching 4/10 AI Platform Economics and Usage-Based Pricing Martin analyzes the industry transition to usage-based billing and challenges the feasibility of outcome-based pricing in complex domains, which Sherman agrees with via test-time compute trends.36:53–40:15 · Guest teaching 3/10 Open Source Strategy and Community Ecosystem Impact Martin highlights that open-weight models do not cannibalize API revenue because running performant inference infrastructure remains the true operational moat.40:15–42:28 · Guest teaching 5/10 Exploring Deep Model Verticalization Across Modalities Martin compares text models to image LoRAs, prompting Sherman to explain the heavy pre-training and post-training bottlenecks that limit deep text verticalization.42:28–45:36 · Guest teaching 4/10 Managing Multimodal Models and World Simulation Infrastructure Martin identifies developing both text and video diffusion models under one roof as an organizational anti-pattern, which Sherman confirms while detailing OpenAI's isolated World Simulation team.45:36–49:06 · Guest teaching 5/10 The Evolution of AI Agents and SOP Workflows Martin brings up developer skepticism regarding low-code agent builders, and Sherman explains why deterministic nodes are essential for procedural standard operating procedures.49:06–53:07 · Guest teaching 2/10 Constraining Agent Execution for Games and Regulated Fields Martin details technical implementation patterns like passing Python pseudocode in prompts to constrain NPC behavior in games and regulated fields, which Sherman enthusiastically praises.0:34–4:59 · Guest disagreement 1/10 The a16z Show Title Card Martin asks conversational questions about Sherman's career transitions from OpenDoor ML pricing to OpenAI's developer platform, drawing on his own background at Livermore.4:59–8:35 · Guest disagreement 0/10 Early Career at Quora and MIT Origins Sherman shares stories about his MIT externship and Quora's early engineering talent, while Martin reminisces about Quora's legendary founding team.8:35–12:06 · Guest disagreement 1/10 Internal Strategy: Horizontal APIs vs. Consumer Products Martin asks whether offering both horizontal APIs and vertical consumer apps creates internal product conflict, prompting Sherman to explain OpenAI's 800M user scale and mission framing.12:06–15:10 · Guest disagreement 1/10 Model Retention and Anti-Disintermediation Mechanics Martin proposes that AI models act as anti-disintermediation technology because users form sticky behavioral habits that resist traditional software abstraction.15:10–19:15 · Guest disagreement 1/10 Developer Workflows in a Multi-Model Landscape Martin cites his hands-on developer experience using Cursor across multiple models to challenge the early consensus that one single AGI model would rule everything.19:15–24:50 · Guest disagreement 1/10 Enterprise Customization: Reinforcement Fine-Tuning (RFT) Martin probes whether fine-tuning is capitulation on general intelligence, leading Sherman to explain how reinforcement fine-tuning (RFT) unlocks state-of-the-art domain capabilities.24:50–27:00 · Guest disagreement 1/10 Context Engineering and the Evolution of RAG Martin sharply critiques primitive RAG architectures as insulting to superintelligence, which Sherman validates while explaining how reasoning models transform context retrieval.27:00–30:51 · Guest disagreement 2/10 Architectural Perspective on AI Agents and Interfaces Martin attempts to categorize agents into classic product versus API categories, but Sherman reframes agents as varied user interfaces for underlying model intelligence.30:51–36:46 · Guest disagreement 1/10 AI Platform Economics and Usage-Based Pricing Martin analyzes the industry transition to usage-based billing and challenges the feasibility of outcome-based pricing in complex domains, which Sherman agrees with via test-time compute trends.36:53–40:15 · Guest disagreement 1/10 Open Source Strategy and Community Ecosystem Impact Martin highlights that open-weight models do not cannibalize API revenue because running performant inference infrastructure remains the true operational moat.40:15–42:28 · Guest disagreement 1/10 Exploring Deep Model Verticalization Across Modalities Martin compares text models to image LoRAs, prompting Sherman to explain the heavy pre-training and post-training bottlenecks that limit deep text verticalization.42:28–45:36 · Guest disagreement 1/10 Managing Multimodal Models and World Simulation Infrastructure Martin identifies developing both text and video diffusion models under one roof as an organizational anti-pattern, which Sherman confirms while detailing OpenAI's isolated World Simulation team.45:36–49:06 · Guest disagreement 1/10 The Evolution of AI Agents and SOP Workflows Martin brings up developer skepticism regarding low-code agent builders, and Sherman explains why deterministic nodes are essential for procedural standard operating procedures.49:06–53:07 · Guest disagreement 0/10 Constraining Agent Execution for Games and Regulated Fields Martin details technical implementation patterns like passing Python pseudocode in prompts to constrain NPC behavior in games and regulated fields, which Sherman enthusiastically praises.0:34–4:59 · The host pushing back 1/10 The a16z Show Title Card Martin asks conversational questions about Sherman's career transitions from OpenDoor ML pricing to OpenAI's developer platform, drawing on his own background at Livermore.4:59–8:35 · The host pushing back 0/10 Early Career at Quora and MIT Origins Sherman shares stories about his MIT externship and Quora's early engineering talent, while Martin reminisces about Quora's legendary founding team.8:35–12:06 · The host pushing back 2/10 Internal Strategy: Horizontal APIs vs. Consumer Products Martin asks whether offering both horizontal APIs and vertical consumer apps creates internal product conflict, prompting Sherman to explain OpenAI's 800M user scale and mission framing.12:06–15:10 · The host pushing back 3/10 Model Retention and Anti-Disintermediation Mechanics Martin proposes that AI models act as anti-disintermediation technology because users form sticky behavioral habits that resist traditional software abstraction.15:10–19:15 · The host pushing back 3/10 Developer Workflows in a Multi-Model Landscape Martin cites his hands-on developer experience using Cursor across multiple models to challenge the early consensus that one single AGI model would rule everything.19:15–24:50 · The host pushing back 3/10 Enterprise Customization: Reinforcement Fine-Tuning (RFT) Martin probes whether fine-tuning is capitulation on general intelligence, leading Sherman to explain how reinforcement fine-tuning (RFT) unlocks state-of-the-art domain capabilities.24:50–27:00 · The host pushing back 4/10 Context Engineering and the Evolution of RAG Martin sharply critiques primitive RAG architectures as insulting to superintelligence, which Sherman validates while explaining how reasoning models transform context retrieval.27:00–30:51 · The host pushing back 3/10 Architectural Perspective on AI Agents and Interfaces Martin attempts to categorize agents into classic product versus API categories, but Sherman reframes agents as varied user interfaces for underlying model intelligence.30:51–36:46 · The host pushing back 3/10 AI Platform Economics and Usage-Based Pricing Martin analyzes the industry transition to usage-based billing and challenges the feasibility of outcome-based pricing in complex domains, which Sherman agrees with via test-time compute trends.36:53–40:15 · The host pushing back 2/10 Open Source Strategy and Community Ecosystem Impact Martin highlights that open-weight models do not cannibalize API revenue because running performant inference infrastructure remains the true operational moat.40:15–42:28 · The host pushing back 2/10 Exploring Deep Model Verticalization Across Modalities Martin compares text models to image LoRAs, prompting Sherman to explain the heavy pre-training and post-training bottlenecks that limit deep text verticalization.42:28–45:36 · The host pushing back 3/10 Managing Multimodal Models and World Simulation Infrastructure Martin identifies developing both text and video diffusion models under one roof as an organizational anti-pattern, which Sherman confirms while detailing OpenAI's isolated World Simulation team.45:36–49:06 · The host pushing back 3/10 The Evolution of AI Agents and SOP Workflows Martin brings up developer skepticism regarding low-code agent builders, and Sherman explains why deterministic nodes are essential for procedural standard operating procedures.49:06–53:07 · The host pushing back 2/10 Constraining Agent Execution for Games and Regulated Fields Martin details technical implementation patterns like passing Python pseudocode in prompts to constrain NPC behavior in games and regulated fields, which Sherman enthusiastically praises.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 27:43 Reframing agents as interfaces rather than separate product lines

Sherman rejects Martin's attempt to isolate agents into traditional API or product categories, reframing them as varied interfaces for underlying model intelligence.

Hardest push from the host ▶ 42:30 Calling dual language and diffusion development an anti-pattern

Martin explicitly pushes back on combining text and video diffusion models in a single company, labeling it an organizational anti-pattern.

Biggest teaching moment ▶ 41:37 Explaining text vs pixel model training bottlenecks

Sherman educates Martin on why text models cannot be verticalized as easily as image diffusion models due to heavy pre-training and post-training compute steps.

The host holds their own ▶ 50:38 Detailing pseudocode prompting patterns for NPCs

Martin demonstrates expert engineering knowledge by describing how developers pass Python pseudocode in prompts to constrain LLM behavior in game NPCs and regulated fields.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
The a16z Show Title Card 3211 Martin asks conversational questions about Sherman's career transitions from OpenDoor ML pricing to OpenAI's developer platform, drawing on his own background at Livermore.
Early Career at Quora and MIT Origins 2100 Sherman shares stories about his MIT externship and Quora's early engineering talent, while Martin reminisces about Quora's legendary founding team.
Internal Strategy: Horizontal APIs vs. Consumer Products 5312 Martin asks whether offering both horizontal APIs and vertical consumer apps creates internal product conflict, prompting Sherman to explain OpenAI's 800M user scale and mission framing.
Model Retention and Anti-Disintermediation Mechanics 7213 Martin proposes that AI models act as anti-disintermediation technology because users form sticky behavioral habits that resist traditional software abstraction.
Developer Workflows in a Multi-Model Landscape 6313 Martin cites his hands-on developer experience using Cursor across multiple models to challenge the early consensus that one single AGI model would rule everything.
Enterprise Customization: Reinforcement Fine-Tuning (RFT) 6513 Martin probes whether fine-tuning is capitulation on general intelligence, leading Sherman to explain how reinforcement fine-tuning (RFT) unlocks state-of-the-art domain capabilities.
Context Engineering and the Evolution of RAG 8414 Martin sharply critiques primitive RAG architectures as insulting to superintelligence, which Sherman validates while explaining how reasoning models transform context retrieval.
Architectural Perspective on AI Agents and Interfaces 5423 Martin attempts to categorize agents into classic product versus API categories, but Sherman reframes agents as varied user interfaces for underlying model intelligence.
AI Platform Economics and Usage-Based Pricing 7413 Martin analyzes the industry transition to usage-based billing and challenges the feasibility of outcome-based pricing in complex domains, which Sherman agrees with via test-time compute trends.
Open Source Strategy and Community Ecosystem Impact 6312 Martin highlights that open-weight models do not cannibalize API revenue because running performant inference infrastructure remains the true operational moat.
Exploring Deep Model Verticalization Across Modalities 6512 Martin compares text models to image LoRAs, prompting Sherman to explain the heavy pre-training and post-training bottlenecks that limit deep text verticalization.
Managing Multimodal Models and World Simulation Infrastructure 7413 Martin identifies developing both text and video diffusion models under one roof as an organizational anti-pattern, which Sherman confirms while detailing OpenAI's isolated World Simulation team.
The Evolution of AI Agents and SOP Workflows 6513 Martin brings up developer skepticism regarding low-code agent builders, and Sherman explains why deterministic nodes are essential for procedural standard operating procedures.
Constraining Agent Execution for Games and Regulated Fields 8202 Martin details technical implementation patterns like passing Python pseudocode in prompts to constrain NPC behavior in games and regulated fields, which Sherman enthusiastically praises.

Statements from this episode (27)

Assertion Supported
Sherman Wu: 10% of the global population uses ChatGPT weekly
“10% of the globe uses it weekly.”
Sherman Wu Nov 28, 2025 ▶ 0:08
Prediction Not checkable as stated
Sherman Wu: AI industry will shift toward specialized models
“It's, like, becoming increasingly clear that there will be room for a bunch of specialized models. There will likely be a proliferation of other types of models.”
Sherman Wu Nov 28, 2025 ▶ 0:16
Disclosure
Sherman Wu: OpenAI deployed a local model at Los Alamos National Lab
“We actually do have a local deployment. At Los Alamos National Labs. It's super cool. I went to visit it. It's like very different than what I'm used to. But yeah, in a like, you know, classified supercomputer with our model running there.”
Sherman Wu Nov 28, 2025 ▶ 1:27
Opinion
Sherman Wu: OpenAI's API pricing is less sophisticated than Opendoor's
“I don't think we do anything as sophisticated as Open Door.”
Sherman Wu Nov 28, 2025 ▶ 3:25
Assertion Partly supported
Sherman Wu: Early Quora team included future Scale AI and Perplexity founders
“A bunch of the perplexity team was there. Dennis, Dennis was on the feed team with me. Johnny Ho Jerry Ma. And then Alexander, the scale, you know, like was there. He was there between high school and college.”
Sherman Wu Nov 28, 2025 ▶ 5:38
Assertion Supported
Sherman Wu: ChatGPT has reached 800 million weekly active users
“The first party app is a really great way to get, you know it was like eight hundred million wows or whatever now.”
Sherman Wu Nov 28, 2025 ▶ 10:04
Assertion Not checkable as stated
Sherman Wu: OpenAI API end-user reach previously exceeded ChatGPT's reach
“One thing we talk about internally sometimes is like, what does our end user reach from the API? Like it's actually, it's like really, really, it's really bright. It might even, it's hard because ChatGPT is growing so quickly, but like it, like at some point i…”
Sherman Wu Nov 28, 2025 ▶ 10:50
Insight
Martin Casado: LLMs naturally resist software disintermediation layers
“Because the models are so hard to abstract away. Like, they're just unruly, right? If you try to, like, have traditional software drive them, they just don't kind of manage very well. So part of me thinks that it's almost like this, like, anti-disintermediatio…”
Martin Casado Nov 28, 2025 ▶ 12:41
Assertion Not checkable as stated
Sherman Wu: Developer retention on OpenAI's API remains surprisingly high
“For or like the retention of people building on our API is like surprisingly high especially when people thought you could just kind of swap things around. You might have, you know, like even tools that help you swap things around. But yeah, the stickiness of …”
Sherman Wu Nov 28, 2025 ▶ 14:56
Disclosure
Sherman Wu: OpenAI previously believed a single model would replace fine-tuning
“I remember like, even with an open AI, the thinking was that there would be like one model that rules them all. And it's like, why would you, I mean, like this kind of goes to the fine tuning API product. It's like, why would you even have a fine tuning produc…”
Sherman Wu Nov 28, 2025 ▶ 17:58
Assertion Partly supported
Sherman Wu: OpenAI RFT elevates domain models to state-of-the-art performance
“The reinforcement fine tuning API... Kind of changes the paradigm away from just, like, small incremental, like, tone improvements, which is what SFT did, to actually improving the model to potentially SOTA level on a particular use case that you know about.”
Sherman Wu Nov 28, 2025 ▶ 23:26
Disclosure
Sherman Wu: OpenAI pilots free training for shared enterprise data
“If you actually build with a reinforcement fine tuning API, you can actually get discounted inference and potentially free training too, if you're willing to share the data”
Sherman Wu Nov 28, 2025 ▶ 24:18
Insight
Sherman Wu: AI development has shifted from prompt to context engineering
“But I think the name of the game now is, is less on like, Prompt engineering as we had thought about it two years ago. It's more of like, it's like the context engineering side where it's like, what are the tools you give it? What is like the data that it pull…”
Sherman Wu Nov 28, 2025 ▶ 25:40
Disclosure
Sherman Wu: OpenAI's o3 model stands out for diligent tool execution
“One of my favorite models is actually O three. Cause it was like one of the most diligent models. It would just like do all these tool calls and it's like really the intelligence itself trying to like do the, you know, tool calls or reg or anything like that o…”
Sherman Wu Nov 28, 2025 ▶ 26:38
Assertion Not checkable as stated
Sherman Wu: OpenAI views ChatGPT and Sora as interfaces for core intelligence
“At the end, like, at the end of the day, OpenAI is like a, an AGI company. It's like an intelligence company. And so, agents are just, like, one way in which this intelligence kind of be manifested. And so, the way that I'd say we actually think about internal…”
Sherman Wu Nov 28, 2025 ▶ 29:40
Assertion Not checkable as stated
Sherman Wu: OpenAI prices its API services using a cost-plus model
“Internally, one thing we do is, is we always make sure that we actually price our usage-based pricing from a, like, cost-plus perspective. Like, we're actually just, like, trying to make sure that we're being responsible from a margin perspective.”
Sherman Wu Nov 28, 2025 ▶ 32:45
Insight
Sherman Wu: Usage-based pricing naturally approximates outcome-based pricing in AI
“On outcome-based pricing it sounds very appealing, like, if it can work, but one thing that we've started realizing is it actually ends up correlating quite a bit with usage-based pricing, especially with test time compute. Like, if the thing is just, like, th…”
Sherman Wu Nov 28, 2025 ▶ 35:53
Disclosure
Sherman Wu: OpenAI has seen zero API cannibalization from open models
“Yeah, I mean, to be clear, like, we have not seen cannibalization at all.”
Sherman Wu Nov 28, 2025 ▶ 39:11
Insight
Sherman Wu: Major AI labs generate revenue from two or three models
“Especially for all these major labs, like they're usually like two or three models where like, that is where you're making all of your impact, all of your revenue.”
Sherman Wu Nov 28, 2025 ▶ 39:39
What-if
Sherman Wu: Replicating OpenAI's flagship model inference performance is extremely hard
“Even if we just like, you know, open source, like if we just literally open sourced GPT-V or something, it would be really, really hard to inference it at the level that we are able to get it to do.”
Sherman Wu Nov 28, 2025 ▶ 39:53
Insight
Sherman Wu: Smaller size of image models drives faster iteration and proliferation
“Image models tend to be way smaller and like you can iterate on it a lot faster. Like that's why you get that crazy cool proliferation of like the image model side.”
Sherman Wu Nov 28, 2025 ▶ 41:43
Insight
Sherman Wu: Heavy compute for text model post-training bottlenecks verticalization
“For the text models, there's always going to be this like really big fat free training step that like you have to invest in here. And then even the post training side is like, You know, it's not the, it's not like the easiest thing. Like it's, you know we all,…”
Sherman Wu Nov 28, 2025 ▶ 41:51
Opinion
Sherman Wu: Combining language and diffusion models is an anti-pattern
“Yeah, I think you're totally right. It's an anti-pattern. It's pretty tough to pull off.”
Sherman Wu Nov 28, 2025 ▶ 42:59
Opinion
Sherman Wu: OpenAI's Sora team has the highest talent concentration he's seen
“For my perspective, I think the biggest thing is I think our like image like our I think called like the world simulation team, like the team that builds Sora and all that under Aditya is just extremely solid. Like they are probably, it's like the highest Conc…”
Sherman Wu Nov 28, 2025 ▶ 43:11
Assertion Not checkable as stated
Sherman Wu: Current AI models cannot reliably execute exact multi-step instructions
“The models today, just like maybe in some future world instruction following would be so good that you just like ask it to do this four step process. And it like always does the four step process exactly. We're still not there yet.”
Sherman Wu Nov 28, 2025 ▶ 46:37
Insight
Sherman Wu: Enterprise AI requires deterministic node workflows for procedural tasks
“There's A huge need on that side to have determinism here. Of which an agent builder with nodes that kind of like helps enforce this thing ends up being very, very helpful. But I think a lot of us, especially in Silicon Valley, don't really appreciate that. Th…”
Sherman Wu Nov 28, 2025 ▶ 48:52
Insight
Martin Casado: Describing complex game logic in natural English prompts fails
“Describing the game logic in English just doesn't work, actually, if you try and do it. And then, like, actually scripting the output doesn't work either if you needed to use it in a game context.”
Martin Casado Nov 28, 2025 ▶ 50:48
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.