Sep 11, 2025 · 1h 8m · bg2-pod

Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod

Sherwin Wu · 31m spoken Olivier Godman · 20m spoken Apoorv Agrawal · 10m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the BG2 Pod, host Apoorv Agrawal interviews OpenAI Platform leaders Sherwin Wu and Olivier Godement to explore OpenAI's enterprise strategy, technical breakthroughs in GPT-5, and real-world deployment case studies across major institutions.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

Brad and Bill as informed peer 4.8 Guest teaching 4.3 Guest disagreement 1.0 Brad and Bill pushing back 1.0
05100:0015:0030:0045:001:00:002:48–6:04 · Brad and Bill as informed peer 4/10 OpenAI's Enterprise Strategy and Core Mission Apoorv opens by framing OpenAI's enterprise efforts against its public-facing ChatGPT persona. Sherwin and Olivier explain how B2B and API distribution directly fulfill OpenAI's core mission of spreading AGI benefits broadly. The exchange is highly collaborative and informative.6:04–11:30 · Brad and Bill as informed peer 6/10 Case Study: T-Mobile and Forward Deployed Engineering Olivier details their deployment at T-Mobile combining voice support and custom evaluation frameworks. Apoorv draws directly on his four-year background in forward deployed engineering at Palantir to probe into integration layers above raw models.11:30–16:46 · Brad and Bill as informed peer 5/10 Case Study: Amgen and Healthcare Innovation Olivier and Sherwin detail deployments at Amgen and Los Alamos National Laboratory. Sherwin shares technical specifics about running an air-gapped o3 model physically on the Venado supercomputer, which Apoorv receives with genuine fascination.16:46–20:17 · Brad and Bill as informed peer 6/10 Keys to Successful Enterprise AI Deployments Apoorv brings up the MIT report asserting that 95% of AI enterprise deployments fail to challenge the guests on failure modes. Olivier counters with pattern matching from hundreds of enterprise accounts, highlighting top-down buy-in, tiger teams, and rigorous bottom-up evaluation sets.20:17–25:38 · Brad and Bill as informed peer 6/10 Physical versus Digital Autonomy and Environment Scaffolding Apoorv presents a paradox comparing physical autonomy in self-driving cars with lagging digital autonomy in web agents. Sherwin and Olivier push back on the timeline and economic metrics, pointing out that physical driving benefits from extensive standardized infrastructure like roads and traffic lights that digital environments currently lack.25:38–29:21 · Brad and Bill as informed peer 5/10 Unpacking GPT-5: Architecture, Thinking, and Latency Trade-offs Apoorv inquires about the development philosophy and benchmark saturation of GPT-5. Olivier and Sherwin explain the deliberate trade-offs between deep reasoning tokens and inference latency for product builders.29:21–33:00 · Brad and Bill as informed peer 4/10 GPT-5 Real-World Performance and Prompting Dynamics Sherwin describes real-world customer reactions to GPT-5, highlighting reduced hallucinations alongside over-literal instruction following. He illustrates how legacy prompt engineering tricks caused the model to produce overly terse outputs until prompts were cleaned up.33:00–38:06 · Brad and Bill as informed peer 6/10 Multimodal Advances and the Realtime API Apoorv probes the architectural difference between stitched speech-to-text-to-speech pipelines and the native end-to-end Realtime API. Olivier and Sherwin detail why native audio-to-audio preserves critical acoustic cues like inflection and emotion.38:06–42:51 · Brad and Bill as informed peer 4/10 Model Customization and Reinforcement Fine-Tuning (RFT) Sherwin demystifies Reinforcement Fine-Tuning (RFT), distinguishing it from traditional Supervised Fine-Tuning by emphasizing verifiable task reward functions. Olivier adds that while base models handle behavior steering well, RFT is mandatory for frontier domain capabilities.42:51–50:22 · Brad and Bill as informed peer 5/10 Rapid Fire: Long and Short Investment Bets In the rapid-fire section, Sherwin offers contrarian bets by going long on esports and short on ephemeral AI tooling frameworks like RL environments. Olivier shares his short on rote memorization in education and long on life sciences administrative automation.50:22–54:08 · Brad and Bill as informed peer 4/10 Favorite Underrated AI Tools and the Evolution of Codex The discussion covers tools like Granola and Codex CLI paired with GPT-5. The group lightly debates the timeline of Codex's release before clarifying the distinction between the legacy Codex model and the modern CLI agent.54:08–59:43 · Brad and Bill as informed peer 4/10 The Future of Software Engineers and AI-Native Youth Sherwin and Olivier discuss the democratization of software development through AI tools, citing a Reddit story of someone building bespoke software for a non-verbal sibling. They advise young students to leverage their innate AI-native fluency and emphasize critical thinking over memorization.59:43–1:05:02 · Brad and Bill as informed peer 3/10 Personal OpenAI Milestones: Roses, Buds, and Thorns The guests reflect on personal highs and lows at OpenAI, discussing the November 2023 board crisis, major API outages, and the launch sprint for GPT-5 and DevDay. Both articulate the specific moments that made them fully 'AGI-pilled.'2:48–6:04 · Guest teaching 3/10 OpenAI's Enterprise Strategy and Core Mission Apoorv opens by framing OpenAI's enterprise efforts against its public-facing ChatGPT persona. Sherwin and Olivier explain how B2B and API distribution directly fulfill OpenAI's core mission of spreading AGI benefits broadly. The exchange is highly collaborative and informative.6:04–11:30 · Guest teaching 4/10 Case Study: T-Mobile and Forward Deployed Engineering Olivier details their deployment at T-Mobile combining voice support and custom evaluation frameworks. Apoorv draws directly on his four-year background in forward deployed engineering at Palantir to probe into integration layers above raw models.11:30–16:46 · Guest teaching 5/10 Case Study: Amgen and Healthcare Innovation Olivier and Sherwin detail deployments at Amgen and Los Alamos National Laboratory. Sherwin shares technical specifics about running an air-gapped o3 model physically on the Venado supercomputer, which Apoorv receives with genuine fascination.16:46–20:17 · Guest teaching 4/10 Keys to Successful Enterprise AI Deployments Apoorv brings up the MIT report asserting that 95% of AI enterprise deployments fail to challenge the guests on failure modes. Olivier counters with pattern matching from hundreds of enterprise accounts, highlighting top-down buy-in, tiger teams, and rigorous bottom-up evaluation sets.20:17–25:38 · Guest teaching 5/10 Physical versus Digital Autonomy and Environment Scaffolding Apoorv presents a paradox comparing physical autonomy in self-driving cars with lagging digital autonomy in web agents. Sherwin and Olivier push back on the timeline and economic metrics, pointing out that physical driving benefits from extensive standardized infrastructure like roads and traffic lights that digital environments currently lack.25:38–29:21 · Guest teaching 4/10 Unpacking GPT-5: Architecture, Thinking, and Latency Trade-offs Apoorv inquires about the development philosophy and benchmark saturation of GPT-5. Olivier and Sherwin explain the deliberate trade-offs between deep reasoning tokens and inference latency for product builders.29:21–33:00 · Guest teaching 4/10 GPT-5 Real-World Performance and Prompting Dynamics Sherwin describes real-world customer reactions to GPT-5, highlighting reduced hallucinations alongside over-literal instruction following. He illustrates how legacy prompt engineering tricks caused the model to produce overly terse outputs until prompts were cleaned up.33:00–38:06 · Guest teaching 5/10 Multimodal Advances and the Realtime API Apoorv probes the architectural difference between stitched speech-to-text-to-speech pipelines and the native end-to-end Realtime API. Olivier and Sherwin detail why native audio-to-audio preserves critical acoustic cues like inflection and emotion.38:06–42:51 · Guest teaching 5/10 Model Customization and Reinforcement Fine-Tuning (RFT) Sherwin demystifies Reinforcement Fine-Tuning (RFT), distinguishing it from traditional Supervised Fine-Tuning by emphasizing verifiable task reward functions. Olivier adds that while base models handle behavior steering well, RFT is mandatory for frontier domain capabilities.42:51–50:22 · Guest teaching 5/10 Rapid Fire: Long and Short Investment Bets In the rapid-fire section, Sherwin offers contrarian bets by going long on esports and short on ephemeral AI tooling frameworks like RL environments. Olivier shares his short on rote memorization in education and long on life sciences administrative automation.50:22–54:08 · Guest teaching 4/10 Favorite Underrated AI Tools and the Evolution of Codex The discussion covers tools like Granola and Codex CLI paired with GPT-5. The group lightly debates the timeline of Codex's release before clarifying the distinction between the legacy Codex model and the modern CLI agent.54:08–59:43 · Guest teaching 4/10 The Future of Software Engineers and AI-Native Youth Sherwin and Olivier discuss the democratization of software development through AI tools, citing a Reddit story of someone building bespoke software for a non-verbal sibling. They advise young students to leverage their innate AI-native fluency and emphasize critical thinking over memorization.59:43–1:05:02 · Guest teaching 4/10 Personal OpenAI Milestones: Roses, Buds, and Thorns The guests reflect on personal highs and lows at OpenAI, discussing the November 2023 board crisis, major API outages, and the launch sprint for GPT-5 and DevDay. Both articulate the specific moments that made them fully 'AGI-pilled.'2:48–6:04 · Guest disagreement 0/10 OpenAI's Enterprise Strategy and Core Mission Apoorv opens by framing OpenAI's enterprise efforts against its public-facing ChatGPT persona. Sherwin and Olivier explain how B2B and API distribution directly fulfill OpenAI's core mission of spreading AGI benefits broadly. The exchange is highly collaborative and informative.6:04–11:30 · Guest disagreement 1/10 Case Study: T-Mobile and Forward Deployed Engineering Olivier details their deployment at T-Mobile combining voice support and custom evaluation frameworks. Apoorv draws directly on his four-year background in forward deployed engineering at Palantir to probe into integration layers above raw models.11:30–16:46 · Guest disagreement 0/10 Case Study: Amgen and Healthcare Innovation Olivier and Sherwin detail deployments at Amgen and Los Alamos National Laboratory. Sherwin shares technical specifics about running an air-gapped o3 model physically on the Venado supercomputer, which Apoorv receives with genuine fascination.16:46–20:17 · Guest disagreement 1/10 Keys to Successful Enterprise AI Deployments Apoorv brings up the MIT report asserting that 95% of AI enterprise deployments fail to challenge the guests on failure modes. Olivier counters with pattern matching from hundreds of enterprise accounts, highlighting top-down buy-in, tiger teams, and rigorous bottom-up evaluation sets.20:17–25:38 · Guest disagreement 3/10 Physical versus Digital Autonomy and Environment Scaffolding Apoorv presents a paradox comparing physical autonomy in self-driving cars with lagging digital autonomy in web agents. Sherwin and Olivier push back on the timeline and economic metrics, pointing out that physical driving benefits from extensive standardized infrastructure like roads and traffic lights that digital environments currently lack.25:38–29:21 · Guest disagreement 0/10 Unpacking GPT-5: Architecture, Thinking, and Latency Trade-offs Apoorv inquires about the development philosophy and benchmark saturation of GPT-5. Olivier and Sherwin explain the deliberate trade-offs between deep reasoning tokens and inference latency for product builders.29:21–33:00 · Guest disagreement 1/10 GPT-5 Real-World Performance and Prompting Dynamics Sherwin describes real-world customer reactions to GPT-5, highlighting reduced hallucinations alongside over-literal instruction following. He illustrates how legacy prompt engineering tricks caused the model to produce overly terse outputs until prompts were cleaned up.33:00–38:06 · Guest disagreement 0/10 Multimodal Advances and the Realtime API Apoorv probes the architectural difference between stitched speech-to-text-to-speech pipelines and the native end-to-end Realtime API. Olivier and Sherwin detail why native audio-to-audio preserves critical acoustic cues like inflection and emotion.38:06–42:51 · Guest disagreement 1/10 Model Customization and Reinforcement Fine-Tuning (RFT) Sherwin demystifies Reinforcement Fine-Tuning (RFT), distinguishing it from traditional Supervised Fine-Tuning by emphasizing verifiable task reward functions. Olivier adds that while base models handle behavior steering well, RFT is mandatory for frontier domain capabilities.42:51–50:22 · Guest disagreement 3/10 Rapid Fire: Long and Short Investment Bets In the rapid-fire section, Sherwin offers contrarian bets by going long on esports and short on ephemeral AI tooling frameworks like RL environments. Olivier shares his short on rote memorization in education and long on life sciences administrative automation.50:22–54:08 · Guest disagreement 2/10 Favorite Underrated AI Tools and the Evolution of Codex The discussion covers tools like Granola and Codex CLI paired with GPT-5. The group lightly debates the timeline of Codex's release before clarifying the distinction between the legacy Codex model and the modern CLI agent.54:08–59:43 · Guest disagreement 1/10 The Future of Software Engineers and AI-Native Youth Sherwin and Olivier discuss the democratization of software development through AI tools, citing a Reddit story of someone building bespoke software for a non-verbal sibling. They advise young students to leverage their innate AI-native fluency and emphasize critical thinking over memorization.59:43–1:05:02 · Guest disagreement 0/10 Personal OpenAI Milestones: Roses, Buds, and Thorns The guests reflect on personal highs and lows at OpenAI, discussing the November 2023 board crisis, major API outages, and the launch sprint for GPT-5 and DevDay. Both articulate the specific moments that made them fully 'AGI-pilled.'2:48–6:04 · Brad and Bill pushing back 0/10 OpenAI's Enterprise Strategy and Core Mission Apoorv opens by framing OpenAI's enterprise efforts against its public-facing ChatGPT persona. Sherwin and Olivier explain how B2B and API distribution directly fulfill OpenAI's core mission of spreading AGI benefits broadly. The exchange is highly collaborative and informative.6:04–11:30 · Brad and Bill pushing back 2/10 Case Study: T-Mobile and Forward Deployed Engineering Olivier details their deployment at T-Mobile combining voice support and custom evaluation frameworks. Apoorv draws directly on his four-year background in forward deployed engineering at Palantir to probe into integration layers above raw models.11:30–16:46 · Brad and Bill pushing back 0/10 Case Study: Amgen and Healthcare Innovation Olivier and Sherwin detail deployments at Amgen and Los Alamos National Laboratory. Sherwin shares technical specifics about running an air-gapped o3 model physically on the Venado supercomputer, which Apoorv receives with genuine fascination.16:46–20:17 · Brad and Bill pushing back 3/10 Keys to Successful Enterprise AI Deployments Apoorv brings up the MIT report asserting that 95% of AI enterprise deployments fail to challenge the guests on failure modes. Olivier counters with pattern matching from hundreds of enterprise accounts, highlighting top-down buy-in, tiger teams, and rigorous bottom-up evaluation sets.20:17–25:38 · Brad and Bill pushing back 3/10 Physical versus Digital Autonomy and Environment Scaffolding Apoorv presents a paradox comparing physical autonomy in self-driving cars with lagging digital autonomy in web agents. Sherwin and Olivier push back on the timeline and economic metrics, pointing out that physical driving benefits from extensive standardized infrastructure like roads and traffic lights that digital environments currently lack.25:38–29:21 · Brad and Bill pushing back 1/10 Unpacking GPT-5: Architecture, Thinking, and Latency Trade-offs Apoorv inquires about the development philosophy and benchmark saturation of GPT-5. Olivier and Sherwin explain the deliberate trade-offs between deep reasoning tokens and inference latency for product builders.29:21–33:00 · Brad and Bill pushing back 0/10 GPT-5 Real-World Performance and Prompting Dynamics Sherwin describes real-world customer reactions to GPT-5, highlighting reduced hallucinations alongside over-literal instruction following. He illustrates how legacy prompt engineering tricks caused the model to produce overly terse outputs until prompts were cleaned up.33:00–38:06 · Brad and Bill pushing back 2/10 Multimodal Advances and the Realtime API Apoorv probes the architectural difference between stitched speech-to-text-to-speech pipelines and the native end-to-end Realtime API. Olivier and Sherwin detail why native audio-to-audio preserves critical acoustic cues like inflection and emotion.38:06–42:51 · Brad and Bill pushing back 0/10 Model Customization and Reinforcement Fine-Tuning (RFT) Sherwin demystifies Reinforcement Fine-Tuning (RFT), distinguishing it from traditional Supervised Fine-Tuning by emphasizing verifiable task reward functions. Olivier adds that while base models handle behavior steering well, RFT is mandatory for frontier domain capabilities.42:51–50:22 · Brad and Bill pushing back 1/10 Rapid Fire: Long and Short Investment Bets In the rapid-fire section, Sherwin offers contrarian bets by going long on esports and short on ephemeral AI tooling frameworks like RL environments. Olivier shares his short on rote memorization in education and long on life sciences administrative automation.50:22–54:08 · Brad and Bill pushing back 1/10 Favorite Underrated AI Tools and the Evolution of Codex The discussion covers tools like Granola and Codex CLI paired with GPT-5. The group lightly debates the timeline of Codex's release before clarifying the distinction between the legacy Codex model and the modern CLI agent.54:08–59:43 · Brad and Bill pushing back 0/10 The Future of Software Engineers and AI-Native Youth Sherwin and Olivier discuss the democratization of software development through AI tools, citing a Reddit story of someone building bespoke software for a non-verbal sibling. They advise young students to leverage their innate AI-native fluency and emphasize critical thinking over memorization.59:43–1:05:02 · Brad and Bill pushing back 0/10 Personal OpenAI Milestones: Roses, Buds, and Thorns The guests reflect on personal highs and lows at OpenAI, discussing the November 2023 board crisis, major API outages, and the launch sprint for GPT-5 and DevDay. Both articulate the specific moments that made them fully 'AGI-pilled.'

speaking balance: gold is Brad and Bill, purple is the guest (3 minute bins)

0:00 · Brad and Bill 0% · guest 100%0:00 · Brad and Bill 0% · guest 100%3:00 · Brad and Bill 0% · guest 100%3:00 · Brad and Bill 0% · guest 100%6:00 · Brad and Bill 0% · guest 100%6:00 · Brad and Bill 0% · guest 100%9:00 · Brad and Bill 0% · guest 100%9:00 · Brad and Bill 0% · guest 100%12:00 · Brad and Bill 0% · guest 100%12:00 · Brad and Bill 0% · guest 100%15:00 · Brad and Bill 0% · guest 100%15:00 · Brad and Bill 0% · guest 100%18:00 · Brad and Bill 0% · guest 100%18:00 · Brad and Bill 0% · guest 100%21:00 · Brad and Bill 0% · guest 100%21:00 · Brad and Bill 0% · guest 100%24:00 · Brad and Bill 0% · guest 100%24:00 · Brad and Bill 0% · guest 100%27:00 · Brad and Bill 0% · guest 100%27:00 · Brad and Bill 0% · guest 100%30:00 · Brad and Bill 0% · guest 100%30:00 · Brad and Bill 0% · guest 100%33:00 · Brad and Bill 0% · guest 100%33:00 · Brad and Bill 0% · guest 100%36:00 · Brad and Bill 0% · guest 100%36:00 · Brad and Bill 0% · guest 100%39:00 · Brad and Bill 0% · guest 100%39:00 · Brad and Bill 0% · guest 100%42:00 · Brad and Bill 0% · guest 100%42:00 · Brad and Bill 0% · guest 100%45:00 · Brad and Bill 0% · guest 100%45:00 · Brad and Bill 0% · guest 100%48:00 · Brad and Bill 0% · guest 100%48:00 · Brad and Bill 0% · guest 100%51:00 · Brad and Bill 0% · guest 100%51:00 · Brad and Bill 0% · guest 100%54:00 · Brad and Bill 0% · guest 100%54:00 · Brad and Bill 0% · guest 100%57:00 · Brad and Bill 0% · guest 100%57:00 · Brad and Bill 0% · guest 100%1:00:00 · Brad and Bill 0% · guest 100%1:00:00 · Brad and Bill 0% · guest 100%1:03:00 · Brad and Bill 0% · guest 100%1:03:00 · Brad and Bill 0% · guest 100%1:06:00 · Brad and Bill 0% · guest 100%1:06:00 · Brad and Bill 0% · guest 100%
Sharpest disagreement ▶ 45:27 Sherwin shorts AI tooling and RL environments

Sherwin takes an unapologetically contrarian stance against popular startup categories, declaring he is short on evals tooling and RL environment startups because rapid model evolution quickly obsoletes them.

Hardest push from Brad and Bill ▶ 16:46 Apoorv challenges enterprise failure rates with MIT report

Apoorv refuses the uncritical enterprise success narrative by confronting the guests with an MIT report finding that 95% of enterprise AI deployments fail.

Biggest teaching moment ▶ 23:14 Sherwin reframes digital autonomy through physical scaffolding

Sherwin educates the host on why physical self-driving autonomy outpaced digital agents by explaining that real-world roads, lane markers, and traffic laws provide standardized scaffolding that digital environments currently lack.

Brad and Bill hold their own ▶ 8:26 Apoorv leverages Palantir background on FDE dynamics

Apoorv draws on his four years as a forward deployed engineer at Palantir to interrogate the technical and integration stack required above the raw model layer.

the scores for every segment, with the reasoning behind each
ChapterTopicBrad and Bill as informed peerGuest teachingGuest disagreementBrad and Bill pushing backWhy
OpenAI's Enterprise Strategy and Core Mission 4300 Apoorv opens by framing OpenAI's enterprise efforts against its public-facing ChatGPT persona. Sherwin and Olivier explain how B2B and API distribution directly fulfill OpenAI's core mission of spreading AGI benefits broadly. The exchange is highly collaborative and informative.
Case Study: T-Mobile and Forward Deployed Engineering 6412 Olivier details their deployment at T-Mobile combining voice support and custom evaluation frameworks. Apoorv draws directly on his four-year background in forward deployed engineering at Palantir to probe into integration layers above raw models.
Case Study: Amgen and Healthcare Innovation 5500 Olivier and Sherwin detail deployments at Amgen and Los Alamos National Laboratory. Sherwin shares technical specifics about running an air-gapped o3 model physically on the Venado supercomputer, which Apoorv receives with genuine fascination.
Keys to Successful Enterprise AI Deployments 6413 Apoorv brings up the MIT report asserting that 95% of AI enterprise deployments fail to challenge the guests on failure modes. Olivier counters with pattern matching from hundreds of enterprise accounts, highlighting top-down buy-in, tiger teams, and rigorous bottom-up evaluation sets.
Physical versus Digital Autonomy and Environment Scaffolding 6533 Apoorv presents a paradox comparing physical autonomy in self-driving cars with lagging digital autonomy in web agents. Sherwin and Olivier push back on the timeline and economic metrics, pointing out that physical driving benefits from extensive standardized infrastructure like roads and traffic lights that digital environments currently lack.
Unpacking GPT-5: Architecture, Thinking, and Latency Trade-offs 5401 Apoorv inquires about the development philosophy and benchmark saturation of GPT-5. Olivier and Sherwin explain the deliberate trade-offs between deep reasoning tokens and inference latency for product builders.
GPT-5 Real-World Performance and Prompting Dynamics 4410 Sherwin describes real-world customer reactions to GPT-5, highlighting reduced hallucinations alongside over-literal instruction following. He illustrates how legacy prompt engineering tricks caused the model to produce overly terse outputs until prompts were cleaned up.
Multimodal Advances and the Realtime API 6502 Apoorv probes the architectural difference between stitched speech-to-text-to-speech pipelines and the native end-to-end Realtime API. Olivier and Sherwin detail why native audio-to-audio preserves critical acoustic cues like inflection and emotion.
Model Customization and Reinforcement Fine-Tuning (RFT) 4510 Sherwin demystifies Reinforcement Fine-Tuning (RFT), distinguishing it from traditional Supervised Fine-Tuning by emphasizing verifiable task reward functions. Olivier adds that while base models handle behavior steering well, RFT is mandatory for frontier domain capabilities.
Rapid Fire: Long and Short Investment Bets 5531 In the rapid-fire section, Sherwin offers contrarian bets by going long on esports and short on ephemeral AI tooling frameworks like RL environments. Olivier shares his short on rote memorization in education and long on life sciences administrative automation.
Favorite Underrated AI Tools and the Evolution of Codex 4421 The discussion covers tools like Granola and Codex CLI paired with GPT-5. The group lightly debates the timeline of Codex's release before clarifying the distinction between the legacy Codex model and the modern CLI agent.
The Future of Software Engineers and AI-Native Youth 4410 Sherwin and Olivier discuss the democratization of software development through AI tools, citing a Reddit story of someone building bespoke software for a non-verbal sibling. They advise young students to leverage their innate AI-native fluency and emphasize critical thinking over memorization.
Personal OpenAI Milestones: Roses, Buds, and Thorns 3400 The guests reflect on personal highs and lows at OpenAI, discussing the November 2023 board crisis, major API outages, and the launch sprint for GPT-5 and DevDay. Both articulate the specific moments that made them fully 'AGI-pilled.'

Statements from this episode (31)

Assertion Supported
Wu: OpenAI's original product was its developer API, not ChatGPT
“When I joined OpenAI around three years ago to work on the API, it was actually the only product that we had. [178] Sherwin Wu: So I think a lot of people actually forget this, where the original product for, from OpenAI actually was not ChatGPT. [183] Sherwin…”
Sherwin Wu Sep 11, 2025 ▶ 2:49
Assertion Supported
Wu: ChatGPT is roughly the fifth largest website in the world
“ChatGPT obviously is really, really, really big now. [233] Sherwin Wu: It's I think like the fifth largest website in the world.”
Sherwin Wu Sep 11, 2025 ▶ 3:51
Assertion Not checkable as stated
Wu: The majority of the startup ecosystem builds on OpenAI's API
“The biggest product that we have is obviously our developer platform, which is our API. [271] Sherwin Wu: You know, many developers, you know, the majority of the startup ecosystem builds on top of this, as well as a lot of digital natives, Fortune 500 enterpr…”
Sherwin Wu Sep 11, 2025 ▶ 4:21
Assertion Supported
Godement: OpenAI models power live voice support features in T-Mobile app
“And so we've been working with T-Mobile pretty much for the past year at that point to basically automate like not only like text support but also voice support. And so today, like, you know, there are like features like in the T-Mobile app That if you call, a…”
Olivier Godman Sep 11, 2025 ▶ 7:25
Assertion Not checkable as stated
Wu: OpenAI used T-Mobile deployment learnings to improve core Realtime models
“And a lot of the improvements that we actually got into the model came out of, you know, the learnings that we have from T-Mobile. It brings in a lot of other change from other customers, but because we were so deeply embedded into T-Mobile and we were able to…”
Sherwin Wu Sep 11, 2025 ▶ 11:03
Disclosure
Godement: OpenAI is working with Amgen to accelerate drug development
“We've been working essentially with AppGen to essentially speed up, like the drug, like development and like communication process.”
Olivier Godman Sep 11, 2025 ▶ 11:50
Insight
Godement: Healthcare AI needs split into R&D data analysis and regulatory documentation
“When I look at those healthcare companies, I feel like there are two big buckets of needs. One is like Pure R&D. It's like, you know, you're seeing like a massive amount of data and like you have super smart scientists who are trying to, you know, combine, tes…”
Olivier Godman Sep 11, 2025 ▶ 12:08
Assertion Supported
Wu: OpenAI deployed o3 on an air-gapped Los Alamos supercomputer
“We actually did a custom on-prem deployment with them onto one of their supercomputers called Venado. And so this actually involves a bunch of, you know very bespoke work with some FDs also with a lot of our developer team. To actually bring one of our reasoni…”
Sherwin Wu Sep 11, 2025 ▶ 14:46
Assertion Supported
Wu: Los Alamos OpenAI supercomputer deployment is shared with Lawrence Livermore and Sandia
“The other cool thing is it's actually being shared between Los Alamos and some of the other labs Lawrence Livermore Sandia as well because it, it's the supercomputer setup where they can all kind of connect with it remotely.”
Sherwin Wu Sep 11, 2025 ▶ 16:35
Insight
Godement: Enterprise AI Deployments Fail Without Subject Matter Experts on Tiger Teams
“The reality is, like, the standard, like, operating procedures, like the SOPs, are largely in people's heads. And so, unless you have that tiger team, like, mix of, like, technical and, like, you know, subject matter experts, Really hard, like, to get somethin…”
Olivier Godman Sep 11, 2025 ▶ 18:43
Insight
Wu: Enterprise AI Evals Must Be Built Bottom-Up by Operators
“And evals also, oftentimes, need to come up bottom up. Right? Because all of these things are kind of in people's heads, in the actual operator's heads. Like, it's actually very hard to have a top-down mandate of, like, you got, like, this is how the evals sho…”
Sherwin Wu Sep 11, 2025 ▶ 19:17
Insight
Wu: Failed Enterprise AI Deployments Usually Lack Data Scaffolding
“My hunch is some of the enterprise deployments that don't actually work out likely don't have the scaffolding or infrastructure for these agents to interact with as well. A lot of the, like, really successful deployments that we've made, a lot of what our FDs …”
Sherwin Wu Sep 11, 2025 ▶ 24:41
Disclosure
Godement: GPT-5 is the first release tuned through months of customer testing
“On the behavior of the model, I think it's the first model, like, large model release, for which we have worked so closely with a bunch of customers for, like, month and month, essentially, to better understand, like, what are the concrete, like, locks?”
Olivier Godman Sep 11, 2025 ▶ 26:49
Assertion Not checkable as stated
Wu: GPT-5 Pro solves previously unsolved problems but takes ten minutes
“These like unsolved problems that none of the other models could handle, you throw out a GPT-V Pro, and it just like one shots it is pretty crazy, but the trade-off here is you're waiting for 10 minutes.”
Sherwin Wu Sep 11, 2025 ▶ 28:22
Insight
Wu: Product users often prefer instant substandard AI answers over 10-minute waits
“As a product builder, there's a latency, there's a real latency trade-off that you have to deal with where, you know, your user might not be happy waiting 10 minutes for, like, the best answer in the world. It might be more okay with the substandard answer and…”
Sherwin Wu Sep 11, 2025 ▶ 28:54
Opinion
Wu: GPT-5 solves coding problems no other AI model can solve
“Especially for like coding use cases, especially at the, you know, at the, when it thinks for a while, it'll usually solve problems that no other models can solve.”
Sherwin Wu Sep 11, 2025 ▶ 29:42
Assertion Not checkable as stated
Wu: GPT-5 hallucinations dropped to near zero on certain benchmark evaluations
“I think there was an eval that showed that hallucinations basically went to zero for a lot of this.”
Sherwin Wu Sep 11, 2025 ▶ 30:00
Disclosure
Godement: OpenAI Has Closed Part of the Intelligence Gap Between Voice and Text
“One of the feedback that we've received is, you know, Because, like, text was so much in the back on intelligence, like, people felt, like, in particular on voice, that the model was somewhat a little less intelligent, and, you know, until you actually see it,…”
Olivier Godman Sep 11, 2025 ▶ 33:35
Insight
Wu: Reinforcement fine-tuning is an order of magnitude more powerful than SFT
“Reinforcement fine tuning introduces like RL or reinforcement learning to this loop. Way more complex, way more finicky, but an order of magnitude more powerful.”
Sherwin Wu Sep 11, 2025 ▶ 40:09
Assertion Supported
Wu: Accordance achieved SOTA results on Tax Bench via OpenAI RFT
“There's another startup called Accordance that's doing this in the tax space. I think they've been targeting an eval called Tax Bench, which looks at, you know, CPA style tasks as well. And because they, because, you know, they're able to turn it into a very g…”
Sherwin Wu Sep 11, 2025 ▶ 41:24
Prediction Not checkable as stated
Godement: Reinforcement fine-tuning will become the norm for pushing AI capability frontiers
“Pushing the frontier on, like, actual capabilities, my hunch is that RFT will pretty much become the norm. Like, you know, if you are actually pushing in your field, like, you know, intelligence, like, you know, to a pretty high point, like, at some point, lik…”
Olivier Godman Sep 11, 2025 ▶ 42:10
Opinion
Wu: Short on the entire AI tooling startup category
“I'm short on the entire category of like tooling around AI AI products.”
Sherwin Wu Sep 11, 2025 ▶ 45:29
Opinion
Wu: Short on reinforcement learning environment startups
“RL environments I think are really big right now as well. Unfortunately, I'm very short on those. not really I don't really see a lot of potential there. See a lot of potential and reinforcement learning and applying it, but I think the startup space around R…”
Sherwin Wu Sep 11, 2025 ▶ 46:01
Prediction Not checkable as stated
Godement: Healthcare will benefit most from AI in next 1–2 years
“Frankly, I think healthcare is probably the industry that will benefit the most from AI in the next, like, year or two.”
Olivier Godman Sep 11, 2025 ▶ 48:23
Opinion
Wu: GPT-5 enables single-generation code completions in the Codex CLI
“The second thing, honestly, is, is GPT-V. Like, I just think GPT-V really allows the product to shine. It's, you know, at the end of the day, this is kind of a, this is a product that really is dependent on the model, underlying model. And when you have to, yo…”
Sherwin Wu Sep 11, 2025 ▶ 52:33
Assertion Not checkable as stated
Wu: OpenAI is currently operating in a GPU crunch
“We are in a GPU crunch, so we'll see how, you know, how long that goes.”
Sherwin Wu Sep 11, 2025 ▶ 54:05
Opinion
Godement: The world suffers from a massive software shortage
“I buy completely the thesis that there is a massive software shortage. Like, in the world.”
Olivier Godman Sep 11, 2025 ▶ 55:51
Prediction Not checkable as stated
Godement: Job skill sets will reconfigure toward far more people coding
“I expect that we'll see, like, way more, a sort of a reconfiguration of, like, people's, like, you know, job and skill set where way more people code. Like, you know, I expect that product managers are going to code, like, more and more, for instance.”
Olivier Godman Sep 11, 2025 ▶ 56:10
Disclosure
Godement: OpenAI PMs stopped writing PRDs to vibe-code prototypes instead
“We started, like, essentially not doing, like, PRDs, like, product requirements documents. You know, classic PM thing. You write, like, five pages, like, my product does that, et cetera. And, you know, PMs have been basically VibeCoding prototypes.”
Olivier Godman Sep 11, 2025 ▶ 56:28
Opinion
Godement: The November 2023 board crisis made OpenAI culturally stronger and more resilient
“I feel it made OpenEye stronger for real now, essentially, when I look after the fact. When I look at, you know, other, like, you know, news, like departures, or, you know, whatever, like, you know, bad news, essentially, I feel the company has built, like, yo…”
Olivier Godman Sep 11, 2025 ▶ 1:00:55
Assertion Supported
Wu: GPT-4 already existed internally at OpenAI by September 2022
“The first one was right when I joined the company in September, 20, 22. We, it was pre-TiGPT. But at the time, GPT-IV already existed internally.”
Sherwin Wu Sep 11, 2025 ▶ 1:06:47
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 40 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.