Sep 11, 2025 · 1h 8m · bg2-pod
Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the BG2 Pod, host Apoorv Agrawal interviews OpenAI Platform leaders Sherwin Wu and Olivier Godement to explore OpenAI's enterprise strategy, technical breakthroughs in GPT-5, and real-world deployment case studies across major institutions.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is Brad and Bill, purple is the guest (3 minute bins)
Sherwin takes an unapologetically contrarian stance against popular startup categories, declaring he is short on evals tooling and RL environment startups because rapid model evolution quickly obsoletes them.
Hardest push from Brad and Bill ▶ 16:46 Apoorv challenges enterprise failure rates with MIT reportApoorv refuses the uncritical enterprise success narrative by confronting the guests with an MIT report finding that 95% of enterprise AI deployments fail.
Biggest teaching moment ▶ 23:14 Sherwin reframes digital autonomy through physical scaffoldingSherwin educates the host on why physical self-driving autonomy outpaced digital agents by explaining that real-world roads, lane markers, and traffic laws provide standardized scaffolding that digital environments currently lack.
Brad and Bill hold their own ▶ 8:26 Apoorv leverages Palantir background on FDE dynamicsApoorv draws on his four years as a forward deployed engineer at Palantir to interrogate the technical and integration stack required above the raw model layer.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Brad and Bill as informed peer | Guest teaching | Guest disagreement | Brad and Bill pushing back | Why |
|---|---|---|---|---|---|---|
| OpenAI's Enterprise Strategy and Core Mission | 4 | 3 | 0 | 0 | Apoorv opens by framing OpenAI's enterprise efforts against its public-facing ChatGPT persona. Sherwin and Olivier explain how B2B and API distribution directly fulfill OpenAI's core mission of spreading AGI benefits broadly. The exchange is highly collaborative and informative. | |
| Case Study: T-Mobile and Forward Deployed Engineering | 6 | 4 | 1 | 2 | Olivier details their deployment at T-Mobile combining voice support and custom evaluation frameworks. Apoorv draws directly on his four-year background in forward deployed engineering at Palantir to probe into integration layers above raw models. | |
| Case Study: Amgen and Healthcare Innovation | 5 | 5 | 0 | 0 | Olivier and Sherwin detail deployments at Amgen and Los Alamos National Laboratory. Sherwin shares technical specifics about running an air-gapped o3 model physically on the Venado supercomputer, which Apoorv receives with genuine fascination. | |
| Keys to Successful Enterprise AI Deployments | 6 | 4 | 1 | 3 | Apoorv brings up the MIT report asserting that 95% of AI enterprise deployments fail to challenge the guests on failure modes. Olivier counters with pattern matching from hundreds of enterprise accounts, highlighting top-down buy-in, tiger teams, and rigorous bottom-up evaluation sets. | |
| Physical versus Digital Autonomy and Environment Scaffolding | 6 | 5 | 3 | 3 | Apoorv presents a paradox comparing physical autonomy in self-driving cars with lagging digital autonomy in web agents. Sherwin and Olivier push back on the timeline and economic metrics, pointing out that physical driving benefits from extensive standardized infrastructure like roads and traffic lights that digital environments currently lack. | |
| Unpacking GPT-5: Architecture, Thinking, and Latency Trade-offs | 5 | 4 | 0 | 1 | Apoorv inquires about the development philosophy and benchmark saturation of GPT-5. Olivier and Sherwin explain the deliberate trade-offs between deep reasoning tokens and inference latency for product builders. | |
| GPT-5 Real-World Performance and Prompting Dynamics | 4 | 4 | 1 | 0 | Sherwin describes real-world customer reactions to GPT-5, highlighting reduced hallucinations alongside over-literal instruction following. He illustrates how legacy prompt engineering tricks caused the model to produce overly terse outputs until prompts were cleaned up. | |
| Multimodal Advances and the Realtime API | 6 | 5 | 0 | 2 | Apoorv probes the architectural difference between stitched speech-to-text-to-speech pipelines and the native end-to-end Realtime API. Olivier and Sherwin detail why native audio-to-audio preserves critical acoustic cues like inflection and emotion. | |
| Model Customization and Reinforcement Fine-Tuning (RFT) | 4 | 5 | 1 | 0 | Sherwin demystifies Reinforcement Fine-Tuning (RFT), distinguishing it from traditional Supervised Fine-Tuning by emphasizing verifiable task reward functions. Olivier adds that while base models handle behavior steering well, RFT is mandatory for frontier domain capabilities. | |
| Rapid Fire: Long and Short Investment Bets | 5 | 5 | 3 | 1 | In the rapid-fire section, Sherwin offers contrarian bets by going long on esports and short on ephemeral AI tooling frameworks like RL environments. Olivier shares his short on rote memorization in education and long on life sciences administrative automation. | |
| Favorite Underrated AI Tools and the Evolution of Codex | 4 | 4 | 2 | 1 | The discussion covers tools like Granola and Codex CLI paired with GPT-5. The group lightly debates the timeline of Codex's release before clarifying the distinction between the legacy Codex model and the modern CLI agent. | |
| The Future of Software Engineers and AI-Native Youth | 4 | 4 | 1 | 0 | Sherwin and Olivier discuss the democratization of software development through AI tools, citing a Reddit story of someone building bespoke software for a non-verbal sibling. They advise young students to leverage their innate AI-native fluency and emphasize critical thinking over memorization. | |
| Personal OpenAI Milestones: Roses, Buds, and Thorns | 3 | 4 | 0 | 0 | The guests reflect on personal highs and lows at OpenAI, discussing the November 2023 board crisis, major API outages, and the launch sprint for GPT-5 and DevDay. Both articulate the specific moments that made them fully 'AGI-pilled.' |