Insight certainty 4/5 debate potential 2/5

Wu: Enterprise AI Evals Must Be Built Bottom-Up by Operators

Sherwin Wu · Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod · Sep 11, 2025 · at 19:17

Sherwin Wu, Head of Engineering for OpenAI Platform, discusses why top-down mandates fail when creating evaluation frameworks for enterprise AI deployments.

0:00 / 0:12exact quote · 12.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“And evals also, oftentimes, need to come up bottom up. Right? Because all of these things are kind of in people's heads, in the actual operator's heads. Like, it's actually very hard to have a top-down mandate of, like, you got, like, this is how the evals should look. A lot of it needs the bottom-up adoption.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Sherwin Wu

Assertion Supported
Wu: OpenAI deployed o3 on an air-gapped Los Alamos supercomputer
“We actually did a custom on-prem deployment with them onto one of their supercomputers called Venado. And so this actually involves a bunch of, you know very bespoke work with some FDs also with a lot of our developer team. To actually bring one of our reasoni…”
Sherwin Wu Sep 11, 2025 ▶ 14:46 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
Assertion Not checkable as stated
Wu: GPT-5 hallucinations dropped to near zero on certain benchmark evaluations
“I think there was an eval that showed that hallucinations basically went to zero for a lot of this.”
Sherwin Wu Sep 11, 2025 ▶ 30:00 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
Opinion
Wu: Short on the entire AI tooling startup category
“I'm short on the entire category of like tooling around AI AI products.”
Sherwin Wu Sep 11, 2025 ▶ 45:29 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
Opinion
Wu: Short on reinforcement learning environment startups
“RL environments I think are really big right now as well. Unfortunately, I'm very short on those. not really I don't really see a lot of potential there. See a lot of potential and reinforcement learning and applying it, but I think the startup space around R…”
Sherwin Wu Sep 11, 2025 ▶ 46:01 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
Assertion Supported
Wu: Los Alamos OpenAI supercomputer deployment is shared with Lawrence Livermore and Sandia
“The other cool thing is it's actually being shared between Los Alamos and some of the other labs Lawrence Livermore Sandia as well because it, it's the supercomputer setup where they can all kind of connect with it remotely.”
Sherwin Wu Sep 11, 2025 ▶ 16:35 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
Insight
Wu: Failed Enterprise AI Deployments Usually Lack Data Scaffolding
“My hunch is some of the enterprise deployments that don't actually work out likely don't have the scaffolding or infrastructure for these agents to interact with as well. A lot of the, like, really successful deployments that we've made, a lot of what our FDs …”
Sherwin Wu Sep 11, 2025 ▶ 24:41 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 40 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.