Insight certainty 4/5 debate potential 2/5

Badam: Coding agents cannot rely solely on offline evaluation datasets

Kiriti Badam · Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon · Jan 11, 2026 · at 42:03

Kiriti Badam discusses AI evaluation methodology and product engineering practices for OpenAI Codex with Lenny Rachitsky.

0:00 / 0:35exact quote · 35.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Coding agents are extremely unique compared to agents for other domains in the sense that These are actually built for customizability and these are built for engineers. So coding agent is not a product which is going to solve like this top five workflows or like top six workflows or whatever, right? It's meant to be customizable in multi different ways. And the implication of that is that your product is going to be used in different integrations and different kinds of tools and different kinds of things. So it gets really hard to build an evaluation data set for all kinds of interactions that Your customers are going to use your product for, right?”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Kiriti Badam

Insight
Badam: Relying solely on either evals or production monitoring is inadequate
“So I feel devals are important. Production monitoring is important, but this notion of only one of them is going to solve things for you. That is completely dismissible in my opinion.”
Kiriti Badam Jan 11, 2026 ▶ 37:49 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Insight
Badam: Relying entirely on fixed evals without team testing fails
“I don't think like if anybody's coming and seeing that, like my, I have this Concrete set of evals that I can, like, bet my life on, and then I don't need to think about anything else. Like, it's not going to work, and every new model that we're going to launc…”
Kiriti Badam Jan 11, 2026 ▶ 44:25 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Insight
Badam: Peer-to-peer multi-agent gossip protocols exceed current AI model capabilities
“If you're building a supervisor agent and there are like sub agents that actually do the work for the super agent, supervisor agent, That is a very successful pattern, but coming with this notion of I'm going to divide the responsibilities based on functionali…”
Kiriti Badam Jan 11, 2026 ▶ 1:02:14 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Insight
Badam: AI fails to create value because it lacks workflow context
“Where is AI failing to create value today, it's mainly about not understanding the context. And the reason that it's not understanding the context is it's not plugged into the right places where actual work is happening.”
Kiriti Badam Jan 11, 2026 ▶ 1:05:42 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Prediction Not checkable as stated
Badam: Future AI products will be proactive background agents
“And now when you extend this to more complex tasks, like a coding agent, which says that like, okay, I have fixed five of your linear tickets and here are the patches just to review them at the start of your day. So I feel that is going to be like extremely us…”
Kiriti Badam Jan 11, 2026 ▶ 1:06:28 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Insight
Badam: AI products should start with high human control before granting agency
“When you don't start with like agents with all the tools and all the context that you have in the company in day one and expect it to work or like, you know, even tinker at that level, you need to be deliberately starting in places where there is minimal impac…”
Kiriti Badam Jan 11, 2026 ▶ 12:05 Why most AI products fail: Lessons from 50+ AI deployments at OpenAI, Google & Amazon
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.