May 15, 2026 · 46m · neon-show

The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO

Ameya Bhatawdekar · 36m spoken Siddhartha Ahluwalia · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Neon Show, host Siddharth Ahluwalia interviews Ameya, Field CTO at Braintrust, on moving beyond chaotic 'vibe coding' to build predictable, production-grade enterprise AI applications. Ameya explains the imperative of continuous evaluation datasets, deep agent observability, automated debugging with meta-agents, and the broader 20-year evolution of machine learning architectures.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

Siddhartha as informed peer 3.9 Guest teaching 4.1 Guest disagreement 1.0 Siddhartha pushing back 1.1
05100:0015:0030:0045:002:21–4:59 · Siddhartha as informed peer 3/10 Understanding Evals and Observability in the Gen AI Paradigm The host asks a straightforward foundational question about definitions of evals and observability. The guest gives a structured technical breakdown contrasting traditional ML training with prompt and context conditioning in Gen AI.5:01–7:53 · Siddhartha as informed peer 3/10 Real-World Agent Deployments and Industrial Applications Siddharth asks for customer case studies illustrating production agent deployments. Ameya provides rich examples spanning construction RFP automation to automated travel and subscription assistants.7:55–11:34 · Siddhartha as informed peer 4/10 How Leading Tech Enterprises Implement AI Evals The host brings up well-known Braintrust adopters like Notion, Stripe, and Zapier, asking why classic APM monitoring is inadequate. Ameya explains the probabilistic nature of LLM interactions and the necessity of full trace logging.11:34–13:51 · Siddhartha as informed peer 2/10 Ameya's Engineering Background and Transition to Braintrust The host inquires about Ameya's role as a Field CTO. Ameya shares his career history across Microsoft, Meta, and Dropbox and how Braintrust solved production chaos.13:54–16:31 · Siddhartha as informed peer 3/10 Leading Post-Sales Engineering and Scaling AI Best Practices The host probes Ameya's choice to move into post-sales field engineering rather than purely internal building. Ameya clarifies the leverage of working across multiple high-profile teams.16:35–21:23 · Siddhartha as informed peer 5/10 Deconstructing Root Causes of AI Agent Failures Siddharth demonstrates domain awareness by referencing Foundation Capital articles on context graphs and context engineering. Ameya elaborates on capturing complete conversational context and agent introspection via tools like Loop.21:26–24:17 · Siddhartha as informed peer 4/10 AI Debugging: Traces, Scorers, and Production Clustering The host asks how debugging functions when agentic architectures execute complex steps, suggesting chat-based inspection. Ameya explains custom deterministic scorers, LLM judges, and cluster analysis across production traces.24:17–27:42 · Siddhartha as informed peer 4/10 Vibe Coding vs. Rigorous Evals in Modern Development Siddharth brings up vibe coding and asks whether teams that move fast with weak evals outperform methodical teams long term. Ameya firmly warns about accumulating technical debt and user churn when relying on vibe-based development.27:46–30:24 · Siddhartha as informed peer 4/10 Enterprise Trust and the Engineering Surrounding Foundation Models The host asks what determines enterprise trust when models underneath are identical commodity models. Ameya details that the surrounding engineering layers, guardrails, and context pipelines are what truly make or break the enterprise product.30:25–36:32 · Siddhartha as informed peer 6/10 Twenty Years of Machine Learning: Classic ML to Neural Networks Both host and guest engage in an extensive retrospective covering classical ML in the 2000s to deep learning in the 2010s. Siddharth shares his direct experience building facial recognition models for the Government of India in 2011.36:32–43:40 · Siddhartha as informed peer 5/10 Transformers, Ambient AI, and the Emergence of Reasoning The conversation moves into transformer architecture and emergent reasoning capabilities. Siddharth presents a detailed case study of Buddy AI in elder care, while Ameya explains how scaling token prediction yields emergent reasoning.2:21–4:59 · Guest teaching 5/10 Understanding Evals and Observability in the Gen AI Paradigm The host asks a straightforward foundational question about definitions of evals and observability. The guest gives a structured technical breakdown contrasting traditional ML training with prompt and context conditioning in Gen AI.5:01–7:53 · Guest teaching 4/10 Real-World Agent Deployments and Industrial Applications Siddharth asks for customer case studies illustrating production agent deployments. Ameya provides rich examples spanning construction RFP automation to automated travel and subscription assistants.7:55–11:34 · Guest teaching 5/10 How Leading Tech Enterprises Implement AI Evals The host brings up well-known Braintrust adopters like Notion, Stripe, and Zapier, asking why classic APM monitoring is inadequate. Ameya explains the probabilistic nature of LLM interactions and the necessity of full trace logging.11:34–13:51 · Guest teaching 3/10 Ameya's Engineering Background and Transition to Braintrust The host inquires about Ameya's role as a Field CTO. Ameya shares his career history across Microsoft, Meta, and Dropbox and how Braintrust solved production chaos.13:54–16:31 · Guest teaching 3/10 Leading Post-Sales Engineering and Scaling AI Best Practices The host probes Ameya's choice to move into post-sales field engineering rather than purely internal building. Ameya clarifies the leverage of working across multiple high-profile teams.16:35–21:23 · Guest teaching 4/10 Deconstructing Root Causes of AI Agent Failures Siddharth demonstrates domain awareness by referencing Foundation Capital articles on context graphs and context engineering. Ameya elaborates on capturing complete conversational context and agent introspection via tools like Loop.21:26–24:17 · Guest teaching 4/10 AI Debugging: Traces, Scorers, and Production Clustering The host asks how debugging functions when agentic architectures execute complex steps, suggesting chat-based inspection. Ameya explains custom deterministic scorers, LLM judges, and cluster analysis across production traces.24:17–27:42 · Guest teaching 4/10 Vibe Coding vs. Rigorous Evals in Modern Development Siddharth brings up vibe coding and asks whether teams that move fast with weak evals outperform methodical teams long term. Ameya firmly warns about accumulating technical debt and user churn when relying on vibe-based development.27:46–30:24 · Guest teaching 4/10 Enterprise Trust and the Engineering Surrounding Foundation Models The host asks what determines enterprise trust when models underneath are identical commodity models. Ameya details that the surrounding engineering layers, guardrails, and context pipelines are what truly make or break the enterprise product.30:25–36:32 · Guest teaching 4/10 Twenty Years of Machine Learning: Classic ML to Neural Networks Both host and guest engage in an extensive retrospective covering classical ML in the 2000s to deep learning in the 2010s. Siddharth shares his direct experience building facial recognition models for the Government of India in 2011.36:32–43:40 · Guest teaching 5/10 Transformers, Ambient AI, and the Emergence of Reasoning The conversation moves into transformer architecture and emergent reasoning capabilities. Siddharth presents a detailed case study of Buddy AI in elder care, while Ameya explains how scaling token prediction yields emergent reasoning.2:21–4:59 · Guest disagreement 1/10 Understanding Evals and Observability in the Gen AI Paradigm The host asks a straightforward foundational question about definitions of evals and observability. The guest gives a structured technical breakdown contrasting traditional ML training with prompt and context conditioning in Gen AI.5:01–7:53 · Guest disagreement 1/10 Real-World Agent Deployments and Industrial Applications Siddharth asks for customer case studies illustrating production agent deployments. Ameya provides rich examples spanning construction RFP automation to automated travel and subscription assistants.7:55–11:34 · Guest disagreement 1/10 How Leading Tech Enterprises Implement AI Evals The host brings up well-known Braintrust adopters like Notion, Stripe, and Zapier, asking why classic APM monitoring is inadequate. Ameya explains the probabilistic nature of LLM interactions and the necessity of full trace logging.11:34–13:51 · Guest disagreement 0/10 Ameya's Engineering Background and Transition to Braintrust The host inquires about Ameya's role as a Field CTO. Ameya shares his career history across Microsoft, Meta, and Dropbox and how Braintrust solved production chaos.13:54–16:31 · Guest disagreement 1/10 Leading Post-Sales Engineering and Scaling AI Best Practices The host probes Ameya's choice to move into post-sales field engineering rather than purely internal building. Ameya clarifies the leverage of working across multiple high-profile teams.16:35–21:23 · Guest disagreement 1/10 Deconstructing Root Causes of AI Agent Failures Siddharth demonstrates domain awareness by referencing Foundation Capital articles on context graphs and context engineering. Ameya elaborates on capturing complete conversational context and agent introspection via tools like Loop.21:26–24:17 · Guest disagreement 1/10 AI Debugging: Traces, Scorers, and Production Clustering The host asks how debugging functions when agentic architectures execute complex steps, suggesting chat-based inspection. Ameya explains custom deterministic scorers, LLM judges, and cluster analysis across production traces.24:17–27:42 · Guest disagreement 2/10 Vibe Coding vs. Rigorous Evals in Modern Development Siddharth brings up vibe coding and asks whether teams that move fast with weak evals outperform methodical teams long term. Ameya firmly warns about accumulating technical debt and user churn when relying on vibe-based development.27:46–30:24 · Guest disagreement 2/10 Enterprise Trust and the Engineering Surrounding Foundation Models The host asks what determines enterprise trust when models underneath are identical commodity models. Ameya details that the surrounding engineering layers, guardrails, and context pipelines are what truly make or break the enterprise product.30:25–36:32 · Guest disagreement 0/10 Twenty Years of Machine Learning: Classic ML to Neural Networks Both host and guest engage in an extensive retrospective covering classical ML in the 2000s to deep learning in the 2010s. Siddharth shares his direct experience building facial recognition models for the Government of India in 2011.36:32–43:40 · Guest disagreement 1/10 Transformers, Ambient AI, and the Emergence of Reasoning The conversation moves into transformer architecture and emergent reasoning capabilities. Siddharth presents a detailed case study of Buddy AI in elder care, while Ameya explains how scaling token prediction yields emergent reasoning.2:21–4:59 · Siddhartha pushing back 1/10 Understanding Evals and Observability in the Gen AI Paradigm The host asks a straightforward foundational question about definitions of evals and observability. The guest gives a structured technical breakdown contrasting traditional ML training with prompt and context conditioning in Gen AI.5:01–7:53 · Siddhartha pushing back 1/10 Real-World Agent Deployments and Industrial Applications Siddharth asks for customer case studies illustrating production agent deployments. Ameya provides rich examples spanning construction RFP automation to automated travel and subscription assistants.7:55–11:34 · Siddhartha pushing back 2/10 How Leading Tech Enterprises Implement AI Evals The host brings up well-known Braintrust adopters like Notion, Stripe, and Zapier, asking why classic APM monitoring is inadequate. Ameya explains the probabilistic nature of LLM interactions and the necessity of full trace logging.11:34–13:51 · Siddhartha pushing back 0/10 Ameya's Engineering Background and Transition to Braintrust The host inquires about Ameya's role as a Field CTO. Ameya shares his career history across Microsoft, Meta, and Dropbox and how Braintrust solved production chaos.13:54–16:31 · Siddhartha pushing back 1/10 Leading Post-Sales Engineering and Scaling AI Best Practices The host probes Ameya's choice to move into post-sales field engineering rather than purely internal building. Ameya clarifies the leverage of working across multiple high-profile teams.16:35–21:23 · Siddhartha pushing back 2/10 Deconstructing Root Causes of AI Agent Failures Siddharth demonstrates domain awareness by referencing Foundation Capital articles on context graphs and context engineering. Ameya elaborates on capturing complete conversational context and agent introspection via tools like Loop.21:26–24:17 · Siddhartha pushing back 1/10 AI Debugging: Traces, Scorers, and Production Clustering The host asks how debugging functions when agentic architectures execute complex steps, suggesting chat-based inspection. Ameya explains custom deterministic scorers, LLM judges, and cluster analysis across production traces.24:17–27:42 · Siddhartha pushing back 1/10 Vibe Coding vs. Rigorous Evals in Modern Development Siddharth brings up vibe coding and asks whether teams that move fast with weak evals outperform methodical teams long term. Ameya firmly warns about accumulating technical debt and user churn when relying on vibe-based development.27:46–30:24 · Siddhartha pushing back 1/10 Enterprise Trust and the Engineering Surrounding Foundation Models The host asks what determines enterprise trust when models underneath are identical commodity models. Ameya details that the surrounding engineering layers, guardrails, and context pipelines are what truly make or break the enterprise product.30:25–36:32 · Siddhartha pushing back 1/10 Twenty Years of Machine Learning: Classic ML to Neural Networks Both host and guest engage in an extensive retrospective covering classical ML in the 2000s to deep learning in the 2010s. Siddharth shares his direct experience building facial recognition models for the Government of India in 2011.36:32–43:40 · Siddhartha pushing back 1/10 Transformers, Ambient AI, and the Emergence of Reasoning The conversation moves into transformer architecture and emergent reasoning capabilities. Siddharth presents a detailed case study of Buddy AI in elder care, while Ameya explains how scaling token prediction yields emergent reasoning.

speaking balance: gold is Siddhartha, purple is the guest (3 minute bins)

0:00 · Siddhartha 0% · guest 100%0:00 · Siddhartha 0% · guest 100%3:00 · Siddhartha 0% · guest 100%3:00 · Siddhartha 0% · guest 100%6:00 · Siddhartha 0% · guest 100%6:00 · Siddhartha 0% · guest 100%9:00 · Siddhartha 0% · guest 100%9:00 · Siddhartha 0% · guest 100%12:00 · Siddhartha 0% · guest 100%12:00 · Siddhartha 0% · guest 100%15:00 · Siddhartha 0% · guest 100%15:00 · Siddhartha 0% · guest 100%18:00 · Siddhartha 0% · guest 100%18:00 · Siddhartha 0% · guest 100%21:00 · Siddhartha 0% · guest 100%21:00 · Siddhartha 0% · guest 100%24:00 · Siddhartha 0% · guest 100%24:00 · Siddhartha 0% · guest 100%27:00 · Siddhartha 0% · guest 100%27:00 · Siddhartha 0% · guest 100%30:00 · Siddhartha 0% · guest 100%30:00 · Siddhartha 0% · guest 100%33:00 · Siddhartha 0% · guest 100%33:00 · Siddhartha 0% · guest 100%36:00 · Siddhartha 0% · guest 100%36:00 · Siddhartha 0% · guest 100%39:00 · Siddhartha 0% · guest 100%39:00 · Siddhartha 0% · guest 100%42:00 · Siddhartha 0% · guest 100%42:00 · Siddhartha 0% · guest 100%45:00 · Siddhartha 0% · guest 100%45:00 · Siddhartha 0% · guest 100%
Sharpest disagreement ▶ 29:39 Vibes are just guessing

Ameya directly dismisses vibe-based development approaches, arguing that shipping without rigorous evals means developers are purely guessing and will end up with demo-ware.

Hardest push from Siddhartha ▶ 9:45 Pushing on necessity of observability

The host challenges why traditional APM monitoring solutions are not enough and forces Ameya to justify why dedicated Gen AI observability is required.

Biggest teaching moment ▶ 2:26 Deconstructing Gen AI paradigm shift

Ameya walks through the fundamental technical shift from training bespoke ML models to conditioning general-purpose foundation models through prompt engineering and evals.

Siddhartha holds their own ▶ 34:21 Demonstrating 2011 neural network deployment

Siddharth demonstrates his technical credibility by recounting his early work building government-scale airport facial recognition systems using neural networks in 2011.

the scores for every segment, with the reasoning behind each
ChapterTopicSiddhartha as informed peerGuest teachingGuest disagreementSiddhartha pushing backWhy
Understanding Evals and Observability in the Gen AI Paradigm 3511 The host asks a straightforward foundational question about definitions of evals and observability. The guest gives a structured technical breakdown contrasting traditional ML training with prompt and context conditioning in Gen AI.
Real-World Agent Deployments and Industrial Applications 3411 Siddharth asks for customer case studies illustrating production agent deployments. Ameya provides rich examples spanning construction RFP automation to automated travel and subscription assistants.
How Leading Tech Enterprises Implement AI Evals 4512 The host brings up well-known Braintrust adopters like Notion, Stripe, and Zapier, asking why classic APM monitoring is inadequate. Ameya explains the probabilistic nature of LLM interactions and the necessity of full trace logging.
Ameya's Engineering Background and Transition to Braintrust 2300 The host inquires about Ameya's role as a Field CTO. Ameya shares his career history across Microsoft, Meta, and Dropbox and how Braintrust solved production chaos.
Leading Post-Sales Engineering and Scaling AI Best Practices 3311 The host probes Ameya's choice to move into post-sales field engineering rather than purely internal building. Ameya clarifies the leverage of working across multiple high-profile teams.
Deconstructing Root Causes of AI Agent Failures 5412 Siddharth demonstrates domain awareness by referencing Foundation Capital articles on context graphs and context engineering. Ameya elaborates on capturing complete conversational context and agent introspection via tools like Loop.
AI Debugging: Traces, Scorers, and Production Clustering 4411 The host asks how debugging functions when agentic architectures execute complex steps, suggesting chat-based inspection. Ameya explains custom deterministic scorers, LLM judges, and cluster analysis across production traces.
Vibe Coding vs. Rigorous Evals in Modern Development 4421 Siddharth brings up vibe coding and asks whether teams that move fast with weak evals outperform methodical teams long term. Ameya firmly warns about accumulating technical debt and user churn when relying on vibe-based development.
Enterprise Trust and the Engineering Surrounding Foundation Models 4421 The host asks what determines enterprise trust when models underneath are identical commodity models. Ameya details that the surrounding engineering layers, guardrails, and context pipelines are what truly make or break the enterprise product.
Twenty Years of Machine Learning: Classic ML to Neural Networks 6401 Both host and guest engage in an extensive retrospective covering classical ML in the 2000s to deep learning in the 2010s. Siddharth shares his direct experience building facial recognition models for the Government of India in 2011.
Transformers, Ambient AI, and the Emergence of Reasoning 5511 The conversation moves into transformer architecture and emergent reasoning capabilities. Siddharth presents a detailed case study of Buddy AI in elder care, while Ameya explains how scaling token prediction yields emergent reasoning.

Statements from this episode (10)

Insight
Bhatawdekar: Gen AI Replaced Custom Model Training with Prompt Conditioning
“After the Gen AI revolution, like the way now we build intelligent applications is we take models off the shelf. These are general purpose models. They can reason on a variety of tasks, and then we condition the models to work a specific way by doing prompt en…”
Ameya Bhatawdekar May 15, 2026 ▶ 2:52
Insight
Bhatawdekar: Gen AI Systems Require Observability Feedback Loops for Evals
“So when you're building Gen AI systems, you really want that feedback loop of observability that helps you build better evals, that helps you ship better AI.”
Ameya Bhatawdekar May 15, 2026 ▶ 4:49
Assertion Supported
Bhatawdekar: Construction companies use AI to create RFP proposals from drawings
“I've seen systems where construction companies are able to now put together effective proposals using complex engineering drawings, architectural plans, specifications to submit proposals for new RFPs.”
Ameya Bhatawdekar May 15, 2026 ▶ 6:24
Assertion Not checkable as stated
Bhatawdekar: Notion, Stripe, and Zapier run automated evals on every system modification
“All of these companies are building agents, intelligent systems that are performing specialized tasks for whatever use cases they have, and so all of them build evals that reflect how their systems are expected to behave in production. They're all following si…”
Ameya Bhatawdekar May 15, 2026 ▶ 8:07
Insight
Bhatawdekar: Production eval datasets must be continuously updated from live logs
“These eval datasets, they are not static. As the teams look at their logs, at how their systems are working in the real world, in the production use cases, they're able to leverage those insights to continually augment their eval datasets. And so these eval da…”
Ameya Bhatawdekar May 15, 2026 ▶ 8:51
Insight
Bhatawdekar: AI systems require full reasoning traces to evaluate response quality
“These AI systems need to log the entire trace of how the AI reasoned on the initial input. What were the tool calls it made? How did it interact with the LLMs? How did it sort of ultimately generate the response? And did that response actually meet the user in…”
Ameya Bhatawdekar May 15, 2026 ▶ 10:51
Insight
Bhatawdekar: Span-level scorers pinpoint errors in AI agent execution
“You can define those as deterministic functions, you know, implemented in code, or you can use LLM as judges, but then you can evaluate like, how did each span perform? And that can give you a fairly good way to zero in on problematic areas of your agents.”
Ameya Bhatawdekar May 15, 2026 ▶ 22:24
Insight
Bhatawdekar: Rigorous evals are existential for AI apps built with 'vibe coding'
“When you're building these intelligent agentic applications using Vibe Coding evals almost become existential. You know, that's the only way you have a high degree of confidence that what you've built is going to work well.”
Ameya Bhatawdekar May 15, 2026 ▶ 24:53
Insight
Ameya Bhatawdekar: Enterprise AI quality comes from surrounding engineering, not just models
“These intelligent systems, these AI systems are not just a model, right? There's a lot of layering that happens on top of these models. These systems have to really deliver specific capabilities or specific experiences that help people do certain specific task…”
Ameya Bhatawdekar May 15, 2026 ▶ 28:02
Insight
Bhatawdekar: AI models natively handle orchestration, replacing complex external engineering frameworks
“People built these very fancy frameworks and systems that were fairly complex and complicated to improve the orchestration capability of the model. And, you know, there were some very impressive engineering feats that happened as a result of that. But now the …”
Ameya Bhatawdekar May 15, 2026 ▶ 44:58
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.