May 15, 2026 · 46m · neon-show
The “Messy State” of AI & How to Fix It | Ameya, Braintrust CTO
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Neon Show, host Siddharth Ahluwalia interviews Ameya, Field CTO at Braintrust, on moving beyond chaotic 'vibe coding' to build predictable, production-grade enterprise AI applications. Ameya explains the imperative of continuous evaluation datasets, deep agent observability, automated debugging with meta-agents, and the broader 20-year evolution of machine learning architectures.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is Siddhartha, purple is the guest (3 minute bins)
Ameya directly dismisses vibe-based development approaches, arguing that shipping without rigorous evals means developers are purely guessing and will end up with demo-ware.
Hardest push from Siddhartha ▶ 9:45 Pushing on necessity of observabilityThe host challenges why traditional APM monitoring solutions are not enough and forces Ameya to justify why dedicated Gen AI observability is required.
Biggest teaching moment ▶ 2:26 Deconstructing Gen AI paradigm shiftAmeya walks through the fundamental technical shift from training bespoke ML models to conditioning general-purpose foundation models through prompt engineering and evals.
Siddhartha holds their own ▶ 34:21 Demonstrating 2011 neural network deploymentSiddharth demonstrates his technical credibility by recounting his early work building government-scale airport facial recognition systems using neural networks in 2011.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Siddhartha as informed peer | Guest teaching | Guest disagreement | Siddhartha pushing back | Why |
|---|---|---|---|---|---|---|
| Understanding Evals and Observability in the Gen AI Paradigm | 3 | 5 | 1 | 1 | The host asks a straightforward foundational question about definitions of evals and observability. The guest gives a structured technical breakdown contrasting traditional ML training with prompt and context conditioning in Gen AI. | |
| Real-World Agent Deployments and Industrial Applications | 3 | 4 | 1 | 1 | Siddharth asks for customer case studies illustrating production agent deployments. Ameya provides rich examples spanning construction RFP automation to automated travel and subscription assistants. | |
| How Leading Tech Enterprises Implement AI Evals | 4 | 5 | 1 | 2 | The host brings up well-known Braintrust adopters like Notion, Stripe, and Zapier, asking why classic APM monitoring is inadequate. Ameya explains the probabilistic nature of LLM interactions and the necessity of full trace logging. | |
| Ameya's Engineering Background and Transition to Braintrust | 2 | 3 | 0 | 0 | The host inquires about Ameya's role as a Field CTO. Ameya shares his career history across Microsoft, Meta, and Dropbox and how Braintrust solved production chaos. | |
| Leading Post-Sales Engineering and Scaling AI Best Practices | 3 | 3 | 1 | 1 | The host probes Ameya's choice to move into post-sales field engineering rather than purely internal building. Ameya clarifies the leverage of working across multiple high-profile teams. | |
| Deconstructing Root Causes of AI Agent Failures | 5 | 4 | 1 | 2 | Siddharth demonstrates domain awareness by referencing Foundation Capital articles on context graphs and context engineering. Ameya elaborates on capturing complete conversational context and agent introspection via tools like Loop. | |
| AI Debugging: Traces, Scorers, and Production Clustering | 4 | 4 | 1 | 1 | The host asks how debugging functions when agentic architectures execute complex steps, suggesting chat-based inspection. Ameya explains custom deterministic scorers, LLM judges, and cluster analysis across production traces. | |
| Vibe Coding vs. Rigorous Evals in Modern Development | 4 | 4 | 2 | 1 | Siddharth brings up vibe coding and asks whether teams that move fast with weak evals outperform methodical teams long term. Ameya firmly warns about accumulating technical debt and user churn when relying on vibe-based development. | |
| Enterprise Trust and the Engineering Surrounding Foundation Models | 4 | 4 | 2 | 1 | The host asks what determines enterprise trust when models underneath are identical commodity models. Ameya details that the surrounding engineering layers, guardrails, and context pipelines are what truly make or break the enterprise product. | |
| Twenty Years of Machine Learning: Classic ML to Neural Networks | 6 | 4 | 0 | 1 | Both host and guest engage in an extensive retrospective covering classical ML in the 2000s to deep learning in the 2010s. Siddharth shares his direct experience building facial recognition models for the Government of India in 2011. | |
| Transformers, Ambient AI, and the Emergence of Reasoning | 5 | 5 | 1 | 1 | The conversation moves into transformer architecture and emergent reasoning capabilities. Siddharth presents a detailed case study of Buddy AI in elder care, while Ameya explains how scaling token prediction yields emergent reasoning. |