Oct 11, 2024 · 1h 56m · latent-space
Production AI Engineering starts with Evals
⌖ your search result is the highlighted band (1:45:02–1:45:39). Playback starts there
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this in-depth interview, Braintrust founder and CEO Ankur Goyal discusses the transition of AI development from academic data science to developer-centric software engineering, drawing on his career at SingleStore, Impira, and Figma. He outlines the architectural foundations of production AI systems, emphasizing declarative evaluations, reliable proxy routing, and simple code-first design over fragile agent frameworks.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Ankur bluntly scoffs at using alternative open source hosting providers for production workloads, sarcastically remarking 'Good luck' and comparing them to unscalable vintage hardware hosts.
Hardest push from the hosts ▶ 1:34:38 Refusing black-box reasoning token billingThe host rejects Ankur's premise that native model reasoning makes external agent workflows obsolete, directly calling out OpenAI charging for hidden reasoning tokens as unacceptable.
Biggest teaching moment ▶ 1:15:30 Deconstructing the standalone vector database thesisAnkur delivers an architectural schooling on database internals, explaining why vector search is merely an algorithm rather than an execution paradigm that justifies a standalone database.
The host holds their own ▶ 1:45:04 Coining the Code Core vs LM Core paradigmThe host demonstrates deep software architecture expertise by extending Ankur's insights into a formal systems design thesis modeled after functional core imperative shells.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| From Microsoft and Research to SingleStore Employee Two | 4 | 2 | 0 | 0 | The host warmly introduces Ankur Goyal and prompts him on his early trajectory through Microsoft, research, and joining SingleStore as employee number two. Ankur shares his candid journey with an agreeable narrative flow. | |
| The HTAP Database Dream and Modern Storage Packaging | 6 | 5 | 2 | 2 | The host asks technical questions regarding HTAP databases and Honeycomb's wide column store approach. Ankur dives deep into database storage architecture, comparing Snowflake's variant type to DuckDB structs and ClickHouse. | |
| Founding Impira and Hard Lessons in Enterprise Sales | 4 | 4 | 1 | 1 | The host prompts Ankur on founding Impira and translating engineering skills to startup sales. Ankur recounts humbling lessons learned regarding enterprise sales execution and selling to business units instead of developers. | |
| The Prioritization Dilemma in Unstructured Data Systems | 5 | 5 | 2 | 1 | The host asks why unstructured data conversion startups struggle despite LLMs. Ankur explains the executive prioritization dilemma where enterprise leaders prefer rebuilding future-proof customer experiences over fixing back-office ingestion. | |
| LLM Disruption and Impira's Acquisition by Figma | 5 | 4 | 1 | 2 | The host asks Ankur to reveal the untold Impira acquisition story by Figma. Ankur openly details how BERT and LLMs demolished bespoke OCR pipelines, forcing an honest strategic pivot and sale. | |
| Inside Figma AI and Designing for UI Engineers | 5 | 4 | 1 | 1 | The host and guest discuss Figma's annual release cadence versus rapid AI iteration cycles. Ankur highlights why vector representation is token-inefficient and frames Figma's AI opportunity around UI code generation. | |
| The AI Engineering Shift and the Role of Evals | 6 | 4 | 2 | 3 | The host challenges the novelty of evals in AI engineering, citing ML history. Ankur clarifies how transformer architectures enable standard software engineers to build AI applications without relying on data science abstractions like NumPy or dataframes. | |
| Braintrust's Evolution: Tracing, Debugging, and IDE Playgrounds | 6 | 3 | 1 | 2 | Ankur breaks down Braintrust's evolution from simple eval metrics to trace logging, debuggers, and interactive IDE playgrounds. The host engages on IDE convergence and analogies to Cursor. | |
| Live Product Demonstration: Sandboxed Evals and Custom Tools | 5 | 3 | 1 | 2 | Ankur performs a live demo showcasing Braintrust's sandboxed Python evaluators and custom TypeScript tool definitions. The host interjects constructively to clarify LLM-as-judge prompt wiring. | |
| Why Specialized Evaluation Platforms Eclipse Spreadsheets | 6 | 5 | 2 | 2 | The host raises DIY solutions like Claude in Sheets and Humanloop. Ankur details why declarative eval data structures, parallel execution, and automated token attribution surpass basic spreadsheet tooling. | |
| Hybrid On-Premise Architecture and Developer-First Design | 5 | 4 | 3 | 2 | Ankur describes Braintrust's contrarian bets: hybrid on-prem architecture, heavy TypeScript SDK support, and ignoring VC pushback that claimed evals were an unviable CI/CD-style market. | |
| AI Engineer World's Fair and Real-Time Customer Collaboration | 6 | 3 | 0 | 1 | The host and guest discuss the AI Engineer World's Fair, Zapier's customer collaboration, and how live conference interactions unblock enterprise proof-of-concepts. | |
| Analyzing the AI Stack and Passing on Vector Databases | 6 | 6 | 3 | 2 | The host asks why a database veteran chose not to build a vector database company. Ankur explains that vector search is a single query type overshadowed by RBAC, joins, and application metadata already managed in primary databases. | |
| Automatic Optimization, Fine-Tuning, and Simple Code | 6 | 5 | 3 | 2 | Ankur argues that fine-tuning is an implementation detail rather than an enduring business outcome, favoring automatic prompt optimization and simpler code over complex PyTorch-style frameworks like DSPy. | |
| Frontier Model Dynamics: OpenAI, Anthropic, and Gateway Routing | 6 | 5 | 2 | 2 | Ankur breaks down production model shares between OpenAI and Anthropic, highlighting Haiku's JSON tool-calling edge and OpenAI's superior endpoint reliability and rate limits over raw hyperscaler deployments. | |
| OpenAI o1 and the Obsolescence of Complex Agent Frameworks | 7 | 5 | 4 | 6 | Ankur predicts OpenAI o1 will make complex graph-based agent frameworks obsolete. The host pushes back vigorously, emphasizing that enterprise developers refuse to pay for invisible reasoning tokens without granular execution control. | |
| Production Realities of Open Source and Inference Economics | 6 | 6 | 5 | 3 | The host questions the low share of open-source models in production. Ankur responds combatively, ridiculing open-source inference provider reliability compared to hyperscalers. | |
| Production Workload Patterns: The Rise of Code-Core Systems | 7 | 4 | 2 | 2 | The host synthesizes real-world AI architecture into 'Code Core vs LM Core', matching Ankur's observation that production workloads favor simple imperative code with selective LLM invocations over sprawling autonomous loops. | |
| Database Passions, Foundation Labs, and the Braintrust Team | 4 | 3 | 1 | 1 | Ankur shares personal reflections on database engineering nostalgia, his Braintrust founding team, and why he would turn down an acquisition offer to stay focused on high-conviction product building. | |
| Investor Relationships, Series A Milestone, and Hiring Call | 3 | 2 | 0 | 0 | The host and guest close by discussing Alana Goyal's hands-on investing methodology, Braintrust's Series A milestone, and active hiring needs in San Francisco. |