Oct 11, 2024 · 1h 56m · latent-space

Production AI Engineering starts with Evals

Ankur Goyal · 1h 28m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this in-depth interview, Braintrust founder and CEO Ankur Goyal discusses the transition of AI development from academic data science to developer-centric software engineering, drawing on his career at SingleStore, Impira, and Figma. He outlines the architectural foundations of production AI systems, emphasizing declarative evaluations, reliable proxy routing, and simple code-first design over fragile agent frameworks.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.4 Guest teaching 4.1 Guest disagreement 1.8 The hosts pushing back 1.9
05100:0020:0040:001:00:001:20:001:40:000:04–3:41 · The hosts as informed peer 4/10 From Microsoft and Research to SingleStore Employee Two The host warmly introduces Ankur Goyal and prompts him on his early trajectory through Microsoft, research, and joining SingleStore as employee number two. Ankur shares his candid journey with an agreeable narrative flow.3:41–8:27 · The hosts as informed peer 6/10 The HTAP Database Dream and Modern Storage Packaging The host asks technical questions regarding HTAP databases and Honeycomb's wide column store approach. Ankur dives deep into database storage architecture, comparing Snowflake's variant type to DuckDB structs and ClickHouse.8:28–13:33 · The hosts as informed peer 4/10 Founding Impira and Hard Lessons in Enterprise Sales The host prompts Ankur on founding Impira and translating engineering skills to startup sales. Ankur recounts humbling lessons learned regarding enterprise sales execution and selling to business units instead of developers.13:34–16:14 · The hosts as informed peer 5/10 The Prioritization Dilemma in Unstructured Data Systems The host asks why unstructured data conversion startups struggle despite LLMs. Ankur explains the executive prioritization dilemma where enterprise leaders prefer rebuilding future-proof customer experiences over fixing back-office ingestion.16:16–25:25 · The hosts as informed peer 5/10 LLM Disruption and Impira's Acquisition by Figma The host asks Ankur to reveal the untold Impira acquisition story by Figma. Ankur openly details how BERT and LLMs demolished bespoke OCR pipelines, forcing an honest strategic pivot and sale.25:25–29:14 · The hosts as informed peer 5/10 Inside Figma AI and Designing for UI Engineers The host and guest discuss Figma's annual release cadence versus rapid AI iteration cycles. Ankur highlights why vector representation is token-inefficient and frames Figma's AI opportunity around UI code generation.29:14–36:00 · The hosts as informed peer 6/10 The AI Engineering Shift and the Role of Evals The host challenges the novelty of evals in AI engineering, citing ML history. Ankur clarifies how transformer architectures enable standard software engineers to build AI applications without relying on data science abstractions like NumPy or dataframes.36:00–41:44 · The hosts as informed peer 6/10 Braintrust's Evolution: Tracing, Debugging, and IDE Playgrounds Ankur breaks down Braintrust's evolution from simple eval metrics to trace logging, debuggers, and interactive IDE playgrounds. The host engages on IDE convergence and analogies to Cursor.41:45–52:05 · The hosts as informed peer 5/10 Live Product Demonstration: Sandboxed Evals and Custom Tools Ankur performs a live demo showcasing Braintrust's sandboxed Python evaluators and custom TypeScript tool definitions. The host interjects constructively to clarify LLM-as-judge prompt wiring.52:05–58:38 · The hosts as informed peer 6/10 Why Specialized Evaluation Platforms Eclipse Spreadsheets The host raises DIY solutions like Claude in Sheets and Humanloop. Ankur details why declarative eval data structures, parallel execution, and automated token attribution surpass basic spreadsheet tooling.58:38–1:06:13 · The hosts as informed peer 5/10 Hybrid On-Premise Architecture and Developer-First Design Ankur describes Braintrust's contrarian bets: hybrid on-prem architecture, heavy TypeScript SDK support, and ignoring VC pushback that claimed evals were an unviable CI/CD-style market.1:06:13–1:11:08 · The hosts as informed peer 6/10 AI Engineer World's Fair and Real-Time Customer Collaboration The host and guest discuss the AI Engineer World's Fair, Zapier's customer collaboration, and how live conference interactions unblock enterprise proof-of-concepts.1:11:08–1:18:22 · The hosts as informed peer 6/10 Analyzing the AI Stack and Passing on Vector Databases The host asks why a database veteran chose not to build a vector database company. Ankur explains that vector search is a single query type overshadowed by RBAC, joins, and application metadata already managed in primary databases.1:18:22–1:24:33 · The hosts as informed peer 6/10 Automatic Optimization, Fine-Tuning, and Simple Code Ankur argues that fine-tuning is an implementation detail rather than an enduring business outcome, favoring automatic prompt optimization and simpler code over complex PyTorch-style frameworks like DSPy.1:24:33–1:31:28 · The hosts as informed peer 6/10 Frontier Model Dynamics: OpenAI, Anthropic, and Gateway Routing Ankur breaks down production model shares between OpenAI and Anthropic, highlighting Haiku's JSON tool-calling edge and OpenAI's superior endpoint reliability and rate limits over raw hyperscaler deployments.1:31:28–1:36:41 · The hosts as informed peer 7/10 OpenAI o1 and the Obsolescence of Complex Agent Frameworks Ankur predicts OpenAI o1 will make complex graph-based agent frameworks obsolete. The host pushes back vigorously, emphasizing that enterprise developers refuse to pay for invisible reasoning tokens without granular execution control.1:36:41–1:39:32 · The hosts as informed peer 6/10 Production Realities of Open Source and Inference Economics The host questions the low share of open-source models in production. Ankur responds combatively, ridiculing open-source inference provider reliability compared to hyperscalers.1:39:32–1:46:05 · The hosts as informed peer 7/10 Production Workload Patterns: The Rise of Code-Core Systems The host synthesizes real-world AI architecture into 'Code Core vs LM Core', matching Ankur's observation that production workloads favor simple imperative code with selective LLM invocations over sprawling autonomous loops.1:46:05–1:52:21 · The hosts as informed peer 4/10 Database Passions, Foundation Labs, and the Braintrust Team Ankur shares personal reflections on database engineering nostalgia, his Braintrust founding team, and why he would turn down an acquisition offer to stay focused on high-conviction product building.1:52:22–1:56:05 · The hosts as informed peer 3/10 Investor Relationships, Series A Milestone, and Hiring Call The host and guest close by discussing Alana Goyal's hands-on investing methodology, Braintrust's Series A milestone, and active hiring needs in San Francisco.0:04–3:41 · Guest teaching 2/10 From Microsoft and Research to SingleStore Employee Two The host warmly introduces Ankur Goyal and prompts him on his early trajectory through Microsoft, research, and joining SingleStore as employee number two. Ankur shares his candid journey with an agreeable narrative flow.3:41–8:27 · Guest teaching 5/10 The HTAP Database Dream and Modern Storage Packaging The host asks technical questions regarding HTAP databases and Honeycomb's wide column store approach. Ankur dives deep into database storage architecture, comparing Snowflake's variant type to DuckDB structs and ClickHouse.8:28–13:33 · Guest teaching 4/10 Founding Impira and Hard Lessons in Enterprise Sales The host prompts Ankur on founding Impira and translating engineering skills to startup sales. Ankur recounts humbling lessons learned regarding enterprise sales execution and selling to business units instead of developers.13:34–16:14 · Guest teaching 5/10 The Prioritization Dilemma in Unstructured Data Systems The host asks why unstructured data conversion startups struggle despite LLMs. Ankur explains the executive prioritization dilemma where enterprise leaders prefer rebuilding future-proof customer experiences over fixing back-office ingestion.16:16–25:25 · Guest teaching 4/10 LLM Disruption and Impira's Acquisition by Figma The host asks Ankur to reveal the untold Impira acquisition story by Figma. Ankur openly details how BERT and LLMs demolished bespoke OCR pipelines, forcing an honest strategic pivot and sale.25:25–29:14 · Guest teaching 4/10 Inside Figma AI and Designing for UI Engineers The host and guest discuss Figma's annual release cadence versus rapid AI iteration cycles. Ankur highlights why vector representation is token-inefficient and frames Figma's AI opportunity around UI code generation.29:14–36:00 · Guest teaching 4/10 The AI Engineering Shift and the Role of Evals The host challenges the novelty of evals in AI engineering, citing ML history. Ankur clarifies how transformer architectures enable standard software engineers to build AI applications without relying on data science abstractions like NumPy or dataframes.36:00–41:44 · Guest teaching 3/10 Braintrust's Evolution: Tracing, Debugging, and IDE Playgrounds Ankur breaks down Braintrust's evolution from simple eval metrics to trace logging, debuggers, and interactive IDE playgrounds. The host engages on IDE convergence and analogies to Cursor.41:45–52:05 · Guest teaching 3/10 Live Product Demonstration: Sandboxed Evals and Custom Tools Ankur performs a live demo showcasing Braintrust's sandboxed Python evaluators and custom TypeScript tool definitions. The host interjects constructively to clarify LLM-as-judge prompt wiring.52:05–58:38 · Guest teaching 5/10 Why Specialized Evaluation Platforms Eclipse Spreadsheets The host raises DIY solutions like Claude in Sheets and Humanloop. Ankur details why declarative eval data structures, parallel execution, and automated token attribution surpass basic spreadsheet tooling.58:38–1:06:13 · Guest teaching 4/10 Hybrid On-Premise Architecture and Developer-First Design Ankur describes Braintrust's contrarian bets: hybrid on-prem architecture, heavy TypeScript SDK support, and ignoring VC pushback that claimed evals were an unviable CI/CD-style market.1:06:13–1:11:08 · Guest teaching 3/10 AI Engineer World's Fair and Real-Time Customer Collaboration The host and guest discuss the AI Engineer World's Fair, Zapier's customer collaboration, and how live conference interactions unblock enterprise proof-of-concepts.1:11:08–1:18:22 · Guest teaching 6/10 Analyzing the AI Stack and Passing on Vector Databases The host asks why a database veteran chose not to build a vector database company. Ankur explains that vector search is a single query type overshadowed by RBAC, joins, and application metadata already managed in primary databases.1:18:22–1:24:33 · Guest teaching 5/10 Automatic Optimization, Fine-Tuning, and Simple Code Ankur argues that fine-tuning is an implementation detail rather than an enduring business outcome, favoring automatic prompt optimization and simpler code over complex PyTorch-style frameworks like DSPy.1:24:33–1:31:28 · Guest teaching 5/10 Frontier Model Dynamics: OpenAI, Anthropic, and Gateway Routing Ankur breaks down production model shares between OpenAI and Anthropic, highlighting Haiku's JSON tool-calling edge and OpenAI's superior endpoint reliability and rate limits over raw hyperscaler deployments.1:31:28–1:36:41 · Guest teaching 5/10 OpenAI o1 and the Obsolescence of Complex Agent Frameworks Ankur predicts OpenAI o1 will make complex graph-based agent frameworks obsolete. The host pushes back vigorously, emphasizing that enterprise developers refuse to pay for invisible reasoning tokens without granular execution control.1:36:41–1:39:32 · Guest teaching 6/10 Production Realities of Open Source and Inference Economics The host questions the low share of open-source models in production. Ankur responds combatively, ridiculing open-source inference provider reliability compared to hyperscalers.1:39:32–1:46:05 · Guest teaching 4/10 Production Workload Patterns: The Rise of Code-Core Systems The host synthesizes real-world AI architecture into 'Code Core vs LM Core', matching Ankur's observation that production workloads favor simple imperative code with selective LLM invocations over sprawling autonomous loops.1:46:05–1:52:21 · Guest teaching 3/10 Database Passions, Foundation Labs, and the Braintrust Team Ankur shares personal reflections on database engineering nostalgia, his Braintrust founding team, and why he would turn down an acquisition offer to stay focused on high-conviction product building.1:52:22–1:56:05 · Guest teaching 2/10 Investor Relationships, Series A Milestone, and Hiring Call The host and guest close by discussing Alana Goyal's hands-on investing methodology, Braintrust's Series A milestone, and active hiring needs in San Francisco.0:04–3:41 · Guest disagreement 0/10 From Microsoft and Research to SingleStore Employee Two The host warmly introduces Ankur Goyal and prompts him on his early trajectory through Microsoft, research, and joining SingleStore as employee number two. Ankur shares his candid journey with an agreeable narrative flow.3:41–8:27 · Guest disagreement 2/10 The HTAP Database Dream and Modern Storage Packaging The host asks technical questions regarding HTAP databases and Honeycomb's wide column store approach. Ankur dives deep into database storage architecture, comparing Snowflake's variant type to DuckDB structs and ClickHouse.8:28–13:33 · Guest disagreement 1/10 Founding Impira and Hard Lessons in Enterprise Sales The host prompts Ankur on founding Impira and translating engineering skills to startup sales. Ankur recounts humbling lessons learned regarding enterprise sales execution and selling to business units instead of developers.13:34–16:14 · Guest disagreement 2/10 The Prioritization Dilemma in Unstructured Data Systems The host asks why unstructured data conversion startups struggle despite LLMs. Ankur explains the executive prioritization dilemma where enterprise leaders prefer rebuilding future-proof customer experiences over fixing back-office ingestion.16:16–25:25 · Guest disagreement 1/10 LLM Disruption and Impira's Acquisition by Figma The host asks Ankur to reveal the untold Impira acquisition story by Figma. Ankur openly details how BERT and LLMs demolished bespoke OCR pipelines, forcing an honest strategic pivot and sale.25:25–29:14 · Guest disagreement 1/10 Inside Figma AI and Designing for UI Engineers The host and guest discuss Figma's annual release cadence versus rapid AI iteration cycles. Ankur highlights why vector representation is token-inefficient and frames Figma's AI opportunity around UI code generation.29:14–36:00 · Guest disagreement 2/10 The AI Engineering Shift and the Role of Evals The host challenges the novelty of evals in AI engineering, citing ML history. Ankur clarifies how transformer architectures enable standard software engineers to build AI applications without relying on data science abstractions like NumPy or dataframes.36:00–41:44 · Guest disagreement 1/10 Braintrust's Evolution: Tracing, Debugging, and IDE Playgrounds Ankur breaks down Braintrust's evolution from simple eval metrics to trace logging, debuggers, and interactive IDE playgrounds. The host engages on IDE convergence and analogies to Cursor.41:45–52:05 · Guest disagreement 1/10 Live Product Demonstration: Sandboxed Evals and Custom Tools Ankur performs a live demo showcasing Braintrust's sandboxed Python evaluators and custom TypeScript tool definitions. The host interjects constructively to clarify LLM-as-judge prompt wiring.52:05–58:38 · Guest disagreement 2/10 Why Specialized Evaluation Platforms Eclipse Spreadsheets The host raises DIY solutions like Claude in Sheets and Humanloop. Ankur details why declarative eval data structures, parallel execution, and automated token attribution surpass basic spreadsheet tooling.58:38–1:06:13 · Guest disagreement 3/10 Hybrid On-Premise Architecture and Developer-First Design Ankur describes Braintrust's contrarian bets: hybrid on-prem architecture, heavy TypeScript SDK support, and ignoring VC pushback that claimed evals were an unviable CI/CD-style market.1:06:13–1:11:08 · Guest disagreement 0/10 AI Engineer World's Fair and Real-Time Customer Collaboration The host and guest discuss the AI Engineer World's Fair, Zapier's customer collaboration, and how live conference interactions unblock enterprise proof-of-concepts.1:11:08–1:18:22 · Guest disagreement 3/10 Analyzing the AI Stack and Passing on Vector Databases The host asks why a database veteran chose not to build a vector database company. Ankur explains that vector search is a single query type overshadowed by RBAC, joins, and application metadata already managed in primary databases.1:18:22–1:24:33 · Guest disagreement 3/10 Automatic Optimization, Fine-Tuning, and Simple Code Ankur argues that fine-tuning is an implementation detail rather than an enduring business outcome, favoring automatic prompt optimization and simpler code over complex PyTorch-style frameworks like DSPy.1:24:33–1:31:28 · Guest disagreement 2/10 Frontier Model Dynamics: OpenAI, Anthropic, and Gateway Routing Ankur breaks down production model shares between OpenAI and Anthropic, highlighting Haiku's JSON tool-calling edge and OpenAI's superior endpoint reliability and rate limits over raw hyperscaler deployments.1:31:28–1:36:41 · Guest disagreement 4/10 OpenAI o1 and the Obsolescence of Complex Agent Frameworks Ankur predicts OpenAI o1 will make complex graph-based agent frameworks obsolete. The host pushes back vigorously, emphasizing that enterprise developers refuse to pay for invisible reasoning tokens without granular execution control.1:36:41–1:39:32 · Guest disagreement 5/10 Production Realities of Open Source and Inference Economics The host questions the low share of open-source models in production. Ankur responds combatively, ridiculing open-source inference provider reliability compared to hyperscalers.1:39:32–1:46:05 · Guest disagreement 2/10 Production Workload Patterns: The Rise of Code-Core Systems The host synthesizes real-world AI architecture into 'Code Core vs LM Core', matching Ankur's observation that production workloads favor simple imperative code with selective LLM invocations over sprawling autonomous loops.1:46:05–1:52:21 · Guest disagreement 1/10 Database Passions, Foundation Labs, and the Braintrust Team Ankur shares personal reflections on database engineering nostalgia, his Braintrust founding team, and why he would turn down an acquisition offer to stay focused on high-conviction product building.1:52:22–1:56:05 · Guest disagreement 0/10 Investor Relationships, Series A Milestone, and Hiring Call The host and guest close by discussing Alana Goyal's hands-on investing methodology, Braintrust's Series A milestone, and active hiring needs in San Francisco.0:04–3:41 · The hosts pushing back 0/10 From Microsoft and Research to SingleStore Employee Two The host warmly introduces Ankur Goyal and prompts him on his early trajectory through Microsoft, research, and joining SingleStore as employee number two. Ankur shares his candid journey with an agreeable narrative flow.3:41–8:27 · The hosts pushing back 2/10 The HTAP Database Dream and Modern Storage Packaging The host asks technical questions regarding HTAP databases and Honeycomb's wide column store approach. Ankur dives deep into database storage architecture, comparing Snowflake's variant type to DuckDB structs and ClickHouse.8:28–13:33 · The hosts pushing back 1/10 Founding Impira and Hard Lessons in Enterprise Sales The host prompts Ankur on founding Impira and translating engineering skills to startup sales. Ankur recounts humbling lessons learned regarding enterprise sales execution and selling to business units instead of developers.13:34–16:14 · The hosts pushing back 1/10 The Prioritization Dilemma in Unstructured Data Systems The host asks why unstructured data conversion startups struggle despite LLMs. Ankur explains the executive prioritization dilemma where enterprise leaders prefer rebuilding future-proof customer experiences over fixing back-office ingestion.16:16–25:25 · The hosts pushing back 2/10 LLM Disruption and Impira's Acquisition by Figma The host asks Ankur to reveal the untold Impira acquisition story by Figma. Ankur openly details how BERT and LLMs demolished bespoke OCR pipelines, forcing an honest strategic pivot and sale.25:25–29:14 · The hosts pushing back 1/10 Inside Figma AI and Designing for UI Engineers The host and guest discuss Figma's annual release cadence versus rapid AI iteration cycles. Ankur highlights why vector representation is token-inefficient and frames Figma's AI opportunity around UI code generation.29:14–36:00 · The hosts pushing back 3/10 The AI Engineering Shift and the Role of Evals The host challenges the novelty of evals in AI engineering, citing ML history. Ankur clarifies how transformer architectures enable standard software engineers to build AI applications without relying on data science abstractions like NumPy or dataframes.36:00–41:44 · The hosts pushing back 2/10 Braintrust's Evolution: Tracing, Debugging, and IDE Playgrounds Ankur breaks down Braintrust's evolution from simple eval metrics to trace logging, debuggers, and interactive IDE playgrounds. The host engages on IDE convergence and analogies to Cursor.41:45–52:05 · The hosts pushing back 2/10 Live Product Demonstration: Sandboxed Evals and Custom Tools Ankur performs a live demo showcasing Braintrust's sandboxed Python evaluators and custom TypeScript tool definitions. The host interjects constructively to clarify LLM-as-judge prompt wiring.52:05–58:38 · The hosts pushing back 2/10 Why Specialized Evaluation Platforms Eclipse Spreadsheets The host raises DIY solutions like Claude in Sheets and Humanloop. Ankur details why declarative eval data structures, parallel execution, and automated token attribution surpass basic spreadsheet tooling.58:38–1:06:13 · The hosts pushing back 2/10 Hybrid On-Premise Architecture and Developer-First Design Ankur describes Braintrust's contrarian bets: hybrid on-prem architecture, heavy TypeScript SDK support, and ignoring VC pushback that claimed evals were an unviable CI/CD-style market.1:06:13–1:11:08 · The hosts pushing back 1/10 AI Engineer World's Fair and Real-Time Customer Collaboration The host and guest discuss the AI Engineer World's Fair, Zapier's customer collaboration, and how live conference interactions unblock enterprise proof-of-concepts.1:11:08–1:18:22 · The hosts pushing back 2/10 Analyzing the AI Stack and Passing on Vector Databases The host asks why a database veteran chose not to build a vector database company. Ankur explains that vector search is a single query type overshadowed by RBAC, joins, and application metadata already managed in primary databases.1:18:22–1:24:33 · The hosts pushing back 2/10 Automatic Optimization, Fine-Tuning, and Simple Code Ankur argues that fine-tuning is an implementation detail rather than an enduring business outcome, favoring automatic prompt optimization and simpler code over complex PyTorch-style frameworks like DSPy.1:24:33–1:31:28 · The hosts pushing back 2/10 Frontier Model Dynamics: OpenAI, Anthropic, and Gateway Routing Ankur breaks down production model shares between OpenAI and Anthropic, highlighting Haiku's JSON tool-calling edge and OpenAI's superior endpoint reliability and rate limits over raw hyperscaler deployments.1:31:28–1:36:41 · The hosts pushing back 6/10 OpenAI o1 and the Obsolescence of Complex Agent Frameworks Ankur predicts OpenAI o1 will make complex graph-based agent frameworks obsolete. The host pushes back vigorously, emphasizing that enterprise developers refuse to pay for invisible reasoning tokens without granular execution control.1:36:41–1:39:32 · The hosts pushing back 3/10 Production Realities of Open Source and Inference Economics The host questions the low share of open-source models in production. Ankur responds combatively, ridiculing open-source inference provider reliability compared to hyperscalers.1:39:32–1:46:05 · The hosts pushing back 2/10 Production Workload Patterns: The Rise of Code-Core Systems The host synthesizes real-world AI architecture into 'Code Core vs LM Core', matching Ankur's observation that production workloads favor simple imperative code with selective LLM invocations over sprawling autonomous loops.1:46:05–1:52:21 · The hosts pushing back 1/10 Database Passions, Foundation Labs, and the Braintrust Team Ankur shares personal reflections on database engineering nostalgia, his Braintrust founding team, and why he would turn down an acquisition offer to stay focused on high-conviction product building.1:52:22–1:56:05 · The hosts pushing back 0/10 Investor Relationships, Series A Milestone, and Hiring Call The host and guest close by discussing Alana Goyal's hands-on investing methodology, Braintrust's Series A milestone, and active hiring needs in San Francisco.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:24:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:27:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:30:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:33:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:36:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:39:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:42:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:45:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:48:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:51:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%1:54:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:37:19 Dismissing open source inference providers

Ankur bluntly scoffs at using alternative open source hosting providers for production workloads, sarcastically remarking 'Good luck' and comparing them to unscalable vintage hardware hosts.

Hardest push from the hosts ▶ 1:34:38 Refusing black-box reasoning token billing

The host rejects Ankur's premise that native model reasoning makes external agent workflows obsolete, directly calling out OpenAI charging for hidden reasoning tokens as unacceptable.

Biggest teaching moment ▶ 1:15:30 Deconstructing the standalone vector database thesis

Ankur delivers an architectural schooling on database internals, explaining why vector search is merely an algorithm rather than an execution paradigm that justifies a standalone database.

The host holds their own ▶ 1:45:04 Coining the Code Core vs LM Core paradigm

The host demonstrates deep software architecture expertise by extending Ankur's insights into a formal systems design thesis modeled after functional core imperative shells.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
From Microsoft and Research to SingleStore Employee Two 4200 The host warmly introduces Ankur Goyal and prompts him on his early trajectory through Microsoft, research, and joining SingleStore as employee number two. Ankur shares his candid journey with an agreeable narrative flow.
The HTAP Database Dream and Modern Storage Packaging 6522 The host asks technical questions regarding HTAP databases and Honeycomb's wide column store approach. Ankur dives deep into database storage architecture, comparing Snowflake's variant type to DuckDB structs and ClickHouse.
Founding Impira and Hard Lessons in Enterprise Sales 4411 The host prompts Ankur on founding Impira and translating engineering skills to startup sales. Ankur recounts humbling lessons learned regarding enterprise sales execution and selling to business units instead of developers.
The Prioritization Dilemma in Unstructured Data Systems 5521 The host asks why unstructured data conversion startups struggle despite LLMs. Ankur explains the executive prioritization dilemma where enterprise leaders prefer rebuilding future-proof customer experiences over fixing back-office ingestion.
LLM Disruption and Impira's Acquisition by Figma 5412 The host asks Ankur to reveal the untold Impira acquisition story by Figma. Ankur openly details how BERT and LLMs demolished bespoke OCR pipelines, forcing an honest strategic pivot and sale.
Inside Figma AI and Designing for UI Engineers 5411 The host and guest discuss Figma's annual release cadence versus rapid AI iteration cycles. Ankur highlights why vector representation is token-inefficient and frames Figma's AI opportunity around UI code generation.
The AI Engineering Shift and the Role of Evals 6423 The host challenges the novelty of evals in AI engineering, citing ML history. Ankur clarifies how transformer architectures enable standard software engineers to build AI applications without relying on data science abstractions like NumPy or dataframes.
Braintrust's Evolution: Tracing, Debugging, and IDE Playgrounds 6312 Ankur breaks down Braintrust's evolution from simple eval metrics to trace logging, debuggers, and interactive IDE playgrounds. The host engages on IDE convergence and analogies to Cursor.
Live Product Demonstration: Sandboxed Evals and Custom Tools 5312 Ankur performs a live demo showcasing Braintrust's sandboxed Python evaluators and custom TypeScript tool definitions. The host interjects constructively to clarify LLM-as-judge prompt wiring.
Why Specialized Evaluation Platforms Eclipse Spreadsheets 6522 The host raises DIY solutions like Claude in Sheets and Humanloop. Ankur details why declarative eval data structures, parallel execution, and automated token attribution surpass basic spreadsheet tooling.
Hybrid On-Premise Architecture and Developer-First Design 5432 Ankur describes Braintrust's contrarian bets: hybrid on-prem architecture, heavy TypeScript SDK support, and ignoring VC pushback that claimed evals were an unviable CI/CD-style market.
AI Engineer World's Fair and Real-Time Customer Collaboration 6301 The host and guest discuss the AI Engineer World's Fair, Zapier's customer collaboration, and how live conference interactions unblock enterprise proof-of-concepts.
Analyzing the AI Stack and Passing on Vector Databases 6632 The host asks why a database veteran chose not to build a vector database company. Ankur explains that vector search is a single query type overshadowed by RBAC, joins, and application metadata already managed in primary databases.
Automatic Optimization, Fine-Tuning, and Simple Code 6532 Ankur argues that fine-tuning is an implementation detail rather than an enduring business outcome, favoring automatic prompt optimization and simpler code over complex PyTorch-style frameworks like DSPy.
Frontier Model Dynamics: OpenAI, Anthropic, and Gateway Routing 6522 Ankur breaks down production model shares between OpenAI and Anthropic, highlighting Haiku's JSON tool-calling edge and OpenAI's superior endpoint reliability and rate limits over raw hyperscaler deployments.
OpenAI o1 and the Obsolescence of Complex Agent Frameworks 7546 Ankur predicts OpenAI o1 will make complex graph-based agent frameworks obsolete. The host pushes back vigorously, emphasizing that enterprise developers refuse to pay for invisible reasoning tokens without granular execution control.
Production Realities of Open Source and Inference Economics 6653 The host questions the low share of open-source models in production. Ankur responds combatively, ridiculing open-source inference provider reliability compared to hyperscalers.
Production Workload Patterns: The Rise of Code-Core Systems 7422 The host synthesizes real-world AI architecture into 'Code Core vs LM Core', matching Ankur's observation that production workloads favor simple imperative code with selective LLM invocations over sprawling autonomous loops.
Database Passions, Foundation Labs, and the Braintrust Team 4311 Ankur shares personal reflections on database engineering nostalgia, his Braintrust founding team, and why he would turn down an acquisition offer to stay focused on high-conviction product building.
Investor Relationships, Series A Milestone, and Hiring Call 3200 The host and guest close by discussing Alana Goyal's hands-on investing methodology, Braintrust's Series A milestone, and active hiring needs in San Francisco.

Statements from this episode (42)

Opinion
Goyal: Snowflake's VARIANT is the best semi-structured data implementation
“It is, without any question, at least in my experience, the best implementation of semi-structured data and sort of solves the problem of storing it very, very efficiently and querying it efficiently, almost as efficiently as if you specified the schema exactl…”
Ankur Goyal Oct 11, 2024 ▶ 6:03
Opinion
Goyal: Observability products would ideally run on Snowflake's VARIANT type
“And I think every observability product in some sort of platonic ideal would be built on top of Snowflake's variant implementation. And have better performance. It would be cheaper. You know, the customer experience would be better. But, you know, alas, it's j…”
Ankur Goyal Oct 11, 2024 ▶ 6:44
Assertion Supported
Goyal: DuckDB struct type lacks true dynamic variant schema flexibility
“DuckDB has a struct type, which is dynamically constructed, but it has all the downsides of traditional structured data types, right? So it's just not Like, for example, if you create, if you infer a bunch of rows with the struct type, and then you present the…”
Ankur Goyal Oct 11, 2024 ▶ 7:30
Assertion Not checkable as stated
Goyal: Selling to business units yielded bigger deals than developer sales
“At Impira, I took kind of the popular advice, which is that developers are a terrible market. So we sold to line of business, and there are a number of benefits to that. Like, we were able to sell six- or seven-figure deals much more easily than We could at Si…”
Ankur Goyal Oct 11, 2024 ▶ 12:35
Insight
Goyal: Founders lacking customer intuition lose to inferior technology
“However, I learned firsthand that if you don't have a very deep, intuitive understanding of your customer, everything becomes harder. Like, you need to throw product managers at the problem. Your own ability to see around corners is much weaker. And, you know,…”
Ankur Goyal Oct 11, 2024 ▶ 12:53
Insight
Goyal: Unstructured data extraction startups remain low-tier enterprise priorities
“It is very, very hard to motivate a large organization to prioritize the problem. And so you're always going to be a second or third tier priority.”
Ankur Goyal Oct 11, 2024 ▶ 15:24
Insight
Goyal: Figma operates like Apple around annual release cycles rather than continuous shipping
“And three, Figma is kind of like Apple, a company that is really optimized around a periodic, like annual release cycle rather than something that's continuous.”
Ankur Goyal Oct 11, 2024 ▶ 25:51
Assertion Supported
Goyal: Figma vector formats consume far more LLM tokens than HTML or JSX
“Vectors are very difficult because they're a data inefficient representation, so the vector format in something like Figma Is choose up like many, many, many, many, many more tokens than HTML and JSX. So it's a very difficult medium to just sort of throw into …”
Ankur Goyal Oct 11, 2024 ▶ 26:58
Insight
Goyal: Designers reject AI designing for them, but AI code generation bridges UI engineering
“In my limited experience and working with designers myself, I think designers do not want AI to design things for them. But there's a lot of things that aren't in the traditional designer toolkit that AI can solve. And I think the biggest one is generating cod…”
Ankur Goyal Oct 11, 2024 ▶ 28:28
Insight
Goyal: Adopting evals resolves engineering stalemates over prompt and model choices
“And I think in the absence of evals, what I saw at Impera, and I see with almost all of our Customers before they start using brain trust is this kind of like stalemate between people on which prompt to use or which model to use or which technique to use that …”
Ankur Goyal Oct 11, 2024 ▶ 31:34
Prediction Not checkable as stated
Goyal: Software engineers will drive AI engineering, but ML tools are unusable for them
“The real gap is that software engineers who have a particular way of thinking, a particular set of biases, a particular type of workflow that they run, are going to be the ones who are doing AI engineering, and that the tools that were built for ML are fantast…”
Ankur Goyal Oct 11, 2024 ▶ 34:05
Insight
Goyal: Continuous evaluation is the foundational workflow for building superior AI software
“Our core belief is that if you embrace evaluation as The sort of core workflow in AI engineering, meaning every time you make a change, you evaluate it, and you use that to drive the next set of changes that you make, then you're able to build much, much bette…”
Ankur Goyal Oct 11, 2024 ▶ 36:21
Insight
Goyal: Matching runtime and eval abstractions eliminates the AI data ETL problem
“If you structure your code so that the same function abstraction that you define to evaluate on equals equals the abstraction that you actually use to run your application, then when you log your application itself, you actually log it in exactly the right for…”
Ankur Goyal Oct 11, 2024 ▶ 37:51
Insight
Goyal contrasts Cursor and Braintrust: AI for software vs software rigor for AI
“Cursor is taking AI and making traditional software engineering like insanely good with AI. And we are taking some of the best things about traditional software engineering and bringing them to building AI software.”
Ankur Goyal Oct 11, 2024 ▶ 40:59
Insight
Goyal: Write prompt evaluations before tweaking prompt text to measure impact
“The idea is like, it's useful to write the eval before you actually like tweak the prompt so that you can measure the impact of the tweak.”
Ankur Goyal Oct 11, 2024 ▶ 44:15
Assertion Not checkable as stated
Goyal: Simple tool-calling prompts cover 80% to 90% of AI use cases
“For probably 80 or 90% of the use cases that we see with people doing this, like very, very simple, I create a prompt, it calls some tools. I can like very ergonomically write the tools, plug into popular services, et cetera, and then just call them kind of li…”
Ankur Goyal Oct 11, 2024 ▶ 50:55
Prediction Not checkable as stated
Goyal: Future of AI engineering centers on reusable tools and tight eval loops
“I think it kind of represents the future of AI engineering, one where You can spend a lot of time writing English and sort of crafting the use case itself. You can reuse tools across different use cases. And then most importantly, the development process is ve…”
Ankur Goyal Oct 11, 2024 ▶ 51:28
Opinion
Goyal: Most AI evaluation tools are merely 'spreadsheet plus plus'
“I would say almost all of the products in the space are spreadsheet plus plus, right? Like, you know, here's a script, generates an eval, I look at the cells, you know, whatever, side by side and compare it.”
Ankur Goyal Oct 11, 2024 ▶ 53:12
Assertion Not checkable as stated
Goyal: Zapier, Coda, and Airtable Required Data to Stay in Cloud
“Zapier was our first user, and then Coda and Airtable quickly followed, and there was just no chance they would be able to use the product unless the data stayed in their cloud.”
Ankur Goyal Oct 11, 2024 ▶ 59:06
Assertion Not checkable as stated
Goyal: Over 75% of Braintrust Eval Users Use TypeScript SDK
“Now I would say every customer and probably north of 75% of the users that are running evals in brain trust are using the TypeScript SDK. It's an overwhelming majority.”
Ankur Goyal Oct 11, 2024 ▶ 1:02:36
Insight
Goyal: Python Dominates AI but TypeScript Dominates Product Building
“At the time, and still, like, AI is sort of at least nominally dominated by Python, but product building is dominated by TypeScript.”
Ankur Goyal Oct 11, 2024 ▶ 1:02:51
Insight
Goyal: Vector Search Hard Part Is Application Permissions, Not Search
“The problem is that the challenge in deploying vector search has very little to do with vector search itself, and much more to do with the data adjacent to vector search. So, for example, if you are at Figma, the Vector search is not actually the hard problem.…”
Ankur Goyal Oct 11, 2024 ▶ 1:15:11
Opinion
Goyal: Vector Search Is Rarely a Storage or Performance Bottleneck
“In almost all cases, vector search is not a storage or performance bottleneck. And in almost all cases, the vector search involves exactly one query, which is, you know, nearest neighbors.”
Ankur Goyal Oct 11, 2024 ▶ 1:16:17
Insight
Goyal: New Databases Succeed Only When Storage and Compilers Rewire Together
“So in my observation, database companies tend to succeed when the storage paradigm Is closely tied to the execution paradigm, and both of those things need to be rewired to work. I think, remember, the databases are not just storage, but they're also compilers…”
Ankur Goyal Oct 11, 2024 ▶ 1:16:44
Insight
Goyal: AI businesses focused solely on fine-tuning are vulnerable to model shifts
“For it to be a business, you need to align with the problem, not the technology. And I think that Automatic optimization is a really great business problem to solve. And I think if you're too fixated on fine tuning as the solution to that problem, then you're …”
Ankur Goyal Oct 11, 2024 ▶ 1:20:37
Assertion Supported
Goyal: In-context learning outperforms fine-tuning in many large-context cases
“There's a lot of cases now, especially with large context models, where in context learning just beats fine tuning.”
Ankur Goyal Oct 11, 2024 ▶ 1:20:55
Assertion Not checkable as stated
Goyal: Fewer Braintrust customers run fine-tuned models in production than six months ago
“I will say in my own experience with customers as of the recording date today, which is September or something, yeah, very few of our customers are currently fine-tuning models. And I think a very, very small fraction of them are running fine-tuned models in p…”
Ankur Goyal Oct 11, 2024 ▶ 1:21:53
Assertion Not checkable as stated
Goyal: Braintrust saw nearly 100% OpenAI market share pre-Claude 3
“Pre-Claude III, it was close to a hundred percent OpenAI.”
Ankur Goyal Oct 11, 2024 ▶ 1:24:59
Assertion Not checkable as stated
Goyal: OpenAI dominates production while Anthropic Sonnet leads side projects
“We still see an overwhelming majority of customers using OpenAI, but almost everyone is using Anthropic for the, and Sonnet specifically for their side projects, whether it's You know, via cursor or prototypes or whatever.”
Ankur Goyal Oct 11, 2024 ▶ 1:26:58
Opinion
Goyal: Running LLM workloads at scale is impractical outside OpenAI
“It's just not practical outside of OpenAI to run use cases at scale in a lot of cases. Like, you can do it, but it requires quite a bit of work. And Because OpenAI is so good at making their models so available, I think they get a lot of credit for the science…”
Ankur Goyal Oct 11, 2024 ▶ 1:27:25
Assertion Not checkable as stated
Goyal: Public clouds fail to match direct OpenAI endpoint experience and capacity
“It hasn't been a smooth journey for people to get the capacity on public clouds that they're able to get through, you know, OpenAI directly. I mean, I think a lot of this is changing, catching up, et cetera. But it hasn't been perfectly smooth. And I think the…”
Ankur Goyal Oct 11, 2024 ▶ 1:28:43
Insight
Goyal: Engineering around LLM limitations guarantees technical obsolescence
“If you make assumptions about the capabilities of models, and you engineer around them, you're almost, like, guaranteed to be screwed.”
Ankur Goyal Oct 11, 2024 ▶ 1:31:45
Prediction Not checkable as stated
Goyal: Agent control flow and graph routing will move into models
“It feels very clear to me that this type of logic is going to be built into the model. Anytime there is control flow complexity or uncertainty complexity, I think the history of AI has been to push more and more into the model.”
Ankur Goyal Oct 11, 2024 ▶ 1:33:10
Prediction Not checkable as stated
Goyal: OpenAI o1 will make agentic frameworks obsolete
“And I think O-one is going to do that to agentic frameworks as well. Hey, I think To me, it seems very unlikely that the, you know, you and me sort of like sipping an espresso and thinking about how, like, different personified roles of people should interact …”
Ankur Goyal Oct 11, 2024 ▶ 1:34:00
Disclosure
Goyal: Under 5% of Braintrust production customers use open source
“Among customers running in production, it's less than five percent.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:00
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:24
Assertion Not checkable as stated
Goyal: GPU inference software companies have high margins and make money
“I don't have any insider information, so I don't know about the hardware companies, but I do know for some of this, excuse me, for some of the software companies, they have high margins and they're making money.”
Ankur Goyal Oct 11, 2024 ▶ 1:38:28
Assertion Not checkable as stated
Goyal: Single-prompt manipulations make up about 50% of Braintrust AI workloads
“I would say about 50% of the use cases that we see are what I would call like single prompt manipulations.”
Ankur Goyal Oct 11, 2024 ▶ 1:39:54
Assertion Not checkable as stated
Goyal: AI workloads are roughly 25% simple agents and 25% advanced agents
“I'd say like probably 25% of the remaining usage is what you could call like a simple agent. Which is probably, you know, a prompt plus some tools. At least one or perhaps the only tool is a rag type of tool, and it is kind of like an enhanced, you know, chatb…”
Ankur Goyal Oct 11, 2024 ▶ 1:41:21
Assertion Not checkable as stated
Goyal: Nearly all Braintrust clients shifted to simple code with LLM calls
“Almost everyone that we work with has gone into this model that, that I, that actually exactly what you said, which is sprinkle intelligence everywhere and make it easy to write dumb code”
Ankur Goyal Oct 11, 2024 ▶ 1:42:27
Insight
Goyal: Systems betting on intrinsic LLM reasoning improvements are more durable
“If you build your system in a way that Kind of assumes LLMs will get better at reasoning and get better at sort of agentic tasks in the LLM itself. Then I think you will build a more durable system.”
Ankur Goyal Oct 11, 2024 ▶ 1:45:53
Assertion Not checkable as stated
Goyal: Observability Giants Built Custom Databases Because Packaged DBs Ignored Variant Types
“My conclusion is that this is a very real problem for a very small number of companies, and that is why Datadog, Splunk, Honeycomb, et cetera, built their own database technology, which is, in some ways, it's sad because all of the technology is a remix of pie…”
Ankur Goyal Oct 11, 2024 ▶ 1:47:01
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.