Oct 8, 2024 · 38m · no-priors

No Priors Ep. 85 | CEO of Braintrust Ankur Goyal

Ankur Goyal · 28m spoken Elad Gil · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, host Elad Gil interviews Ankur Goyal, founder and CEO of Braintrust, to explore the evolution of enterprise AI evaluation, data infrastructure shifts, and production engineering practices. Goyal explains why enterprises are converging on deterministic control flows, TypeScript, and proprietary frontier models while outlining Braintrust's journey toward becoming a universal platform for AI development.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.2% of the talking time here. How this is scored →

The hosts as informed peer 4.4 Guest teaching 5.3 Guest disagreement 1.7 The hosts pushing back 1.0
05100:0010:0020:0030:000:45–5:44 · The hosts as informed peer 5/10 Genesis of Braintrust: Lessons from Impira to Figma Elad shares his early personal observations from customer calls with Braintrust and framing why customers demanded commercial software over open source. Ankur explains why evals are deceptive and how early pioneers realized it was harder than a simple for loop.5:45–8:57 · The hosts as informed peer 4/10 Enterprise Adoption Trends: RAG vs. Fine-Tuning Ankur reframes the debate around fine-tuning, explaining that fine-tuning is merely a technique while automatic optimization is the real outcome, and points out that nearly all production customers have migrated back to instruction tuning.8:57–12:43 · The hosts as informed peer 5/10 Open Source vs. Proprietary Model Economics in Production Elad contextually grounds Ankur's insights as reflecting live production systems at scale rather than hobbyist benchmarks. Ankur details the economic realities of proprietary APIs beating open-source self-hosting on developer velocity and ROI.12:43–17:19 · The hosts as informed peer 4/10 Embeddings, Semantic Search, and Analytical Workload Disruption Drawing on his database background from SingleStore, Ankur delivers deep insights dismissing the standalone vector database debate as trivial while arguing that semantic search and embeddings will completely disrupt traditional OLAP warehouse workloads.17:19–19:28 · The hosts as informed peer 5/10 Architectural Realism: Moving from Autonomous Agents to Deterministic Code Elad outlines the typical journey enterprise startups make before reaching maturity. Ankur validates this by breaking down how pioneering companies walked back from autonomous agent while-loops in favor of deterministic code with LLM calls throughout.19:28–25:12 · The hosts as informed peer 5/10 Evolution of AI Teams and the Dominance of TypeScript Elad probes into why traditional ML engineers resisted LLMs, and Ankur offers an honest psychological explanation before making a bold contrarian case for TypeScript dominating AI over Python due to strong typing and product engineering dominance.25:12–30:14 · The hosts as informed peer 5/10 Tooling Shifts: Framework Abandonment and Hyperscaler Consolidation Ankur points out how framework usage has been abandoned in favor of pervasive native software engineering alongside massive vendor consolidation around AWS and OpenAI. Elad draws comparisons to other coding CEOs like Tobi Lutke and Jensen Huang.30:15–33:07 · The hosts as informed peer 4/10 Operational Edge: In-Person Work, Fast Feedback Loops, and GTM Strategy Ankur discusses their strict in-office requirement, interrupt-driven fast iteration speed, and high-touch GTM strategy. Elad adds his perspective on how startups commonly misdefine their initial customer envelopes.33:07–38:04 · The hosts as informed peer 3/10 Future of Braintrust: Universal Developer Platform and Automated Evals Ankur explains how customer pull transformed Braintrust from pure evals into an observability and prompt engineering platform, and outlines how LLMs are increasingly used to evaluate logs containing private PII.0:45–5:44 · Guest teaching 4/10 Genesis of Braintrust: Lessons from Impira to Figma Elad shares his early personal observations from customer calls with Braintrust and framing why customers demanded commercial software over open source. Ankur explains why evals are deceptive and how early pioneers realized it was harder than a simple for loop.5:45–8:57 · Guest teaching 6/10 Enterprise Adoption Trends: RAG vs. Fine-Tuning Ankur reframes the debate around fine-tuning, explaining that fine-tuning is merely a technique while automatic optimization is the real outcome, and points out that nearly all production customers have migrated back to instruction tuning.8:57–12:43 · Guest teaching 5/10 Open Source vs. Proprietary Model Economics in Production Elad contextually grounds Ankur's insights as reflecting live production systems at scale rather than hobbyist benchmarks. Ankur details the economic realities of proprietary APIs beating open-source self-hosting on developer velocity and ROI.12:43–17:19 · Guest teaching 7/10 Embeddings, Semantic Search, and Analytical Workload Disruption Drawing on his database background from SingleStore, Ankur delivers deep insights dismissing the standalone vector database debate as trivial while arguing that semantic search and embeddings will completely disrupt traditional OLAP warehouse workloads.17:19–19:28 · Guest teaching 5/10 Architectural Realism: Moving from Autonomous Agents to Deterministic Code Elad outlines the typical journey enterprise startups make before reaching maturity. Ankur validates this by breaking down how pioneering companies walked back from autonomous agent while-loops in favor of deterministic code with LLM calls throughout.19:28–25:12 · Guest teaching 6/10 Evolution of AI Teams and the Dominance of TypeScript Elad probes into why traditional ML engineers resisted LLMs, and Ankur offers an honest psychological explanation before making a bold contrarian case for TypeScript dominating AI over Python due to strong typing and product engineering dominance.25:12–30:14 · Guest teaching 5/10 Tooling Shifts: Framework Abandonment and Hyperscaler Consolidation Ankur points out how framework usage has been abandoned in favor of pervasive native software engineering alongside massive vendor consolidation around AWS and OpenAI. Elad draws comparisons to other coding CEOs like Tobi Lutke and Jensen Huang.30:15–33:07 · Guest teaching 4/10 Operational Edge: In-Person Work, Fast Feedback Loops, and GTM Strategy Ankur discusses their strict in-office requirement, interrupt-driven fast iteration speed, and high-touch GTM strategy. Elad adds his perspective on how startups commonly misdefine their initial customer envelopes.33:07–38:04 · Guest teaching 6/10 Future of Braintrust: Universal Developer Platform and Automated Evals Ankur explains how customer pull transformed Braintrust from pure evals into an observability and prompt engineering platform, and outlines how LLMs are increasingly used to evaluate logs containing private PII.0:45–5:44 · Guest disagreement 1/10 Genesis of Braintrust: Lessons from Impira to Figma Elad shares his early personal observations from customer calls with Braintrust and framing why customers demanded commercial software over open source. Ankur explains why evals are deceptive and how early pioneers realized it was harder than a simple for loop.5:45–8:57 · Guest disagreement 2/10 Enterprise Adoption Trends: RAG vs. Fine-Tuning Ankur reframes the debate around fine-tuning, explaining that fine-tuning is merely a technique while automatic optimization is the real outcome, and points out that nearly all production customers have migrated back to instruction tuning.8:57–12:43 · Guest disagreement 2/10 Open Source vs. Proprietary Model Economics in Production Elad contextually grounds Ankur's insights as reflecting live production systems at scale rather than hobbyist benchmarks. Ankur details the economic realities of proprietary APIs beating open-source self-hosting on developer velocity and ROI.12:43–17:19 · Guest disagreement 2/10 Embeddings, Semantic Search, and Analytical Workload Disruption Drawing on his database background from SingleStore, Ankur delivers deep insights dismissing the standalone vector database debate as trivial while arguing that semantic search and embeddings will completely disrupt traditional OLAP warehouse workloads.17:19–19:28 · Guest disagreement 2/10 Architectural Realism: Moving from Autonomous Agents to Deterministic Code Elad outlines the typical journey enterprise startups make before reaching maturity. Ankur validates this by breaking down how pioneering companies walked back from autonomous agent while-loops in favor of deterministic code with LLM calls throughout.19:28–25:12 · Guest disagreement 3/10 Evolution of AI Teams and the Dominance of TypeScript Elad probes into why traditional ML engineers resisted LLMs, and Ankur offers an honest psychological explanation before making a bold contrarian case for TypeScript dominating AI over Python due to strong typing and product engineering dominance.25:12–30:14 · Guest disagreement 1/10 Tooling Shifts: Framework Abandonment and Hyperscaler Consolidation Ankur points out how framework usage has been abandoned in favor of pervasive native software engineering alongside massive vendor consolidation around AWS and OpenAI. Elad draws comparisons to other coding CEOs like Tobi Lutke and Jensen Huang.30:15–33:07 · Guest disagreement 1/10 Operational Edge: In-Person Work, Fast Feedback Loops, and GTM Strategy Ankur discusses their strict in-office requirement, interrupt-driven fast iteration speed, and high-touch GTM strategy. Elad adds his perspective on how startups commonly misdefine their initial customer envelopes.33:07–38:04 · Guest disagreement 1/10 Future of Braintrust: Universal Developer Platform and Automated Evals Ankur explains how customer pull transformed Braintrust from pure evals into an observability and prompt engineering platform, and outlines how LLMs are increasingly used to evaluate logs containing private PII.0:45–5:44 · The hosts pushing back 1/10 Genesis of Braintrust: Lessons from Impira to Figma Elad shares his early personal observations from customer calls with Braintrust and framing why customers demanded commercial software over open source. Ankur explains why evals are deceptive and how early pioneers realized it was harder than a simple for loop.5:45–8:57 · The hosts pushing back 1/10 Enterprise Adoption Trends: RAG vs. Fine-Tuning Ankur reframes the debate around fine-tuning, explaining that fine-tuning is merely a technique while automatic optimization is the real outcome, and points out that nearly all production customers have migrated back to instruction tuning.8:57–12:43 · The hosts pushing back 1/10 Open Source vs. Proprietary Model Economics in Production Elad contextually grounds Ankur's insights as reflecting live production systems at scale rather than hobbyist benchmarks. Ankur details the economic realities of proprietary APIs beating open-source self-hosting on developer velocity and ROI.12:43–17:19 · The hosts pushing back 1/10 Embeddings, Semantic Search, and Analytical Workload Disruption Drawing on his database background from SingleStore, Ankur delivers deep insights dismissing the standalone vector database debate as trivial while arguing that semantic search and embeddings will completely disrupt traditional OLAP warehouse workloads.17:19–19:28 · The hosts pushing back 1/10 Architectural Realism: Moving from Autonomous Agents to Deterministic Code Elad outlines the typical journey enterprise startups make before reaching maturity. Ankur validates this by breaking down how pioneering companies walked back from autonomous agent while-loops in favor of deterministic code with LLM calls throughout.19:28–25:12 · The hosts pushing back 1/10 Evolution of AI Teams and the Dominance of TypeScript Elad probes into why traditional ML engineers resisted LLMs, and Ankur offers an honest psychological explanation before making a bold contrarian case for TypeScript dominating AI over Python due to strong typing and product engineering dominance.25:12–30:14 · The hosts pushing back 1/10 Tooling Shifts: Framework Abandonment and Hyperscaler Consolidation Ankur points out how framework usage has been abandoned in favor of pervasive native software engineering alongside massive vendor consolidation around AWS and OpenAI. Elad draws comparisons to other coding CEOs like Tobi Lutke and Jensen Huang.30:15–33:07 · The hosts pushing back 1/10 Operational Edge: In-Person Work, Fast Feedback Loops, and GTM Strategy Ankur discusses their strict in-office requirement, interrupt-driven fast iteration speed, and high-touch GTM strategy. Elad adds his perspective on how startups commonly misdefine their initial customer envelopes.33:07–38:04 · The hosts pushing back 1/10 Future of Braintrust: Universal Developer Platform and Automated Evals Ankur explains how customer pull transformed Braintrust from pure evals into an observability and prompt engineering platform, and outlines how LLMs are increasingly used to evaluate logs containing private PII.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 31.7% · guest 68.3%0:00 · the hosts 31.7% · guest 68.3%3:00 · the hosts 27.3% · guest 72.7%3:00 · the hosts 27.3% · guest 72.7%6:00 · the hosts 14.2% · guest 85.8%6:00 · the hosts 14.2% · guest 85.8%9:00 · the hosts 29.2% · guest 70.8%9:00 · the hosts 29.2% · guest 70.8%12:00 · the hosts 21.3% · guest 78.7%12:00 · the hosts 21.3% · guest 78.7%15:00 · the hosts 23.2% · guest 76.8%15:00 · the hosts 23.2% · guest 76.8%18:00 · the hosts 25.1% · guest 74.9%18:00 · the hosts 25.1% · guest 74.9%21:00 · the hosts 10.7% · guest 89.3%21:00 · the hosts 10.7% · guest 89.3%24:00 · the hosts 11.6% · guest 88.4%24:00 · the hosts 11.6% · guest 88.4%27:00 · the hosts 24.4% · guest 75.6%27:00 · the hosts 24.4% · guest 75.6%30:00 · the hosts 25.6% · guest 74.4%30:00 · the hosts 25.6% · guest 74.4%33:00 · the hosts 11.1% · guest 88.9%33:00 · the hosts 11.1% · guest 88.9%36:00 · the hosts 4.7% · guest 95.3%36:00 · the hosts 4.7% · guest 95.3%
Sharpest disagreement ▶ 24:05 TypeScript superior to Python for AI systems

Ankur takes a firm contrarian stance against Python advocates, highlighting that despite getting trolled on Twitter, TypeScript is objectively better suited for uncertain AI data shapes.

Hardest push from the hosts ▶ 20:33 Elad probes ML resistance motives

Elad challenges whether ML practitioner skepticism of LLMs was really due to technical problem-set mismatches rather than general dismissal, prompting Ankur to explain the emotional disruption.

Biggest teaching moment ▶ 14:00 Debunking vector DB hype and explaining OLAP disruption

Ankur dismisses the popular discourse around vector databases as silly and educates the audience on why semantic search fundamentally breaks traditional relational OLAP data warehouse architectures.

The host holds their own ▶ 29:35 Elad connects Jensen Huang's CEO architecture to startup execution

Elad synthesizes Jensen Huang's management thesis on structuring companies around founder strengths with standard operational traps like reinventing sales compensation.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Genesis of Braintrust: Lessons from Impira to Figma 5411 Elad shares his early personal observations from customer calls with Braintrust and framing why customers demanded commercial software over open source. Ankur explains why evals are deceptive and how early pioneers realized it was harder than a simple for loop.
Enterprise Adoption Trends: RAG vs. Fine-Tuning 4621 Ankur reframes the debate around fine-tuning, explaining that fine-tuning is merely a technique while automatic optimization is the real outcome, and points out that nearly all production customers have migrated back to instruction tuning.
Open Source vs. Proprietary Model Economics in Production 5521 Elad contextually grounds Ankur's insights as reflecting live production systems at scale rather than hobbyist benchmarks. Ankur details the economic realities of proprietary APIs beating open-source self-hosting on developer velocity and ROI.
Embeddings, Semantic Search, and Analytical Workload Disruption 4721 Drawing on his database background from SingleStore, Ankur delivers deep insights dismissing the standalone vector database debate as trivial while arguing that semantic search and embeddings will completely disrupt traditional OLAP warehouse workloads.
Architectural Realism: Moving from Autonomous Agents to Deterministic Code 5521 Elad outlines the typical journey enterprise startups make before reaching maturity. Ankur validates this by breaking down how pioneering companies walked back from autonomous agent while-loops in favor of deterministic code with LLM calls throughout.
Evolution of AI Teams and the Dominance of TypeScript 5631 Elad probes into why traditional ML engineers resisted LLMs, and Ankur offers an honest psychological explanation before making a bold contrarian case for TypeScript dominating AI over Python due to strong typing and product engineering dominance.
Tooling Shifts: Framework Abandonment and Hyperscaler Consolidation 5511 Ankur points out how framework usage has been abandoned in favor of pervasive native software engineering alongside massive vendor consolidation around AWS and OpenAI. Elad draws comparisons to other coding CEOs like Tobi Lutke and Jensen Huang.
Operational Edge: In-Person Work, Fast Feedback Loops, and GTM Strategy 4411 Ankur discusses their strict in-office requirement, interrupt-driven fast iteration speed, and high-touch GTM strategy. Elad adds his perspective on how startups commonly misdefine their initial customer envelopes.
Future of Braintrust: Universal Developer Platform and Automated Evals 3611 Ankur explains how customer pull transformed Braintrust from pure evals into an observability and prompt engineering platform, and outlines how LLMs are increasingly used to evaluate logs containing private PII.

Statements from this episode (24)

Assertion Not checkable as stated
Gil: Early Enterprise Prospects Asked Braintrust Not to Open Source
“I remember in the early conversations we had around the company or the idea, I should say, it was meant to even potentially be open source. And as the first time that I was involved with some sort of customer call and people would say, we don't want you to ope…”
Elad Gil Oct 8, 2024 ▶ 2:43
Assertion Not checkable as stated
Goyal: 50% of enterprise AI production use cases involve RAG
“Unambiguously, people are doing rag. So that one is, you know, it's like simple and obvious. Probably around 50% of the use cases that we see in production involve rag of some sort.”
Ankur Goyal Oct 8, 2024 ▶ 6:17
Assertion Not checkable as stated
Goyal: Nearly all Braintrust customers have abandoned fine-tuned models
“Almost if not all of our customers have moved off of fine-tuned models onto instruction-tuned models and are seeing really good performance.”
Ankur Goyal Oct 8, 2024 ▶ 7:30
Assertion Not checkable as stated
Goyal: Anthropic's Claude 3.5 Sonnet has really taken off
“Especially, you know, Claude III-V Sonnet has really taken off.”
Ankur Goyal Oct 8, 2024 ▶ 9:09
Assertion Not checkable as stated
Goyal: Practical adoption of open-source models remains very limited
“So we see very limited practical adoption of open source models, but I think more interest than ever.”
Ankur Goyal Oct 8, 2024 ▶ 9:23
Insight
Goyal: Open-source models will struggle until they improve UX or iteration speed
“So I think until open source can really move the needle on one of those two axes, it's going to be tough for it to be adopted broadly.”
Ankur Goyal Oct 8, 2024 ▶ 10:32
Insight
Goyal: Internet-trained LLMs outperform models trained on internal enterprise data
“And I think the big insight or the crazy, you know, non-intuitive thing about LLMs is that something trained on the internet outperforms what an enterprise can produce with their own data trained on data in a data warehouse.”
Ankur Goyal Oct 8, 2024 ▶ 11:25
Prediction Not checkable as stated
Goyal: Enterprise AI data infrastructure will move away from data warehouse ETL
“And I think the way that enterprises will collect data and leverage it into, you know, these AI processes does not look like doing ETL on a data warehouse that's, you know, running in, in Amazon or something like that. I think it's gonna totally change.”
Ankur Goyal Oct 8, 2024 ▶ 12:03
Prediction Not checkable as stated
Goyal: Embeddings and LLMs will replace relational indexes for querying data
“LLMs and, you know, specifically embeddings are going to be core to how people actually query data, not, you know, traditional algebraic relational indexes.”
Ankur Goyal Oct 8, 2024 ▶ 14:02
Assertion Supported
Goyal: Relational databases are fully capable of adding HNSW vector indices
“Relational databases are perfectly capable of adding HNSW indices to them.”
Ankur Goyal Oct 8, 2024 ▶ 14:22
Prediction Not checkable as stated
Goyal: AI semantic search will disrupt OLAP far more than OLTP
“What will really be disrupted is the OLAP workload. So relational, you can't just slap you know, semantic search and stuff into the architecture of a traditional data warehouse. I think that is actually a much deeper set of things that will need to change than…”
Ankur Goyal Oct 8, 2024 ▶ 14:28
Disclosure
Goyal: Braintrust requires front-end engineering candidates to write C++
“Actually, for example, if you do a front-end interview at Braintrust, one of the questions involves writing some C++, and we lose a lot of candidates because of that question but it's a good signal that maybe Braintrust isn't the right place for you to work.”
Ankur Goyal Oct 8, 2024 ▶ 15:39
Insight
Goyal: Pioneering AI companies are abandoning free-form autonomous agents
“Probably the most consistent thing I've seen is companies kind of walking back from the illusion that totally free form agents will solve all of their problems. So I think maybe like two or three months ago, Many of the pioneering companies went way down the a…”
Ankur Goyal Oct 8, 2024 ▶ 18:23
Assertion Not checkable as stated
Goyal: Impira's document extraction tech became totally irrelevant with LLMs
“Well, I went through this myself watching the technology that we built to do document extraction at Impura become, you know, totally irrelevant.”
Ankur Goyal Oct 8, 2024 ▶ 20:41
Disclosure
Goyal: Vast majority of Braintrust customers now use TypeScript over Python
“First of all a vast majority of our customers use TypeScript and, you know, early on, some of our customers were dealing with, like, should we use TypeScript or Python? And some teams were using TypeScript, some teams were using Python. Now, almost everyone, i…”
Ankur Goyal Oct 8, 2024 ▶ 23:43
Opinion
Goyal: TypeScript's type system makes it better suited for AI than Python
“Another thing is that TypeScript as a language is inherently better suited for AI workloads because of the type system. So the type system basically allows you to launder, you know, the crazy stuff that comes out of an AI model into a well-defined structure th…”
Ankur Goyal Oct 8, 2024 ▶ 24:25
Insight
Goyal: Software engineering teams are abandoning specialized AI application frameworks
“The biggest thing I've seen over the past six months is, People dropping the use of frameworks.”
Ankur Goyal Oct 8, 2024 ▶ 25:22
Opinion
Goyal: AWS regained its mojo by hosting Anthropic Claude on Bedrock
“AWS has its mojo back now that they have Anthropic on bedrock and Anthropic is, you know, especially cloud three and three, five are really, really good.”
Ankur Goyal Oct 8, 2024 ▶ 26:24
Assertion Not checkable as stated
Goyal: Some enterprises consolidated AI stacks to OpenAI, AWS, and Braintrust
“There's some companies that we talked to and their AI vendors are, it's literally OpenAI, AWS, and Braintrust and pretty much everything else has consolidated away.”
Ankur Goyal Oct 8, 2024 ▶ 26:51
Disclosure
Goyal: Braintrust embraces in-office work and an interrupt-driven engineering culture
“Another thing that we're really bullish on at Braintrust is people being in the office and being really comfortable being interrupt-driven.”
Ankur Goyal Oct 8, 2024 ▶ 30:22
Disclosure
Goyal: Braintrust's early growth came from targeting 50 key AI innovators
“Really the thing that we did was we made that list of, like, 50 people who we thought were leading the way in AI and said, you know, let's try to figure out a way to get to these people and either get, recruit them as investors or as customers. And I think tha…”
Ankur Goyal Oct 8, 2024 ▶ 31:51
Insight
Goyal: AI observability exists to collect datasets for evaluations and fine-tuning
“In AI, the whole point of observability is to collect data into data sets that you can use to do evals, and then again, eventually fine tune models or, you know, more advanced things.”
Ankur Goyal Oct 8, 2024 ▶ 33:50
Insight
Goyal: Validating LLM outputs is significantly easier for frontier models than generation
“It's way easier for an LLM, especially a frontier model, to look at the work of you know, itself or another LLM and accurately assess it.”
Ankur Goyal Oct 8, 2024 ▶ 36:58
Assertion Not checkable as stated
Goyal: Over half of evaluations run on Braintrust are LLM-based
“I think probably more than half of the evals that people do in Braintrust are LLM based.”
Ankur Goyal Oct 8, 2024 ▶ 37:11
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.