The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 25 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Not checkable as stated
Goyal: Nearly all Braintrust clients shifted to simple code with LLM calls
“Almost everyone that we work with has gone into this model that, that I, that actually exactly what you said, which is sprinkle intelligence everywhere and make it easy to write dumb code”
Ankur Goyal Oct 11, 2024 ▶ 1:42:27 Production AI Engineering starts with Evals
Disclosure
Goyal: Under 5% of Braintrust production customers use open source
“Among customers running in production, it's less than five percent.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:00 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Goyal: Software engineers will drive AI engineering, but ML tools are unusable for them
“The real gap is that software engineers who have a particular way of thinking, a particular set of biases, a particular type of workflow that they run, are going to be the ones who are doing AI engineering, and that the tools that were built for ML are fantast…”
Ankur Goyal Oct 11, 2024 ▶ 34:05 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Fewer Braintrust customers run fine-tuned models in production than six months ago
“I will say in my own experience with customers as of the recording date today, which is September or something, yeah, very few of our customers are currently fine-tuning models. And I think a very, very small fraction of them are running fine-tuned models in p…”
Ankur Goyal Oct 11, 2024 ▶ 1:21:53 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Simple tool-calling prompts cover 80% to 90% of AI use cases
“For probably 80 or 90% of the use cases that we see with people doing this, like very, very simple, I create a prompt, it calls some tools. I can like very ergonomically write the tools, plug into popular services, et cetera, and then just call them kind of li…”
Ankur Goyal Oct 11, 2024 ▶ 50:55 Production AI Engineering starts with Evals
Prediction Not checkable as stated
Levie: Every enterprise will adopt AI agent evaluation and observability tools
“I've been fully convinced that the whole agent observability and eval space is going to be a massive space. I'm super excited for what brain trust is doing, excited for, you know, Langsmith, all the things. And I think what you're going to, I mean, this is lik…”
Aaron Levie Mar 5, 2026 ▶ 34:38 Why Every Agent Needs a Box — Aaron Levie, Box
Insight
Goyal: Creating Golden Datasets for AI Evals Is Wasted Effort
“People don't really want to create golden data sets. It's, I think it's often a wasted effort to the point that you're making. I think the best teams view offline evals as a mechanism of reconciling what they see in production with real users who are using the…”
Ankur Goyal Dec 7, 2025 ▶ 10:17 The Great Evals Debate — Ankur Goyal & Malte Ubl
Insight
Goyal: Matching runtime and eval abstractions eliminates the AI data ETL problem
“If you structure your code so that the same function abstraction that you define to evaluate on equals equals the abstraction that you actually use to run your application, then when you log your application itself, you actually log it in exactly the right for…”
Ankur Goyal Oct 11, 2024 ▶ 37:51 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Single-prompt manipulations make up about 50% of Braintrust AI workloads
“I would say about 50% of the use cases that we see are what I would call like single prompt manipulations.”
Ankur Goyal Oct 11, 2024 ▶ 1:39:54 Production AI Engineering starts with Evals
Insight
Goyal contrasts Cursor and Braintrust: AI for software vs software rigor for AI
“Cursor is taking AI and making traditional software engineering like insanely good with AI. And we are taking some of the best things about traditional software engineering and bringing them to building AI software.”
Ankur Goyal Oct 11, 2024 ▶ 40:59 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: AI workloads are roughly 25% simple agents and 25% advanced agents
“I'd say like probably 25% of the remaining usage is what you could call like a simple agent. Which is probably, you know, a prompt plus some tools. At least one or perhaps the only tool is a rag type of tool, and it is kind of like an enhanced, you know, chatb…”
Ankur Goyal Oct 11, 2024 ▶ 1:41:21 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Braintrust saw nearly 100% OpenAI market share pre-Claude 3
“Pre-Claude III, it was close to a hundred percent OpenAI.”
Ankur Goyal Oct 11, 2024 ▶ 1:24:59 Production AI Engineering starts with Evals
Disclosure
Sands: Stripe chose Braintrust for AI evaluations over two dozen vendors
“So we actually had like more than two dozen applicants for this evals RFP. Shawn 'Swyx' Wang: There's no way you can evaluate all of them. Emily Glassberg Sands: Well, we actually, we did. So they wrote like nice one pagers. We read them all. We narrowed it do…”
Emily Glassberg Sands Oct 30, 2025 ▶ 1:13:40 The Agents Economy Backbone - with Emily Glassberg Sands, Head of Data & AI at Stripe
Assertion Not checkable as stated
OpenAI's o3 outperforms GPT-4o on Convex evals by a small margin
“You know, oh, three does do better than four. Oh, I mean, we use brain trust for tracking all this quantitatively, but I can't remember off the top of my head, but it's not like a slam dunk.”
Sujay Jayakar Mar 19, 2025 ▶ 17:56 Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
Prediction Not checkable as stated
Goyal: Future of AI engineering centers on reusable tools and tight eval loops
“I think it kind of represents the future of AI engineering, one where You can spend a lot of time writing English and sort of crafting the use case itself. You can reuse tools across different use cases. And then most importantly, the development process is ve…”
Ankur Goyal Oct 11, 2024 ▶ 51:28 Production AI Engineering starts with Evals
Insight
Goyal: Continuous evaluation is the foundational workflow for building superior AI software
“Our core belief is that if you embrace evaluation as The sort of core workflow in AI engineering, meaning every time you make a change, you evaluate it, and you use that to drive the next set of changes that you make, then you're able to build much, much bette…”
Ankur Goyal Oct 11, 2024 ▶ 36:21 Production AI Engineering starts with Evals
Insight
Goyal: Adopting evals resolves engineering stalemates over prompt and model choices
“And I think in the absence of evals, what I saw at Impera, and I see with almost all of our Customers before they start using brain trust is this kind of like stalemate between people on which prompt to use or which model to use or which technique to use that …”
Ankur Goyal Oct 11, 2024 ▶ 31:34 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Zapier, Coda, and Airtable Required Data to Stay in Cloud
“Zapier was our first user, and then Coda and Airtable quickly followed, and there was just no chance they would be able to use the product unless the data stayed in their cloud.”
Ankur Goyal Oct 11, 2024 ▶ 59:06 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Over 75% of Braintrust Eval Users Use TypeScript SDK
“Now I would say every customer and probably north of 75% of the users that are running evals in brain trust are using the TypeScript SDK. It's an overwhelming majority.”
Ankur Goyal Oct 11, 2024 ▶ 1:02:36 Production AI Engineering starts with Evals
Assertion Not checkable as stated
Goyal: Selling to business units yielded bigger deals than developer sales
“At Impira, I took kind of the popular advice, which is that developers are a terrible market. So we sold to line of business, and there are a number of benefits to that. Like, we were able to sell six- or seven-figure deals much more easily than We could at Si…”
Ankur Goyal Oct 11, 2024 ▶ 12:35 Production AI Engineering starts with Evals
Assertion Supported
Krentsel: The Exo Harness Is Running in Production at Braintrust
“The EXO harness and agents built over top of it is running in production at Braintrust.”
Alex Krentsel Aug 15, 2026 ▶ 34:19 Exo: Harnesses should see their own code and logs — Alex Krentsel, UC Berekeley / Google Research
Assertion Not checkable as stated
Goyal: Braintrust sees surge in PMs and designers joining eval process
“We've seen like a massive surge of product manager, product managers and designers getting interested in participating in the eval process among our customers.”
Ankur Goyal Dec 7, 2025 ▶ 18:24 The Great Evals Debate — Ankur Goyal & Malte Ubl
Disclosure
Ambience Healthcare uses Braintrust for on-premise LLM observability
“One of the cool tools that we use internally is Braintrust. I think they're doing some incredible work over there, building like a tool for domain experts. They do give some built in observability. They let you deploy on premise. It's a fantastic technology. H…”
Brendan Fortuna Jul 29, 2025 ▶ 15:28 ⚡️Using RFT to Build Clinical Superintelligence
Disclosure
Crivello: Lindy will likely switch to Braintrust for AI evaluations
“We're most likely going to switch to Braintrust.”
Florent Crivello Nov 15, 2024 ▶ 28:44 Agents @ Work: Lindy.ai (with live demo!)
Insight
Goyal: Write prompt evaluations before tweaking prompt text to measure impact
“The idea is like, it's useful to write the eval before you actually like tweak the prompt so that you can measure the impact of the tweak.”
Ankur Goyal Oct 11, 2024 ▶ 44:15 Production AI Engineering starts with Evals
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.