Krentsel: The Exo Harness Is Running in Production at Braintrust
“The EXO harness and agents built over top of it is running in production at Braintrust.”
Steve Jobs Was Excluded From Pixar's Braintrust Due to His Overwhelming Presence
“Steve had such a powerful voice that it didn't matter when he spoke,
Ed Catmull: He was going to have this extremely strong effect on the dynamics of the room, and he understood that.
Ed Catmull: So, it was the reason he didn't come.
Ed Catmull: I said, you…”
Catmull: Serious creative problems require small groups rather than large meetings
“If things are going well, you like a bigger group because you're training other people to feel comfortable. In this way of working. But the problem is really serious when you go to a small group where they're for past that stuff and they can focus on the probl…”
Bhatawdekar: AI systems require full reasoning traces to evaluate response quality
“These AI systems need to log the entire trace of how the AI reasoned on the initial input. What were the tool calls it made? How did it interact with the LLMs? How did it sort of ultimately generate the response? And did that response actually meet the user in…”
Bhatawdekar: Span-level scorers pinpoint errors in AI agent execution
“You can define those as deterministic functions, you know, implemented in code, or you can use LLM as judges, but then you can evaluate like, how did each span perform? And that can give you a fairly good way to zero in on problematic areas of your agents.”
Ameya Bhatawdekar: Enterprise AI quality comes from surrounding engineering, not just models
“These intelligent systems, these AI systems are not just a model, right? There's a lot of layering that happens on top of these models. These systems have to really deliver specific capabilities or specific experiences that help people do certain specific task…”
Levie: Every enterprise deploying AI agents will require automated evaluation tooling
“And then I'm like, oh, actually everybody on the entire planet, if you're putting agents into an enterprise workflow needs evals. Because you need to know if all of a sudden your agent just stopped producing, you know, like you know, loan origination documents…”
Levie: Every enterprise will adopt AI agent evaluation and observability tools
“I've been fully convinced that the whole agent observability and eval space is going to be a massive space. I'm super excited for what brain trust is doing, excited for, you know, Langsmith, all the things. And I think what you're going to, I mean, this is lik…”
Goyal: Creating Golden Datasets for AI Evals Is Wasted Effort
“People don't really want to create golden data sets. It's, I think it's often a wasted effort to the point that you're making. I think the best teams view offline evals as a mechanism of reconciling what they see in production with real users who are using the…”
Goyal: Braintrust sees surge in PMs and designers joining eval process
“We've seen like a massive surge of product manager, product managers and designers getting interested in participating in the eval process among our customers.”
Sands: Stripe chose Braintrust for AI evaluations over two dozen vendors
“So we actually had like more than two dozen applicants for this evals RFP. Shawn 'Swyx' Wang: There's no way you can evaluate all of them. Emily Glassberg Sands: Well, we actually, we did. So they wrote like nice one pagers. We read them all. We narrowed it do…”
Ambience Healthcare uses Braintrust for on-premise LLM observability
“One of the cool tools that we use internally is Braintrust. I think they're doing some incredible work over there, building like a tool for domain experts. They do give some built in observability. They let you deploy on premise. It's a fantastic technology. H…”
Early AI teams preferred managed evals over brittle open-source tools
“We actually heard early on from people and allergic, no reaction to that. They were like, Hey, we were using open source stuff and it's super brittle and it breaks all the time. And we just evals suck and observability sucks and we just don't want to deal with…”
Braintrust built a terrible go-to-market motion to test product strength
“I wanted at Braintrust to, again, in the spirit of skepticism, Build a terrible go-to-market motion early on, but make the product, make Braintrust only successful if the product was so good that despite being grossly incompetent at selling and marketing our p…”
Intellectually fascinating technical problems efficiently drive peer-to-peer customer acquisition
“I think that people without us having to force them to or ask them to, they are just interested in solving the problem around evals and they find it intellectually fascinating on their own. And I think that leads to a lot of organic knowledge sharing. A lot of…”
Overhyped failure is much better for startups than dying in obscurity
“I think there's only one right answer, which is it's much better to be overhyped and then fail. I mean, obviously within the realms of morality and integrity, but I was, I sort of learned like, I don't want to be an obscure product that no one cares about.”
AI observability logs average 50KB per row compared to 900B traditionally
“Every span in brain trust land, which is like a row of something that you'd log in brain trust, the average size is
50 kilobytes.
In traditional observability, it's 900 bytes.”
Tantivi was the only search library to survive Braintrust's 100x benchmark
“There's one library called Tantivi, which is a Rust re-implementation of the Lucene search index, which is a very popular and well-regarded and used search technology. It's the technology behind Elasticsearch as well. And that was the only piece of code that w…”
Braintrust deliberately excluded traditional enterprises early to focus on product engineers
“We say our product is for product engineering teams that are trying to incorporate AI into their core products and services. And that description excludes and excluded a lot of, ah, sort of traditional enterprises, for example, early on, and we were okay with …”
Braintrust operates with only one company-wide meeting per week
“Like, we don't have, we have one meeting per week as a company, and that's it.”
Goyal cuts off all meetings at noon to protect afternoon focus
“So I cut off meetings at noon every day. I don't take any meetings afternoon.”
Requiring 40-to-50-person startup teams to work weekends is very challenging
“I think personally, my perspective is even if you're trying to hire like a pretty good, well-rounded and, you know, skilled group of people that's even, you know, 40 or 50 people, it's very challenging to get everyone to work on weekends. And so we don't, that…”
Alana Goyal and Elad Gil were the first investors in Braintrust
“Honestly, I think there's two people I'd point to and these were the first two people that invested in BrainTrust and really helped us. The first is Alana”
Pre-PMF startups must hire salespeople accustomed to working without PMF
“A classic example of this is when companies hire salespeople from extremely successful product-led companies, and they expect the salesperson to be good for their product-led company that does not have product market fit. That is not the right You actually wan…”
OpenAI's o3 outperforms GPT-4o on Convex evals by a small margin
“You know, oh, three does do better than four. Oh, I mean, we use brain trust for tracking all this quantitatively, but I can't remember off the top of my head, but it's not like a slam dunk.”
Zhang reveals reciprocal angel investing among Math Olympiad AI startup founders
“I angel invested in a lot of the companies you just listed. A lot of their founders are angel investors in our company.”
Crivello: Lindy will likely switch to Braintrust for AI evaluations
“We're most likely going to switch to Braintrust.”
Goyal: Selling to business units yielded bigger deals than developer sales
“At Impira, I took kind of the popular advice, which is that developers are a terrible market. So we sold to line of business, and there are a number of benefits to that. Like, we were able to sell six- or seven-figure deals much more easily than We could at Si…”
Goyal: Adopting evals resolves engineering stalemates over prompt and model choices
“And I think in the absence of evals, what I saw at Impera, and I see with almost all of our Customers before they start using brain trust is this kind of like stalemate between people on which prompt to use or which model to use or which technique to use that …”
Goyal: Software engineers will drive AI engineering, but ML tools are unusable for them
“The real gap is that software engineers who have a particular way of thinking, a particular set of biases, a particular type of workflow that they run, are going to be the ones who are doing AI engineering, and that the tools that were built for ML are fantast…”
Goyal: Continuous evaluation is the foundational workflow for building superior AI software
“Our core belief is that if you embrace evaluation as The sort of core workflow in AI engineering, meaning every time you make a change, you evaluate it, and you use that to drive the next set of changes that you make, then you're able to build much, much bette…”
Goyal: Matching runtime and eval abstractions eliminates the AI data ETL problem
“If you structure your code so that the same function abstraction that you define to evaluate on equals equals the abstraction that you actually use to run your application, then when you log your application itself, you actually log it in exactly the right for…”
Goyal contrasts Cursor and Braintrust: AI for software vs software rigor for AI
“Cursor is taking AI and making traditional software engineering like insanely good with AI. And we are taking some of the best things about traditional software engineering and bringing them to building AI software.”
Goyal: Write prompt evaluations before tweaking prompt text to measure impact
“The idea is like, it's useful to write the eval before you actually like tweak the prompt so that you can measure the impact of the tweak.”
Goyal: Simple tool-calling prompts cover 80% to 90% of AI use cases
“For probably 80 or 90% of the use cases that we see with people doing this, like very, very simple, I create a prompt, it calls some tools. I can like very ergonomically write the tools, plug into popular services, et cetera, and then just call them kind of li…”
Goyal: Future of AI engineering centers on reusable tools and tight eval loops
“I think it kind of represents the future of AI engineering, one where You can spend a lot of time writing English and sort of crafting the use case itself. You can reuse tools across different use cases. And then most importantly, the development process is ve…”
Goyal: Zapier, Coda, and Airtable Required Data to Stay in Cloud
“Zapier was our first user, and then Coda and Airtable quickly followed, and there was just no chance they would be able to use the product unless the data stayed in their cloud.”
Goyal: Over 75% of Braintrust Eval Users Use TypeScript SDK
“Now I would say every customer and probably north of 75% of the users that are running evals in brain trust are using the TypeScript SDK. It's an overwhelming majority.”
Goyal: Fewer Braintrust customers run fine-tuned models in production than six months ago
“I will say in my own experience with customers as of the recording date today, which is September or something, yeah, very few of our customers are currently fine-tuning models. And I think a very, very small fraction of them are running fine-tuned models in p…”
Goyal: Braintrust saw nearly 100% OpenAI market share pre-Claude 3
“Pre-Claude III, it was close to a hundred percent OpenAI.”
Goyal: Under 5% of Braintrust production customers use open source
“Among customers running in production, it's less than five percent.”
Goyal: Single-prompt manipulations make up about 50% of Braintrust AI workloads
“I would say about 50% of the use cases that we see are what I would call like single prompt manipulations.”
Goyal: AI workloads are roughly 25% simple agents and 25% advanced agents
“I'd say like probably 25% of the remaining usage is what you could call like a simple agent. Which is probably, you know, a prompt plus some tools. At least one or perhaps the only tool is a rag type of tool, and it is kind of like an enhanced, you know, chatb…”
Goyal: Nearly all Braintrust clients shifted to simple code with LLM calls
“Almost everyone that we work with has gone into this model that, that I, that actually exactly what you said, which is sprinkle intelligence everywhere and make it easy to write dumb code”
Gil: Early Enterprise Prospects Asked Braintrust Not to Open Source
“I remember in the early conversations we had around the company or the idea, I should say, it was meant to even potentially be open source. And as the first time that I was involved with some sort of customer call and people would say, we don't want you to ope…”
Goyal: 50% of enterprise AI production use cases involve RAG
“Unambiguously, people are doing rag. So that one is, you know, it's like simple and obvious. Probably around 50% of the use cases that we see in production involve rag of some sort.”
Goyal: Nearly all Braintrust customers have abandoned fine-tuned models
“Almost if not all of our customers have moved off of fine-tuned models onto instruction-tuned models and are seeing really good performance.”
Goyal: Practical adoption of open-source models remains very limited
“So we see very limited practical adoption of open source models, but I think more interest than ever.”
Goyal: Braintrust requires front-end engineering candidates to write C++
“Actually, for example, if you do a front-end interview at Braintrust, one of the questions involves writing some C++, and we lose a lot of candidates because of that question but it's a good signal that maybe Braintrust isn't the right place for you to work.”
Goyal: Pioneering AI companies are abandoning free-form autonomous agents
“Probably the most consistent thing I've seen is companies kind of walking back from the illusion that totally free form agents will solve all of their problems. So I think maybe like two or three months ago, Many of the pioneering companies went way down the a…”
Goyal: Vast majority of Braintrust customers now use TypeScript over Python
“First of all a vast majority of our customers use TypeScript and, you know, early on, some of our customers were dealing with, like, should we use TypeScript or Python? And some teams were using TypeScript, some teams were using Python. Now, almost everyone, i…”
Goyal: Software engineering teams are abandoning specialized AI application frameworks
“The biggest thing I've seen over the past six months is, People dropping the use of frameworks.”
Goyal: Some enterprises consolidated AI stacks to OpenAI, AWS, and Braintrust
“There's some companies that we talked to and their AI vendors are, it's literally OpenAI, AWS, and Braintrust and pretty much everything else has consolidated away.”
Goyal: Braintrust embraces in-office work and an interrupt-driven engineering culture
“Another thing that we're really bullish on at Braintrust is people being in the office and being really comfortable being interrupt-driven.”
Goyal: Braintrust's early growth came from targeting 50 key AI innovators
“Really the thing that we did was we made that list of, like, 50 people who we thought were leading the way in AI and said, you know, let's try to figure out a way to get to these people and either get, recruit them as investors or as customers. And I think tha…”
Goyal: Over half of evaluations run on Braintrust are LLM-based
“I think probably more than half of the evals that people do in Braintrust are LLM based.”
Pixar's Braintrust holds no authority to mandate changes to directors
“The brain trust has no authority. The director does not have to follow any of the specific suggestions given. The brain trust notes are intended to bring the true causes of a problem to the surface, not to demand a specific remedy.”
Catmull: Pixar's Braintrust could suggest ideas but never override directors
“What they were supposed to do was to point out the problems, discuss it, make suggestions, but they couldn't tell people how to solve the problem. They couldn't override the director, and all this frankly was just to get it so that the director was free to lis…”
Catmull: Detaching contributors from their ideas makes creative feedback work
“For me, the magic means that people throw out ideas and they're not attached to them. If they help and they're accepted, good. If they don't help and they're not used, that's okay too, because you're not attached to them. And what we're trying to do is to figu…”
Jackson: BTRST cash price does not affect Braintrust network operations
“The cash value of the brain trust token is actually not relevant to the network, right? It's one token, one vote, the more tokens you have, the more influence you have over the network. So whether that token's worth 10 cents or a dollar is not actually that re…”