The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Ankur Goyal no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q conversations, you'd talk to them and they'd say, oh, we don't need this. And then three months later they'd call and say, okay, we really need this. And it was always roughly the same timeframe. Are you seeing any common patterns today in terms of, okay, companies that are now a year or 18 months into their journey using LLMs, like they always have Have the same thing come up?

A There's a couple things. So one is companies that are fairly deep into their journey. They, they have like one or two North star products that are pretty mature and they're trying to figure out how to get those products to the next stage. Probably the most consistent thing I've seen is companies kind of walking back from the, uh, illusion that totally free form agents will solve all of their problems. So I think maybe like two or three months ago, Many of the pioneering companies went way down the agent rabbit hole. Um, and they kind of realized like, wow, this is actually not, this is, this is not, not the right, um, approach. It's so hard to control performance. Um, the error rates are really high and they compound really quickly. Um, and so, you know, most of those companies have kind of walked back and, um, tried to, to, to build a different architecture where the control flow is, is actually managed deterministically by their code. Um, but they, Um, make LLM calls kind of like throughout the entire, uh, architecture of the product. Um, and so that's, that's probably the biggest thing that we're seeing now is, um, I, I don't, I don't know if there's a good term for it yet, but maybe this kind of pervasive, uh, AI engineering throughout a product rather than trying to shove everything into the, you know, um, while loop of an agent.

AI assessment note: “most consistent thing I've seen is companies kind of walking back from the, uh, illusion”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q really surprising. Like people really pushed on, we want this to exist for a long time. We want to be able to pay for it. Um, and so there was that kind of really interesting market pull. Why, why do you think there was so much interest or need for this or demand for it? Or, you know, what does Braintrust do and how does that really impact your customers?

A You know, many of our customers had actually built, uh, early customers had built like internal versions of Braintrust Before we engaged with them. And, uh, there's a couple of things that sort of came out of that. One is it helped them gain an appreciation for how hard the problem is. Evals sound really easy. Oh, it's just a for loop, you know, and then I look at, I console.log the, you know, for loop as I go and I look at the results. Um, but the reality is like, uh, you know, the faster you can eval, the faster you can look at eval results, which start to get really complicated as you start doing things with agents and so on. Um, the faster you can actually iterate and build stuff. It is actually a pretty hard problem to, to do evals well. And many of our early customers, um, who were kind of like the pioneers in, in AI engineering, um, had learned that the hard way. Um, and I think the other problem is that, uh, you know, folks, especially folks, you know, like Brian, for example, they saw that AI would be a pervasive technology throughout the whole org. Not just a project that, you know, Brian might babysit and, and work on with one team. And, um, having a really consistent and, um, standardized, you know, way of doing things was really important. I remember early on, um, Brian pointed me to the Vercel docs, and he said, one of the things I love about this is that, um, whe…

AI assessment note: “it helped them gain an appreciation for how hard the problem is”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q you spend a lot of your career is on sort of databases and data infrastructure and things like that. So you had the BP engineering at single store, which I think, um, was renowned for really having an exceptional, like, um, database centric team. How do you think about the data infrastructure that exists for the AI world today? What's, what's needed? What's lacking? What works well? What doesn't work?

A The, the, the shift is that people have hoarded lots and lots of semi-useful data in data warehouses. Prior to LLMs, uh, there was actually this whole industry around AI where, you know, companies like DataRobot, for example, would come in and help you train models based on these Proprietary structured data that you've collected in your super proprietary data warehouse. And I think the, the, the big insight or the crazy, you know, non-intuitive thing about, um, LLMs is that something trained on the internet outperforms, uh, what an enterprise can produce with their own data trained, um, on data in a data warehouse. And I think Not only, um, is the nature of, like, the data processing problem different, but the value of data is actually, uh, you know, and how we think about the value of data is very, very different. Like, just hoarding data about your, you know, claims history or transaction history, it might not actually be that useful. Um, the real question is, like, how do you, you know, uh, construct a model which is really good at reasoning about the problems that you're working on, and I think the way that enterprises will Um, collect data and leverage it into, you know, these AI processes does not look like doing ETL on a data warehouse that's, you know, running in, in Amazon or something like that. I think, um, it's gonna totally change. And, and I've seen, um, you know,…

AI assessment note: “does not look like doing ETL on a data warehouse that's, you know, running in, in Amazon”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you think that was just like a different problem set in terms of traditional ML and the applications of it are different from what JNI can do, or do you think it was something else?

A Well, I went through this myself, uh, watching the technology that we built to do document extraction at Impura become, you know, totally irrelevant. Um, and, uh, Personally, I, I think it's an emotional thing. Like, you, you try GPT-III for the first time, and first of all, you know, back then at least, um, it was kind of snarky, and so that was a little bit irritating, um, and it, it was also just way better at everything than anything you could possibly train, and I think that is so fundamentally disruptive to, you know, a lot of companies, a lot of people's individual identity, Um, it, it just is not easy to wrap your head around if, uh, you've been doing AI and ML for a while. So I, I think it was largely an emotional thing. You could argue that there's a cost, security, privacy, whatever element of it, but the companies that were sort of on the leading edge, they were able to figure that out pretty quickly. Um, you know, now I think more companies have come along the journey, and I've seen a lot of really smart ML and data science people embrace LLMs and bring a lot of the sort of rigor That is still relevant around evals and measurement and, you know, um, prototyping and so on and become these like AI platform teams. Usually it's a combination of people with product engineering backgrounds and, um, you know, a few folks with statistics or data science backgrounds. Um, an…

AI assessment note: “Personally, I, I think it's an emotional thing.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Have you seen any other, um, shifts in terms of, uh, usage of specific languages or, um, tooling or other things that's happened with this wave of AI?

A Yeah, I think the, the biggest thing I've seen over the past six months is, People dropping the use of frameworks. So early on, I think people thought that AI is this, you know, really unique thing. And just like, you know, Ruby on Rails or whatever, um, we're gonna need to build new kinds of applications, uh, with new kinds of frameworks to be able to, to build AI software. And, uh, really, I think people have walked back from that and they now think of AI as kind of like a core part of their software engineering, um, as a whole. And so, AI is now kind of, like, pervasively spreading throughout people's code base, um, and it's not constrained to what you can create with, you know, a single framework.

AI assessment note: “the biggest thing I've seen over the past six months is, People dropping the use of frameworks.”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Could you tell me a little bit more about what you view as a future of brain trust? How does it evolve as like a product and platform? And then how does it change as AI changes? Is all eval eventually done by machines or, you know, what, what does the future hold for us?

A Yeah. I asked myself that question, you know, every month or so and surprisingly little changes, but, um, you know, brain trust, we started out by solving the eval problem and I think we did that really well. And what we realized is that there's actually this whole platform that people want. One of our customers, uh, actually air table early on, they used our evals product to do observability. So they literally Uh, would create experiments every day as if they were evals and just dump their logs into, um, into those experiments. That's, you know, it's pretty obvious when someone starts doing that, that they're trying to do observability in your product, and we dug into why, and it turns out that in AI, the whole point of observability is to collect data, um, into data sets that you can use to do evals, and then, and then again, eventually fine tune models or, you know, more advanced things. But still, you know, evals is the, is the most important element there. And, and the next thing that happened is that I, you know, some of our customers said, Hey, actually, um, I'm already doing, you know, observability and evals and stuff in Braintrust. I'm spending so much time in this product. Why do I have to go back to my IDE? Which by the way, it knows nothing about my evals. It knows nothing about my logs. Um, can I work on prompts in Braintrust? Can I repro what I'm seeing live? Can…

AI assessment note: “what we realized is that there's actually this whole platform that people want”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.