The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mitch Trojanowski no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So you guys created that concept of behavior specs. So walk us through what that is. Uh, very practically, is that a marked end file? What, what, what does it look like?

A Yeah. So the, the original idea, actually, actually my co-founder Matt came up with the idea, uh, literally about two years ago, we were talking about auto AGI. Back when we were, Um, uh, even back when we were, had agents that weren't fully, I guess, agentic as you think of them today, and they, you know, had restricted choices, even back then you still wanted to think about, okay, what kind of choice do you want it to make at this, at this fork in the road? And so Matt, we actually used to call it internally meta behaviors, uh, because the idea was that it was a, you're defining, uh, a behavior, but it's at a meta level because it's all the, across all the behaviors agent will have, you know, in all the different trajectories. And, uh, so the idea was that instead of trying to write the prompt, you have to first align on what the meta behavior is. Uh, and so that was actually the first purpose of this concept. I swear to God, literally two years ago. Um, and over time that kind of evolved, um, and we ended up calling it behaviors just cause it's, it's a bit simpler. Uh, and the idea is that you have a, a markdown file in which you write down, how do you want an agent to behave? Simple as that. It could be, it could be at varying degrees of granularity. So let's say you have something that's like very specific, like, um, you need to go to like, look at the primary sources. May…

AI assessment note: “you have a, a markdown file in which you write down, how do you want an agent to behave?”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And did you pick accounting because of how interesting that was from a agent building perspective or the other way around?

A It's probably the other way around, but I do think it is actually quite interesting from an agent perspective. Accounting is interesting for a lot of reasons. It is, one, one of, if not the largest knowledge work profession in the country. There are over three million, you know, combined kind of accountants in the country. And what I think is so cool about accounting, actually, is that most people don't really think about accounting. They don't think like, well, yeah, why is that even there? You know, probably most listeners have never thought, like, why does it even exist? And, um, I know we're gonna talk about agents, maybe quickly, 30 seconds, just to convince everyone how cool accounting is. Uh, if you think about the real world, uh, so much stuff happens. Economic activity, right? Like, you know, I was just drinking a water bottle there, like, you know, someone, um, uh, that bottler had to choose to, like, go, uh, uh, buy from that factory or that supplier, um, or decide to open, you know, some additional, uh, store or hire a salesperson. And these are all economic decisions that stem from understanding The real world. What's in the real world? You know, money moves hands, someone signs a contract, someone delivers the inventory. It's like all these events that occur. And so much of modern capitalism relies on the ability of all these actors to make decisions on these even…

AI assessment note: “It's probably the other way around, but I do think it is actually quite interesting”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So do you think of, um, doing your own RL as some kind of moat against being swallowed up by model performance? I mean, as, as you think about applied AI companies of the future, will, will they all be RL labs of some sort?

A So generally, and if anyone's currently trying to found a company, I recommend thinking this way. Technical moats are not real moats. Like, I, I, there's no portion of basis's long-term terminal value that stems from some, you know, secret RL trick we found that nobody else found. So that doesn't matter. Um, uh, what matters is that right now we are obviously very good at building long horizon agents, um, that, you know, can be reliable and deployed in production, and we'll continue to be the best at that. And that allows us to win market share and get deeply embedded. And that's why we move really fast, because the work we're trying to do is to, um, uh, go and proliferate maybe before, like, you get to AGI, whatever you want to call that. So I think that Most of the moats that will exist will be business moats. That is true in the AGI era. I would argue that's also been true in the pre-AGI era. Like, I don't think that, you know, Salesforce can write a better SQL query than I can. Like, the mode that Salesforce has is not related to their technology, right? It's related to their business position. You know, it's the, it's the powers. It's the, it's the, it's the, the workflows that they own. It's, it's all the, it's so many of these different things that, uh, come with being, you know, embedded in, and that's what matters, uh, not the technology. The technology is a, Is a, uh,…

AI assessment note: “Technical moats are not real moats. Like, I, I, there's no portion of basis's”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q What does that even mean, designing an ontology? An ontology practically is, what is this, it's a graph database, it's a series of relationships?

A Yeah, it could be a lot of different, it could be, there are different formats. I think the, the simplest way to think about it is honestly just a file system, um, uh, in which, you know, you have like some structure. Obviously, most people, you do virtual file systems, and so you have a lot of flexibility there, so there could be other types of metadata associated with the files and the folders, right? You, you, you could, there could be, you know, uh, connectors, uh, in nodes and like some GraphDB if you wanted. Obviously there could also be embeddings. Like there's so many different, you know, sources of, of data that you can get to like help enrich. Um, uh, and now, by the way, as models are getting cheaper and cheaper, more and more of that actually can just be done, um, using inference instead of using determined, like things like graphs or things like embeddings. You know, if you have, if Luna costs is free, then suddenly you could run Lunas across your entire ontology and summarize stuff for the agent up or things like that. Um, and so that, that's kind of how I would think about it. And maybe also one more piece. It's not just the, the ontology doesn't just mean the structure of the folders and ontology traditionally, it's the, it's also like the language, right? Like what are the objects and the concepts? Because at the end of the day, if you're being trained at runti…

AI assessment note: “the simplest way to think about it is honestly just a file system”

Answered raw tape D 5 · C 3 · P 3 · Cm 3 3.60

Q Okay. So, uh, why do agents struggle? Uh, you mentioned three reasons, maybe, maybe, maybe you mentioned what those reasons are, and then we'll go into them turn by turn.

A Yeah, so I think agents, um, uh, struggle for a lot of reasons. Um, I think, one, they struggle because, uh, they don't necessarily know what good looks like. Um, I think they struggle because it may not be easy to verify yourself, uh, at, at runtime. Um, as we were talking about with coding, um, I think another part is that, uh, And this is maybe not an agent struggle, but maybe it's a UX thing, which is that for coding engineers are just very in the weeds of it. And so there's a kind of a difference where if you were like running a long horizon agent for coding, if the engineer was not, you know, engineering and like in the weeds of the code, if instead they were more abstracted away, your maybe level of, uh, but quality and, you know, how you make decisions probably, you maybe need a, need a higher bar than you would for coding. You know, coders are okay with lower bars. That's been true forever. Um, and so I think all these things add up, um, in, in, in making it, and even now with coding, like, the agents are not yet, uh, they're not, they're not, they're not human level at being coherent over long periods of time. That's obvious because they can't code like a junior engineer on a project for two weeks. So they can't, they're, that's worse than a human, uh, even to start, so.

AI assessment note: “one, they struggle because, uh, they don't necessarily know what good looks like.”

Redirected raw tape D 2 · C 3 · P 2 · Cm 2 2.30

Q uh, guessing and inferring from how the model behaves through like artifacts and judging from tool calls and how does that work? Um, Uh, and, um, and then maybe walk us through, uh, each time a model changes or the next version of the model, you know, gets released. Like, do you have to then look at, um, everything that you've been doing in the light of that new model?

A Essentially maybe it goes back to the, like, you know, Opus three, O one, O three, because I, I do think one thing that's really important when you're building agents, but it's definitely building a company around it is, You should, you shouldn't be that surprised. Um, you should have a model of the world and as things change, you should update your model. Um, but, uh, to be successful, you know, you can't just update your model all the time. You need to be right a little bit. Um, and I, I do think that, uh, If you really internalize some concepts about this, right? That like, you now have this, forgetting about the internals for a second, forget about this LM, you have this magic box and you have this, this magic box or this alien, I like to call it sometimes, and you could send in huge amounts of data into this alien and it will be able to, uh, uh, reason and learn at inference time within that magic box, and then come back to you with output that, Uh, you know, now that tool calling obviously works and whatnot, you can plug into the rest of the system. That's kind of all you really need to know, and I think once you really appreciate what that means, and you take it to its logical conclusion, a lot of stuff starts to fall out of that. Because you start to understand, it's like, okay, well wait, like, if I have this magic box that can do this, does that mean that it could dec…

AI assessment note: “forgetting about the internals for a second, forget about this LM, you have this magic box”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.