Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q How do you make the other engineers not feel like you're not special? I think that's something that I hear a lot. It's like, hey, you know, why aren't these people working on all the cool LLM things? And like, I'm stuck working on, you know, the KYC integration with whatever. You know what I mean? It's like, how do you build that culture?
A You know, it's interesting. I, I thought that that would be more of a problem, but the benefit of having really optimized Our engineering culture around business impact actually causes it to cut in the other direction where for folks, some folks don't want to work on the AI products because it doesn't have as much clear direct like business impact right now. Doesn't, doesn't impact revenues directly. And so I, uh, I think folks for the most part, uh, we've, we've enabled folks who have a strong desire to work on, on, um, AI products to, to join that team. Like somebody, somebody transferred out of our expense management organization to come over there because they're really passionate about Taking like their knowledge of like policy evaluation and, and bringing it into the, the AI, uh, uh, team. But the most part, I think everybody understands like how their work, uh, ladders up and maybe there's some like friendly rivalry because like the folks who say we're kind of a card product, they, they drive 60% of our direct revenue. And so they, they're pretty happy with that. And, uh, and they don't feel like they're being left out. Uh, and I will also say, um, as you probably saw in this, this piece that we, we, uh, put out with, uh, first round. There is a lot of smaller applications of LLMs peppered throughout all of our product and operations teams. It's just some of the more nov…
AI assessment note: “the benefit of having really optimized Our engineering culture around business impact”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So do you want to expand on that? Yeah.
A And this is an area where we've only scratched the surface here, but, but a big, a big challenge that, that we face is that the world knowledge or the knowledge that's built into the model about, uh, about, you know, what, GPT-V thinks Brex does and how it thinks our business operates is actually quite different from what our business offers today or how our product works. And so we've had to, to work on building a corpus of sort of product documentation, process documentation, and like curate this set of information to basically ground a variety of our LLM applications, including like that Brex assistant, which is like the You know, the assistant that employees, uh, will, will talk to is like, we don't want it to, to hallucinate features that we don't have, or like give, give wrong information there. And similarly, like, uh, some of the operational, um, uh, agents need to be grounded on, um, like what our ICP is, because if you ask, uh, you know, ChatGPT five right now, like what types of businesses does Brex, uh, onboard or like what types of businesses does Brex serve? It might not give an accurate Uh, explanation to that, to that question. It might, it might say we're a corporate car for startups, which is what we did, you know, seven years ago. And it might say we're only, we only serve enterprises. And so that has been an interesting challenge. And I think we're, what we'…
AI assessment note: “a big challenge that, that we face is that the world knowledge”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What about, um, evals? How do you build evals? Who manages them?
A Well, it depends on, uh, it depends on the application. So on the, on the operational AI side, um, those evals are basically baked into the, in the platform around every, um, every prompt or every agent. And for the most part, I think most of these use cases kind of come online, like the V one of like our, our, um, commercial underwriting agent or the V one of our, our startup KYC agent are co-developed between like a subject matter expert in ops and like an engineer. And they're going to kind of co-develop Um, uh, an initial eval set. But then from there, generally in ops, you're always doing QA, be it like on humans or on, uh, on, on the LLM, uh, decisions. And so whenever, like as part of our QA feedback loop, whenever there is, uh, a mistake, that's usually almost always gonna result in like, uh, another eval being written as like a regression test. Uh, so all of that within Ops AI is pretty, pretty straightforwardly managed. On the product AI side, that's where it starts getting a little bit more challenging because the multi, multi-agent network, um, is quite challenging to evaluate. And so what we do there is we try to adopt some of the state of the art for multi-turn evals where we will, um, we'll basically have a, an agent embody the user and like, you know, have basically the, um, the end user agent is given an objective and then we basically have it run a multi-turn,…
AI assessment note: “co-developed between like a subject matter expert in ops and like an engineer”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. Maybe run people through the Brex agent platform. We'll put the diagram in the video where you had the LLM gateway, you have like the whole MCP layer. We just had David, the creator of MCP. Right before you. So this is very timely. Um, yeah. How did you start building that? What's the architecture?
A Yeah, the architecture, you know, I, I think simple is, uh, is elegant and we, we've had basically an LL gateway and, and, uh, a basic hand rolled platform, uh, from the very early days. In fact, right before being tapped to become CTO, I was leading, uh, like an AI, uh, labs team internally, uh, in the wake of like the announcement of ChatGPT, you know, everybody saw this through technology and said, Hey, what are we going to do with it? And so one of the first things that we did, um, I think January, 20, 23, that would have been, uh, was Try to put together some internal infrastructure that made it possible for us to deploy, deploy, manage version and eval prompts, uh, and then be able to manage, uh, like data egress and model routing and, uh, have some very basic like observability and cost monitoring, uh, in an LLM gateway. So that's, that's infrastructure that we stood up and it still continues to power a lot of those smaller, uh, more, let's say like precise applications of LLM. So like, for instance, we've, uh, we set up a completely automated, uh, Pipeline for, um, evaluating, uh, customer applications to get them onboarded instantly to Brex, which is something that used to require, um, human intervention either for underwriting or KYC. But now we basically have a series of, of agents and, um, and particularly like research agents that will go and do the work that human…
AI assessment note: “put together some internal infrastructure that made it possible for us to deploy, manage version and eval prompts”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Are all the evals supposed to pass or do you have a set of evals that are like someday the model will be good enough and like how to change over time?
A Yeah, it's interesting. Um, I don't know if we have any that, that are like, oh, someday I hope it'll be good enough to do this, but it's like there, there are the evals that are, are blocking because they would indicate like a, a regression, an unacceptable regression. So these tend to be just accuracy related, um, evals, but then there are others that are more about like, Tone and coherency and these types of things where they're, they're more subjective and we were just looking at those over time as a, as a metric. But the, the team is actually interesting. I think we're going to get a big update on like how the team is thinking about evals tomorrow and like our Friday, our Friday review. So it's, this is an area where I'd say the largest challenge, like the largest change we needed to make and how we were executing Sort of as like a lab or an incubator back, uh, earlier this year to like where we are now where we've, we've shipped and like we're trying to, to increase the rigor has been around, uh, like avoiding regressions and having more and more increasingly robust devals.
AI assessment note: “there are the evals that are, are blocking because they would indicate like a, a regression”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Yeah. Any final call to action for things that you want to buy? Like what should people build for you? Like problems you're trying to solve that you would love people to reach out for to, to help?
A The call that I'd make is for folks who are interested in, in multi-agent networks to, to get in touch with us, because I, I do feel like this is something where, where we're, we're innovating in, in service of, of our customers and where I, I feel like the frameworks, the tooling, um, and the, the research is, is, is there. There's actually quite a lot of like interesting papers and things that we lean on. Uh, but I would love to, uh, would love to see more of that like encoded in the, um, in the, What's available at large in the industry, because I feel like my intuition has been that trying to craft LLMs into deterministic workflows and DAGs is, is kind of underselling like the power that they have to actually plan and execute more in a more sophisticated, like fluid way. And, and I, and I just want to see like the industry lean in more, um, on, uh, on these agent to agent, uh, Uh, interactions.
AI assessment note: “call that I'd make is for folks who are interested in, in multi-agent networks”