The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

640exchanges match
640on raw tape
31redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q as like potential retrieval artifacts for the model? Like, or do I have like the edge cases, it's like a good example, right? It's like, If you're building systems, you already have in your mind specific edge cases depending on it, but now you have to, like, every time repeat it. Like, are you having people spend a lot more time writing out more generic things to bring back or?

A Um, I mean, I do think well-written guides of, of how to do good software engineering are going to be useful because they can be used as input to models or, you know, read by other developers so that their prompts are You know, more clear about what the underlying software system should, should be doing. Um, you know, I think It may not be that you need to create a custom one for every situation. If you have general guides and put those into, you know, the context of a coding agent, that, that can be helpful. Like in, you can imagine one for distributed systems. You could say, okay, think about failures of these kinds of things. And these are some techniques you can deal with failures. You know, you can have, uh, you know, Paxos like replication, or, you know, you can, uh, Send the request to two places and tolerate failure because you only need one of them to come back. You know, a little description of 20 techniques like that in building distributed systems probably would go a long way to having a coding agent be able to sort of cobble up more reliable and robust distributed systems.

AI assessment note: “well-written guides of, of how to do good software engineering are going to be useful”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q How do you make the other engineers not feel like you're not special? I think that's something that I hear a lot. It's like, hey, you know, why aren't these people working on all the cool LLM things? And like, I'm stuck working on, you know, the KYC integration with whatever. You know what I mean? It's like, how do you build that culture?

A You know, it's interesting. I, I thought that that would be more of a problem, but the benefit of having really optimized Our engineering culture around business impact actually causes it to cut in the other direction where for folks, some folks don't want to work on the AI products because it doesn't have as much clear direct like business impact right now. Doesn't, doesn't impact revenues directly. And so I, uh, I think folks for the most part, uh, we've, we've enabled folks who have a strong desire to work on, on, um, AI products to, to join that team. Like somebody, somebody transferred out of our expense management organization to come over there because they're really passionate about Taking like their knowledge of like policy evaluation and, and bringing it into the, the AI, uh, uh, team. But the most part, I think everybody understands like how their work, uh, ladders up and maybe there's some like friendly rivalry because like the folks who say we're kind of a card product, they, they drive 60% of our direct revenue. And so they, they're pretty happy with that. And, uh, and they don't feel like they're being left out. Uh, and I will also say, um, as you probably saw in this, this piece that we, we, uh, put out with, uh, first round. There is a lot of smaller applications of LLMs peppered throughout all of our product and operations teams. It's just some of the more nov…

AI assessment note: “the benefit of having really optimized Our engineering culture around business impact”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So do you want to expand on that? Yeah.

A And this is an area where we've only scratched the surface here, but, but a big, a big challenge that, that we face is that the world knowledge or the knowledge that's built into the model about, uh, about, you know, what, GPT-V thinks Brex does and how it thinks our business operates is actually quite different from what our business offers today or how our product works. And so we've had to, to work on building a corpus of sort of product documentation, process documentation, and like curate this set of information to basically ground a variety of our LLM applications, including like that Brex assistant, which is like the You know, the assistant that employees, uh, will, will talk to is like, we don't want it to, to hallucinate features that we don't have, or like give, give wrong information there. And similarly, like, uh, some of the operational, um, uh, agents need to be grounded on, um, like what our ICP is, because if you ask, uh, you know, ChatGPT five right now, like what types of businesses does Brex, uh, onboard or like what types of businesses does Brex serve? It might not give an accurate Uh, explanation to that, to that question. It might, it might say we're a corporate car for startups, which is what we did, you know, seven years ago. And it might say we're only, we only serve enterprises. And so that has been an interesting challenge. And I think we're, what we'…

AI assessment note: “a big challenge that, that we face is that the world knowledge”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What about, um, evals? How do you build evals? Who manages them?

A Well, it depends on, uh, it depends on the application. So on the, on the operational AI side, um, those evals are basically baked into the, in the platform around every, um, every prompt or every agent. And for the most part, I think most of these use cases kind of come online, like the V one of like our, our, um, commercial underwriting agent or the V one of our, our startup KYC agent are co-developed between like a subject matter expert in ops and like an engineer. And they're going to kind of co-develop Um, uh, an initial eval set. But then from there, generally in ops, you're always doing QA, be it like on humans or on, uh, on, on the LLM, uh, decisions. And so whenever, like as part of our QA feedback loop, whenever there is, uh, a mistake, that's usually almost always gonna result in like, uh, another eval being written as like a regression test. Uh, so all of that within Ops AI is pretty, pretty straightforwardly managed. On the product AI side, that's where it starts getting a little bit more challenging because the multi, multi-agent network, um, is quite challenging to evaluate. And so what we do there is we try to adopt some of the state of the art for multi-turn evals where we will, um, we'll basically have a, an agent embody the user and like, you know, have basically the, um, the end user agent is given an objective and then we basically have it run a multi-turn,…

AI assessment note: “co-developed between like a subject matter expert in ops and like an engineer”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And is it like, I am a fortune 500, I need advisors on objective analysis and I call you guys and you pull up a custom report for me. You come into my office and give me a workshop. What, what, what kind of engagement is that?

A So we have a benchmark and insight subscription, which looks like standardized reports that cover key topics or key challenges enterprises face when looking to understand AI and choose between all the technologies. And so, for instance, one of the report is a model deployment report. How to think about choosing between serverless inference, managed deployment solutions, or leasing chips and running inference yourself is, is an example kind of decision that big enterprises, Uh, face, and it's hard to, hard to reason through. Like, this AI stuff is, is really new to, to everybody, and so we try and help with our reports and insight subscription companies navigate that. We also do custom private benchmarking, and, um, so that's very different from the public benchmarking, um, that we publicize, and there's no commercial model around that, but for private benchmarking, well, at times, Create benchmarks, run benchmarks to specs that enterprises want. And we'll also do that sometimes for AI companies who have built things and we help them understand what they've built with private benchmarking, um, you know, through the expertise, mainly that we've developed through trying to support everybody, uh, publicly, uh, with our public benchmarks.

AI assessment note: “we have a benchmark and insight subscription... We also do custom private benchmarking”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q What tools and what data connections come to mind when you say, what's interesting? What, what, what's notable work that people have done?

A Oh, okay. So, my favorite example on this is that until very recently, I would argue that it was Basically impossible to get an LLM to draft an email for me in any useful way, because most times you're sending an email, you're not just writing something for the sake of writing it. Chances are context required is a whole bunch of historical emails. Maybe it's notes that you've made. Maybe it's meeting notes. Maybe it's, um, pulling something from your, um, any of like wherever you at work store stuff So for me, like Google drive, one drive, um, and our super best databases, if we need to do some analysis or some data or something, Preferably. Model can be plugged into all of those things and can go do some useful work. Based on it, the things that, like, I find most impressive currently that I am somewhat surprised work really well in late 25 are that I can have models use Superbase MCP to read only, of course, run a whole bunch of SQL queries to do pretty significant data analysis and make charts and stuff, and can read my Gmail and my Notion.

AI assessment note: “models use Superbase MCP to read only... and can read my Gmail and my Notion.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q like your, like, how do you hire a sales team like this? Like we have founders listening, they're building interesting ad products. They don't really know how to go to market. Do you have to offer an arm and a leg to hire your first sales leader? Do you have to only work with Kleiner to do that? Like, what is the actual principle that you advise founders to follow?

A Uh, I'll give you some anti-patterns. The first is do not just go on their LinkedIn and look at all the fancy logos that they have gone and worked at and immediately assume that because they were at Snowflake or because they were at Databricks, they must be good for your AI company. It just doesn't work that way. In fact, in many cases, it's the inverse is true, where if you had to sell the number three product in a market, and you had to fight tooth and nail, and you were still successful there, You're probably, like, if you go to a great company, gonna have a much higher proclivity to do well, right? Whereas if you were, I don't know, if you joined Snowflake at a hundred million of ARR, and you join, like, their enterprise team in the Bay Area, it's like, yeah, I get it, but, like, that's not that impressive. No offense to anybody that joined Snowflake at that time, there were some diamonds in the rough. So I think that's, that's one.

AI assessment note: “do not just go on their LinkedIn and look at all the fancy logos”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q that you don't know that you don't know that is going to catch you off guard. Like some competitor is going to come in and eat your lunch or whatever, right? So to me, that is the, I think the debate of like why people are moving from offline to online or online to vibechecks is because it costs a lot of effort to do offline and it probably lasts

A Maybe six months if you're lucky. Well, I think, I think you're conflating two different things. One thing is sitting in a room and enumerating a set of scenarios or tests that you think represent a workload. And then the other thing is choosing to run those tests outside of production. And both of those things are activities that can happen offline. But I don't think anyone who's legit who's doing offline evals is doing the former anymore. Like, In fact, when we started BrainTrust and we talked to customers, we're, um, sort of very excited about BrainTrust helping them create golden data sets. And now I actually think that term is kind of a dirty term. People don't really want to create golden data sets. It's, I think it's, it's, it's often a wasted effort to, to the point that you're making. I think the best teams view offline evals as a mechanism of reconciling what they see in production with real users who are using the product with Tests and iteration that they can do offline. Like the best teams that we work with on a daily or even more frequently sometimes basis are discovering use cases from logs online and then pulling them into their environments and then playing with them. And I think the difference between doing offline evals and online evals or offline evals and just pure vibe checks is I think the pure vibe check or pure online version of this is you Observe some…

AI assessment note: “Well, I think, I think you're conflating two different things.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q get my guests to do a better job of is just brag. Could you brag a bit? Yeah. Just like some really impressive project that you accomplished, just, just the opens people's minds. Like, let's get, let's get specific without maybe naming the exact client, unless you can. Um, and then also like, what's the highest hourly rate that one of the engineers has made since you're technically on path?

A Yeah. So I'll answer the last one, uh, or the second one first. We will probably have more than one engineer make million dollars cash next year based on this model. And that is just with story point compensation. It's very likely that we will have more than a handful of folks make more than a million dollars next year. Um, The answer to the first question, like, for example, one project we built, so we work with this company that's a, they, they build, they work, they partner with retailers to basically make cameras in the business more valuable, and the way that they do that is they deploy what was historically like a gen four raspberry pi to the stores, and they would, they would run like one model on that device. We basically took some off-the-shelf models and trained some models ourselves, and then quantized them down so they could actually run on that Um, on that for, but also on jets and nanos, and we got them to all run in parallel. So now basically what these models allow you to do is as a store, you can get a heat map. You can see where the lines and the cues are forming in your store. You can even get pictures of shelves and understand what needs to be stocked. And you can do things like theft detection because we have body analysis and we can understand things like things are crossing arms, right? And This took our team two weeks to put together early prototypes, an…

AI assessment note: “We will probably have more than one engineer make million dollars cash next year”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah. That's amazing. Okay. So, uh, like quick question on just the, uh, the stack that you guys have landed on, like, is there a house stack, what are you guys finding in terms of like the various coding agents and all that?

A Yeah. Um, we do work in a number of different stacks, a number of different languages and stuff, but We feel pretty strongly in like high structure allows for agents to work autonomously for longer. And so our default stack is TypeScript front end TypeScript back end with a shared file where, or a shared folder where all of our shared types and schemas and things like that live and typically react front end, or even something as simple as like express on the back end. Like we don't really care about the frameworks. It's more just like TypeScript allows us to have that flexibility to, um, like the flexibility of JavaScript, but the, but the constraints of TypeScript. And then those error messages allow the cloud code or cursor agents or whatever to iterate on themselves and, and run things, see the errors and, and continue in terms of the actual like AI engineering stack and what coding agents and things like that we're using. I always tell clients this, like our team doesn't have a favorite coding agent of the year or of the month or even of the week. Like if I go over there to our team right now and I ask them what model is performing the best for coding? Right now. They'll say today at four 42, we're noticing that Claude code is actually performing better because of X, Y, Z reason. But yesterday codex was outperforming Claude code on object on, on activities like X, Y, and Z.…

AI assessment note: “our default stack is TypeScript front end TypeScript back end with a shared file”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So concrete example, Goodfire is like the most, the most interesting one. Mechanistic interpretability. I didn't even think that was a market that was worth investing in, but obviously Anthropic does. Uh, and, uh, they seem like they have good vibes. What, what's the, I guess the, the summary of your, of like your take on the company?

A The way I think about the company is right now, almost all frontier and some many non-frontier AI models are complete black boxes. You don't understand why they produce the outputs they produce. All of the eval and studies on them are empirical studies, not intrinsic to the model. So it's like, Hey, here's the outputs we saw. And therefore this is the benchmark score, or this is how we think it did. If we believe as a society that Five and 10 years later in the future, these models are going to be critically important for making pretty heavy decisions, whether it's, I call it anything from whether somebody should get a loan or insurance or a legal decision, then I don't think that the black box approach is long-term scalable. It's just not how society can function, where it's, you say, you throw your hands up and say, well, this is what the model said. And then I asked it, explain yourself. And it said this other stuff. Great. Like that's kind of what we have today. That's the best thing that we have. Mechanistic interpretability is really going into the weights of the model and trying to figure out why did the model do what it did? And one of the more concrete and relatable examples of this that, you know, you guys may be aware of is GPT-IVO had this phase of sycophancy that, um, a lot of users really liked, but It's kind of one of those things that's not as easily detectable …

AI assessment note: “The way I think about the company is right now, almost all frontier... models are complete black boxes”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q the scale changed with AI? So before, you know, I used to run some software website and we would have the same issue. People buy the software and then gets charged back, but it's like 20 bucks. Like today, you could, like, you know, use the credit card and sign up for the OpenAI API and spend 10,000 dollars, 15,000 dollars, like much what's the shape of the fraud today?

A So friendly fraud is like not stolen card credentials, but something like non-payment abuse, free trial abuse, refund abuse. So it's me, they're my credentials, but I'm not actually creating a creative revenue for the business. This has happened for a while, and actually, if you, if you ask business leaders, like, I think something like 47%, payments leaders, like 47% of them will say that their biggest fraud challenge is friendly fraud. I would say this was, like, just much less of an issue for SaaS for two reasons. One, like, what were you stealing? You weren't stealing computer inference or whatever. And two, more importantly, even if you were stealing some service, like, the marginal cost of providing that good or service for Salesforce or whomever was like near zero, and so it didn't totally crush your unit economics. Now we're in the world where GPUs are expensive, inference costs are high, and free trial abuse or refund abuse or general non-payment abuse, right, you, you rack up these charges and you never pay, is like existentially threatening for AI businesses. I was talking to a small AI founder the other day because we're, we're building sort of a suite of Radar extensions that are explicitly targeted at this type of fraud. And everyone tells me it's a huge issue, and so with every company I talk to, I try to dig in on, for you, what exactly is the issue? And there's…

AI assessment note: “Now we're in the world where GPUs are expensive, inference costs are high”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So this is like from the outside how people perceive AI as Stripe. What about Insight? So you mentioned 3.5 was kind of like the moment you took it seriously. What were the first internal use cases and then how do you use AI as Stripe today?

A Yeah. First internal use cases were, you know, bottoms up experimentation, right? So we created, we call it go LLM, but it's like just a chat GPT like interface where you can engage with a bunch of different models. Um, it was the very, very first version actually wasn't like an LLM proxy where you could build production grade systems. It was literally just like chat GPT like stuff. And then we had this like preset feature, which was like prompt, Saving and sharing. And so you could share your temp. Oh, you know, this is how I figure out what customers to reach out to and generate reach outs or, um, you know, rewrite my marketing content in striped tone or whatever. And you had sort of hundreds of presets that came on like overnight because, uh, everyone was into it. And then we generated, so then LLM proxy was like, okay, now production grade access for engineers to these LLMs. And a lot of the early use cases there were actually around merchant understanding. So I mentioned a little bit ago, but we have like thousands of merchants that come under Stripe every day and we have to understand who are they? What are they selling? Is it supportable through the card networks? Like, are they credit worthy? Are they fraudulent? Um, and there's a lot that LLMs can do there. So those were some of our earlier, um, earlier use cases. Fast forward to today. I mean, you know, I, I actually …

AI assessment note: “First internal use cases were, you know, bottoms up experimentation... Fast forward to today”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q transition and that, that, that, uh, is the perfect intersection of financial infrastructure and AI. So maybe, uh, could you tell this, the story of ACP, right? Like, I, I think, uh, this is one of the, the, the biggest launches of, I guess, like in the, in, in, in the second half of the year. And, and like, I guess a really important strategic move between OpenAI and Stripe.

A Yeah. So, you know, we talked a bunch about AI companies in general. One important slice of AI companies is AI commerce, agentic commerce. And, you know, I think just zooming back, like we're all spending more and more time in some combination of broad consumer-based tools like ChatGPT and AI dev tools like Replit or Vercel or whatever. And We want those agents, those tools to increasingly take action on our behalf. And I think, you know, we saw an early version of this in chat GPT with operator, but an important area we want them to take action is buying on our behalf. You know, sometimes it's recommending products, but often it's like literally getting it all the way over the wall. So, um, a couple of weeks ago, we announced our agentic commerce protocol, which is joint with OpenAI. And it's basically just a shared standard for how businesses can talk to agents. So if you think about it, like, it used to be that a human was buying from a business. Now there's an agent that's sitting in the middle. And that fundamentally needs to change how the financial infrastructure works. Like, checkout needs to look different. Fraud checks need to look different. Payment flows need to look different. But also, merchants are trying to figure out How they can efficiently expose their product catalog, their inventory, their brand, their pricing through a range of agents to have access to tha…

AI assessment note: “a couple of weeks ago, we announced our agentic commerce protocol, which is joint with OpenAI”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah. Yeah. Uh, cool. Uh, do you, uh, just a side follow-up, uh, did you actually follow the sort of, uh, schedule free optimizer stuff from last year with Aaron Defazio? Anything came out of that?

A I don't have a strong opinion. I tried it, uh, like a few days after it was released because I was, uh, we were writing a paper about the WSD schedule. And so this was kind of an alternative to it. And, uh, it didn't perform as well as like, uh, Simple Adam W and it was more sensible to some, to the beta one and beta two. But I think like with this kind of stuff, you, you really need to have this, uh, this knowledge on how to optimize, uh, those hyperparameters for each optimizer to really get the best performance of each one. And, uh, so I don't want to, to say something bad about it because basically at 19%, uh, there is a 90% chance that it's just me that used the wrong hyperparameters.

AI assessment note: “I tried it, uh, like a few days after it was released”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah. So you left YC, you spent a year in the, kind of, the wilderness. You went through YC, uh, as 23. Uh, what's that journey like? What's the...

A You know, I was very excited about AI things in general. Um, this was, so I left YC, I guess, uh, beginning of twenty-twenty-two, and I was trying out a bunch of different things. Um, ended up landing on what turned into OpenPipe in early twenty-twenty-three. This was, uh, let's see, so I'd been working, so my, my co-founder is my brother, um, my little brother, which has been a fun journey on its own. We were looking at different ideas, and one thing we realized was we actually started the company immediately after the GPT-IV launch, and what we saw as the opportunity in the market at the time, which has changed since then, was GPT-IV was insanely expensive and extremely powerful, but there was an opportunity to distill, like, specific workflows from GPT-IV down to much smaller, much cheaper models, and there was, like, a very clear value prop there. Given how expensive GPT-IV was, it was hard to deploy in production, But you could sort of like take those abilities and deploy them much more cheaply. So, so that was kind of the first thing we built was this kind of very managed, very clean, um, distillation flow.

AI assessment note: “I left YC, I guess, uh, beginning of twenty-twenty-two, and I was trying out”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Um, what's your, uh, Vibevel? So as soon as you get a new model, what's like the first two, three things that you do with it?

A It's such a good question. I have three, uh, and, and the, the team got really tired of me by the end of this, uh, Sonnet four foot five process. Cause like every, literally every snapshot, you know, I would drop everything I was doing and run these. So one is, um, the virtual boy, I had a, you know, I actually never owned a virtual boy, but the local video game store in Brazil when I was growing up had one. And then if you remember, this is like a doomed Nintendo console that had these like a stereoscopic wireframe Red and black, three D graphics is a very early product. It totally failed. Um, but I always, I like to have, uh, in cloud AI, like create me a like virtual boy style, three D shooter game. And it's really funny to see the, the, all the checkpoints from like early Sonnet 4.5, where I was like, I don't know, guys, this is like, it's pretty early. I know there's a lot of RL left, but it's not looking that good to, you know, about a week ago, I was like, okay, great. This is like officially good. It's like better than Opus at this. It's like, It generated this, like, great split-screen stereoscopic thing, three-dimensional, like, thing. So anyway, you could really see it evolve. So that's, that's one I always, uh, shoot for. Um, within Cloud Code, there's a particular, uh, sort of change to our code base that I, I, I, like, once did in Cloud Code. I was like, oh, this …

AI assessment note: “I have three, uh, and, and the, the team got really tired of me”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q the final one with Sonnet 4.5, it has the chat history right on the side, versus on the actual Cloud AI, you have to click through to go to the history. Like, do you use these models to, like, think about all the different permutations of, like, how to build these UIs and products, or do you feel like the models are still, you know, pretty median results so far?

A One of the better projects that we did here was get a lot of our kind of products into artifacts, or at least into a harness that we could iterate on them really quickly. So one really fun use, um, Nate Parrott is one of the designers on our team. He, he got Claude to be able to prototype Claude code UIs, even though it's a terminal UI, Claude can sort of imagine what that's like. And it's very valuable to say like, all right, well, what it would look like if we revamp settings in this way and not have to go and necessarily code the end. So I think they can be very useful in sort of exploring that space of What, how could this product evolve? What does it mean to do this, uh, differently? And they can also actually just implement the changes as well, but even sort of from a prototype phase, it's valuable to have it sort of iterate quickly over ideas for non-engineers on the team.

AI assessment note: “I think they can be very useful in sort of exploring that space”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I can also go into CLI and use some code. Was that an easy choice? Like, was there a lot of discussion on we should just do one of the modes? Like supporting both is obviously more work, right? And a lot of these products don't support both. So what was that initial design choice of the structure of the product? And then we'll dive into the models as well.

A So we started with the VS Code extension because it was the easiest thing to get off the ground. Like when you have a VS Code extension, you have a marketplace, you can ship this, you can update it 15 times every day. You don't have to think about updating stuff. You also are next to the editor. And looking back, you know, it's been six months, the editor might be dying or you might do a lot of coding outside the editor. Back then it sounded much more radical than it does sound right now. So we started with like, let's explore this and having the thing next to your editor is a good place to start. And we could, you know, you can see the cursor, you can do selection and whatnot, but we were really like, um, from the start, we didn't want to have like a deeply integrated thing. It was always like, ah, let's keep the feature small. We got to be able to move fast. And then we build up the CLI on the side as like a different client, which also It also gives us the ability to abstract like the core and the client stuff. So that's a nice boundary to have. But then to be 100% honest, we were also surprised by how many people were fine with using a CLI for Cloud Code, for example. Like if you had asked me half a year ago, I would have said no way, like a CLI tool. And what we realized is, well, a CLI is not just, you know, It's a UI, sure, but also it's a CLI program. That means you can…

AI assessment note: “we started with the VS Code extension because it was the easiest thing”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q we can sort of pull it up on screen now. I always like to give people the sort of visual aid, uh, as well while we talked. Yeah. I mean, just a quick, you know, while we're putting it on screen, why, why did you have to rename the company? Like, why can't you just be like, owner by Gitpod, right? You know, like what, what, um, why so drastic?

A I mean, so what we understood as we initially also thought about just naming the agent owner, what we quickly understood is that it's a core part of the experience and it's a core part of the platform. Those two things are not separate. Right. And, um, that is, that was ultimately then the driver of the decision to say, Hey, um, let's, let's embrace that. Gitpod is also as a name, very technical. So we kind of have outgrown that name. To a certain extent, like with Git and Pod, it's like, it's a Kubernetes-based architecture. We released a blog post a couple of months ago that was called We're Leaving Kubernetes. So it just did not really fit. So generally you can say our ambition, product architecture, and also the value that we deliver to our customer has outgrown the core, you know, um, name of GitPod.

AI assessment note: “our ambition, product architecture, and also the value that we deliver to our customer has outgrown”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah, totally. Uh, what is, what is other challenges? I guess, like, people always talk about, like, memory. Do you talk about, like, which model you're using? Any discoveries from, like, using GPT-V versus Cloud-IV, you know, anything like that?

A Yeah. So we, right now, um, essentially live on top of Sonnet IV, and, uh, we've tried a bunch of other models. We found that to work really, really well for us. The, like one of the, the trickier parts to get right actually was file it. It's surprisingly like the, you know, we, we tried different iterations. We tried to come up with our own. We started very simple and we're like, you know, here produce diff. We're gonna apply diff that didn't go very far. We tried to tell it, use SED to make edits and then had a lot of, um, problems with it getting back slashes wrong. And we tried essentially a fork of Klein's apply edit tool for a long time. And at this point we've landed at, uh, essentially anthropic's standard, uh, string replace built-in tool, and that's working decent, but it's obviously model specific. And you know, so the moment we, uh, like for the other models that we support, we will need to find something else, but file edit. Then no, but it's really, really hot and it's very much at the heart of it. Like if that fails, your agent's gonna go astray and burn tokens for no good reason very quickly.

AI assessment note: “essentially live on top of Sonnet IV, and, uh, we've tried a bunch”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So the keywords for that data curation is a service, data efficiency, all those terms. In the pre-chat before we started recording, you mentioned that there's a cool story around how you got into data in the first place, right? You were at GDM. You were at Meta as a research scientist. Describe how like that became an interest.

A My PhD is actually in neuroscience. Uh, so I come much more from an empirical science sort of background. I actually spent time trying to teach mice how to count and then analyze the activity of thousands of neurons in the brain while mice did count and try to understand how did that actually happen? How, what were the neural dynamics that enabled that? Um, and that's actually initially how I got into machine learning was as a means to analyze my, my neural datasets. I also started my PhD, So Alex came right after that, Tari DQN right after that. Lots of evidence that AI was going to be very, very exciting, which, which led to me transitioning. But as a result, because I had this kind of somewhat different background of being trained as an empirical scientist rather than as a computer scientist, my real first mission when I, when I joined AI was to try to build more of a science of deep learning. Something that I think, you know, is still true today in many cases is that deep learning is an empirical science, but most people that have computer science backgrounds were trained more In the context of a branch of theory, right? Everything was very provable. That was the initial pushback to deep learning, actually, was that you couldn't prove anything in it. But deep learning is, at its core, an empirical science, right? We have to run large experiments. We understand the rules for…

AI assessment note: “that's actually initially how I got into machine learning was as a means to analyze”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I'll push you a bit on this. Yeah. Um, you know, I think a popular view is post training is elicitation. Of capabilities that you already trained in pre-training. So what dependencies can you have that feedback into the, into pre-training?

A So, so I'm inclined to, to agree with that view. And, and I think that that view would lead very strongly to the fact that you should be trying to optimize your pre-training data to make post-training processes more effective. So you should try to figure out how do I optimize my pre-training data so that the slope of the test time compute curve, or so that the slope of the RL curve is as steep as you possibly can be. Um, or alternatively, how do I optimize my pre-training data so that the slope of the jailbreaking curve Is as shallow as possible, right? Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to take it out. If it's really hard to put it in, it's really hard to take it out. That's just like a truism of models, right? So if you do alignment during pre-training, you'll actually end up with models that are, I think, largely impossible to misalign without putting a massive amount of data into them. Um, I think there are a lot of benefits to that. Um, and I think we've also seen evidence for this, like looking at the difference between Lama and Quen with respect to their ability to be post-trained, right? It's much easier to RL Quen than it is to do Lama. Likely that has t…

AI assessment note: “optimize your pre-training data to make post-training processes more effective”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Amazing. And then like, so, you know, I think there's, there's a lot of questions about the core agent harness agent loop. Um, yeah, let's just like talk about the technical lesson on the journey that you went through. Where did you start and where have you ended up?

A Yeah. So I, I obviously didn't know about any of how any of this stuff worked until we're like, maybe like about two months into development. So roughly two months ago, I knew that all an agent was, was a loop running somewhere. And in that process of doing this, I've learned some of the details about that. Like I said, we don't try to innovate on how that works. We look at cloud code. We dump all their system prompts. We dump all their tool descriptions. We dump all their tool schemas. We re-implement the tools. And when you're using an anthropic model, we basically have the exact same implementation. Some of the details of how to figure, like some of the prompt caching stuff was tricky to understand how that worked. But at the end of the day, like it's just a loop where you send the prompt with a bunch of tool descriptions. It tells you what tools to call. You call it tools and the results back. And there's really not too much magic there.

AI assessment note: “didn't know about any of how any of this stuff worked until we're like”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q I mean, I think either is like up there, but also like, this is curious, like since you build in public, Probably this is public, but like, what's an example of a thing that, you know, locally temporarily improved things, but actually you regretted it two weeks later?

A Yeah, so I did this thing. So the edit tool, the way the tools work on most of these tools now is you ask the LM for the file you want to change the old string to look for and the new string to replace. And now there's some percentage of the time where it nails it. You search the old string, you find it, you replace it. Sometimes it kind of messes up what the old string looks like, and you search for and it's not there. But it's usually A lot of times it's a very simple mistake. Like it used tab spacing when you use spaces. Um, so you can have a series of fallback strategies and the client team published all of their fallback strategies, which I then ported to open code. And then I was like, that was easy. I just had open code port it. Let me go crazy. I went to Gemini's CLI and I was like, let me take all their strategies too. And Gemini ones were not good. It was, uh, initially it looked good, but then I noticed it was getting caught in these crazy loops and it was like editing my files in these totally messed up ways. And I should have known because you can tell when you're looking through the Gemini code base, it's like really rushed. So I ended up dropping those strategies, but there was at least like two or three weeks where That was definitely a regression. And I only knew because I was using the tool.

AI assessment note: “there was at least like two or three weeks where That was definitely a regression.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q actually using transformers for at least the mercury models. Yeah. Um, and basically does it have to be trained end to end, or can you just take any existing open model backbone and, uh, uh, and, you know, do this process? Like, do you have to initialize from zero basically, or can you just, can you lean, can you sort of leverage off of existing, uh, other, other language models?

A There have been, uh, attempts in the literature, um, where people have tried basically using, um, existing, uh, open source models and then kind of like fine tune them. The challenge is that, yeah, the training objective is quite different because you are training, uh, based on denoising as opposed to next token prediction. Diffusion models are not causal and that is also kind of problematic. Uh, I mean, it's a big advantage of diffusion models that you don't have to be causal. You can actually look at the whole context to the left and to the right. As you decide, kind of like the edits that you want to do, which is giving, you know, it's a, it's a very powerful thing if you think about how you generate objects, but that makes it quite different, quite difficult to adapt kind of like models that have been trained with, with causal masking to, to this, to this, to this, to this new task. And so there have been attempts in the literature, annealing the, the, the masks, uh, the attention mask, like there have been various, Tricks and attempts at doing that, but it's not straightforward.

AI assessment note: “makes it quite difficult to adapt kind of like models that have been trained”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And to wrap on maybe the pre-training phase, is there also post-training in the same way that people do with normal LLMs? Well, not normal, but you know what I mean? Like instruction following and things like that. Does that work the same as those models or is it different in diffusion?

A So our, uh, pipelines, yeah, we have something similar in the sense that we also use, uh, we can do basically SFT and we can also do, we do reinforcement learning, RLHF. We do reinforcement learning from German feedback. Again, some of the algorithms have to change. For example, we have a DPO algorithm specialized for diffusion language models. I was one of the original authors of the DPO paper that was originally designed for LLMs, autoregressive LLMs. I worked on a paper extending DPO for diffusion models, more continuous diffusion models, like image generation, video generation. And that's also been quite impactful Like if you look at some of the papers that have been released on some of the leading open source image generative models based on diffusion, you can see that the DPO piece using human feedback to really align the models to human preferences is really important if you want to improve quality. And we've extended a variant of DPO to diffusion language models. So we're actually also able to use preference data to, to fine tune the models. And, uh, but yeah, basically we have pipelines that are compatible with the Existing SFT and RLHF data sets. And we, we do that with customers. Like we often work with customers that have their own internal data sets. They want us to fine tune the models on their proprietary data, and we can easily do that. It's all pretty easy once…

AI assessment note: “we have something similar in the sense that we also use, uh, we can do basically SFT”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q cognizant of the Serving requirements between autoregressive and diffusion. A lot of people rely on the batching characteristics in, in production. You would not have that, right? Are you aiming to, like, uh, you know, are you aiming to be, to abstract that problem away from people? Or are you just going to release the models open and open ways and then people will deal with it however they want?

A So we don't have a plan at the moment to release models or to open source any model. And in fact, it's a little bit tricky even to, even if we wanted to release, let's say a small, maybe less capable version of the models, we'd have to release the inference code, which is also proprietary. Unlike an autoregressive model where the inference code is kind of like straightforward, there's not that many ways you can do inference. In an autoregressive model, you can tweak a little bit the sampling, but it's, it's pretty straightforward. In a diffusion language model, there is a lot more knobs and a lot more innovation that can happen on the inference side. And so it's very hard for us to open source a model because we'd have to release the inference code. And by the way, we have our own inference engine. So just like you would normally serve an LLM using a VLLM or SGLang or a Tensor or TLLM, we have built our own inference engine. And so we are supporting production traffic already with our own inference engine. We support continuous batch and quantization. We support a lot of the features. That are supported in these serving engines, caching, like prefix caching, like a lot of that stuff is already handled by our own serving engine, but yeah, it gets a little bit tricky if we wanted to, to open source because a lot of the IP is around on the inference side as well.

AI assessment note: “we don't have a plan at the moment to release models or to open source”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Part of me thinks that this, the sort of LLM centric problem is very different from the machine learning problem that you had at Cruise, but maybe it isn't because What's different? What's the same?

A Machine learning is still very frustrating. It's always been frustrating. Um, that's a great question. I think like over the last, you know, four years, there's been huge transformation in the industry and how machine learning has been practiced. Um, you know, before you might focus at Cruise on labeling massive data sets, millions of images, right? Or LIDAR and, and fine tuning models that are like a lot smaller. These days, the models are, are massive. Um, they generalize extremely well. You can use techniques like prompting to kind of few shot them. Even fine tuning itself is very sample efficient. So it's definitely a lot faster to iterate. It's a lot faster to get into production. That's actually really, really exciting. The thing that I'll say that they have in common, um, is it's still a data engine. It's still very data-driven, you know, um, machine learning where, what data do you have? Like, what's the distribution? How do you collect really good annotations? How do you evaluate success? Some of these fundamentals haven't changed.

AI assessment note: “The thing that I'll say that they have in common, um, is it's still a data engine.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah. What was the model evaluation at the time? So I'm sure part of, like, the We Need Plan and Act is, like, maybe the models are not able to do it end to end. When you started working on that, what were the model limitations, what were the best models, and then how has that evolved over time?

A Yeah, when I first started working on Klein, this was, I think, 10 days after Cloud through Five Sonic came out. I was reading Anthropics Model Card Addendum, and there was this section about agentic coding and how it was so much better at this step-by-step Accomplishing tasks. And they talked about running this internal test where they let the model run in this loop where it could call tools. Um, and it was obvious to me that, okay, they have some version and they have some application internally. That's really different from how the, you know, the other things at the time were things like copilot and cursor and Adr. They didn't do this for like step by step reasoning and accomplishing tasks. They were more suited for the Q and A and, and one shot prompting paradigm. Uh, at the time, I think it was, uh, June, 24. Anthropic was doing a build with cloud hackathon. So I thought, okay, this is a really cool new capability that none of the models have really been capable of doing before. And, uh, I think being able to create something from the ground up and take advantage of kind of like the nuances of how much the models improved in that point in time. So for example, cloud through five was also really good at this test called needle in a haystack, where if it has A lot of context in its context window. For example, you know, 90% of its 200 K context window is filled up. It's real…

AI assessment note: “when I first started working on Klein, this was, I think, 10 days after Cloud through Five Sonic”

← previous page 2 next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.