Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q give you a hand on that. I know in the US there's not that many dialects, there's like accents, but like most of the language is like, like the words that people use are similar. Because I know, for example, Spanish is like, you know, Spanish spoken in Argentina is like very different than Spanish spoken in Mexico. How do you kind of adjust for that? Or maybe you don't.
A So I would say that, for example, currently we teach American English, standard American English. We don't really teach other accents or other dialects. For now, given how small we are, we just have to be pragmatic and Teach in the direction that most people want and most of our users know. So we've made those decisions like on the contacting side for American Spanish, every language that we're teaching. But I do expect that in the future, we're going to get a lot more sharply differentiated. Like if you want to learn British English, then we'll teach you British English. We'll teach you how to pronounce it, et cetera. I think all of that feels like something that superhuman language, you know, tutors should be able to do.
AI assessment note: “currently we teach American English... We don't really teach other accents or other dialects.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q is cool. And we can talk about that as well. But the other thing I think is that around about seven, eight B, maybe four B is when you start switching from like a single GPU setup to like a distributed setup. And I'm wondering, like, do grad students get HPC training? How much do they teach you of like, just how to work with like large clusters of stuff?
A Oh, to be clear, they don't teach you anything, like anything, like if you see a paper coming out from even, you know, Stanford, they're probably the best school in AI if you had to choose. And it's not like they're learning how to do like multi-node distributed FSTP training, like with whatever deep speed, you have to learn that from the internet and from other people. And like, there's no classes that really do that. I mean, it's, that's hard to facilitate, like, As one person, I would say most grad students are doing stuff on single GPU. Some people are doing multi-GPU training. There's probably basically no grad students doing multi-node training. I mean, there's probably a few, especially if they have like company affiliations, but that's really unusual, I think.
AI assessment note: “Oh, to be clear, they don't teach you anything, like anything”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q is like, what's valuable to me more than like the Terminal itself. Why people like cloud code is to bring it around, you know? And like with warp code, it's like, I would love to bring it around, but you can imagine somebody saying, I'm not getting anything from warp outside of warp code. So why do I have to only use it in warp? You know what I mean?
A Doing like a headless, like CLI version is interesting. My take on that is that there is a lot of value in what we can do at the UX that if it's a pure CLI app, you just can't do. So like, There's no way for cloud code running in ghosty, let's say to have an editable diff. Uh, now you might say like, well, why do I need an editable diff? Like if it does it perfectly, even if you don't want to edit the diffs, I think honestly, like the next thing that we'll probably be building is like basically a way in warp of code reviewing what the agent is produced. Cause I think that you totally want, like if you're having an agent write code, I find it very, very cumbersome to like push all the changes up to GitHub. Review it myself, then come back into my tool and have the, you know, then retell it to do it. And so there's all these like native UX things that I believe you want if your primary workflow is like launching and managing agents. There is not available for your CLI app. I would say like our likely path forward on that is like, we will probably at some point make a warp CLI app where if you really want to run it headless, like you want to run it in CI. Cause I do think this is an advantage of like cloud code. You can run it there, but I think for like the hands-on keys, like I'm a developer coming into work and I want to like use these tools to delegate work. I think it's a muc…
AI assessment note: “there is a lot of value in what we can do at the UX”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Do you have any public numbers that we can reference of, like, What are you saying regarding growth trajectory of Warp just, just to establish in people's minds if they've never heard of you?
A You know, it's like almost 600,000 active users on Warp right now. Our revenue since the beginning of the year is growing, like, depends on the week, between five to 15% every week, and so, like, it's actually a pretty fast revenue ramp. Now, we're not nearly at the scale, just to be clear, of, like, Curse or even Windsurf, but, like, we have pretty Awesome product market fit with people who discover the coding and paid parts of warp. So the other thing is we're in an amazing market right now is the other thing I would say is like the, um, and this wasn't true a year ago, but like now it's, uh, it's very expected to pay for developer tools. And for the first time we have companies that are coming to us being like, my developers aren't using the latest AI tools. Like how do we get them to use it? And that's like a totally new, Development, which is very good for us. So we're growing really, really well, and we have this very big existing user base, you know, not big like relative to like chat GPT is like five hundred million users or whatever, but for a developer product, like pretty substantial, and that's also growing pretty quickly.
AI assessment note: “almost 600,000 active users on Warp right now. Our revenue since the beginning of the year”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What's your day to day? Like you have a lot of meetings. You have just a few.
A Yeah. I have a lot of, you're talking about my, my, my lived life. Um, a normal weekday for me is, um, I wake up about seven, 15, my wife and I get the kids out the door and she drives them to school at about eight. I usually do a half hour, a 45 minute walk with my dogs, get exercise, strenuous up and downhill. Heart rate goes up, which is good. Listen podcasts, for example, yours. So that's, you know, why, why I have time to actually follow Exciting, cool things that other people are doing. Get to work at nine. And so work nine to six or seven or something like that. And most of that is meetings. And so that's me trying to solve whatever the problems are of the day, get home, dinner with the kids. I insist on eating with my family, hang out with them until bedtime and then crush through for like two, three more hours until I pass out. And do that regularly.
AI assessment note: “work nine to six or seven or something like that. And most of that is meetings.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Last but not least, the question that everybody wants the answer to. Have you finished building the Lego robotics table for your kids that you mentioned last time? And what's the, do you have a new project that you're working on?
A Massively forgot about this. So we still use Lego robotics table. So this is a big four bait sheet of plywood with a bunch of, uh, two by threes around the edge, but a four by eight sheet applied was pretty hard to work with. And so it breaks apart into three decomposable modular sections. And so it's super great. And the kids are getting way better at programming and Legos and all this kind of stuff. Um, gosh, what's my most recent project? I was just building swords in the shop with kids on a bandsaw. And so you take a piece of wood, you have a bandsaw, give it to a kid, and you say, don't cut your fingers off. Turns out that the bands, I mean, I'm, I'm obviously joking a little bit. They get a lot of, they get a lot of oversight, but a bandsaw is actually a very safe tool. And so the reason for that is that a bandsaw, which if you, probably a lot of people have never seen a bandsaw, you can do a Google search for, you have two wheels and then you have a blade that goes around the wheels. And you've got a table, and the cool thing about a bandsaw is that you can crank down the opening towards the blade, so it's just, just as big as a piece of wood, and also, it pulls the wood into the table, and so, because it pulls the wood into the table, the risk of, like, something called kickback, and there's a lot of other things like this, is very low, and so you can basically tell a k…
AI assessment note: “So we still use Lego robotics table... what's my most recent project? I was just building swords”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. How did that start? Did you have a master plan to build cloud code or?
A There's no master plan. Uh, when I joined Anthropic, I was experimenting with different ways to use the model kind of in different places. And the way I was doing that was through the public API, the same API that everyone else has access to. And one of the really weird experiments was this cloud that runs in a terminal. And I was using it for kind of weird stuff. I was using it to like, look at what music I was listening to and React to that and then, you know, like screenshot my, you know, video player and explain what's happening there and things like this. And this was like kind of a pretty quick thing to build and it was pretty fun to play around with. And then at some point I gave it access to the terminal and the ability to code. And suddenly it just felt very useful. Like I was using this thing every day. It kind of expanded from there. We gave the core team access and they all started using it every day, which was pretty surprising. Uh, and then we gave all the engineers and researchers at Anthropic access and pretty soon everyone was using it every day. And I remember we had this DAU chart for internal users and I was just watching it and it was, it was vertical like for days and we're like, all right, there's something here. We got to give this to external people so everyone else can try this too. Yeah. Yeah. That's where it came from.
AI assessment note: “There's no master plan. Uh, when I joined Anthropic, I was experimenting”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q and whatnot. But I know you recently released like, you know, compacting context features and all that. How do you decide how thick it needs to be on top of the CLI? So that's kind of the share interface. And at what point Are you deciding between, okay, this should be a part of clock code versus this is just something for the IDE people to figure out, for example?
A Yeah, there's kind of three layers at which we can build something. So the, you know, being an AI company, the most natural way to build anything is to just build it into the model and have the model do the behavior. The next layer is probably scaffolding on top, so it's like cloud code itself, and then the layer after that is using cloud code as a tool in a broader workflow, so to compose defend. You know, so for example, a lot of people use code with, you know, tmux, for example, to manage a bunch of windows and a bunch of sessions happening in parallel. We don't need to build all of that in. Um, compact, it's sort of this thing that kind of has to live in the middle because it's something that we want to work when you use code. You shouldn't have to pull in extra tools on top of it. And rewriting memory in this way isn't something the model can do today. So you have to use a tool for it. And so it, it kind of has to live in that, that middle area. We tried a bunch of different options for compacting, you know, like rewriting, uh, old tool calls and, uh, truncating old messages and not new messages. And then the end, we actually just did the simplest thing, which is ask Claude to summarize the, you know, the previous messages and just return that and that's it. And it's funny with, when the model is so good, the simple thing usually works. You don't have to over-engineer it.
AI assessment note: “there's kind of three layers at which we can build something.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q How can writing a file ever be unsafe if you have version control? I think that's
A Yeah, I think it's, I think there's, like, a few different probably, like, aspects of safety to think about, so it could be useful just to break that out a little bit. So for file editing, it's actually less, I think, about safety, although there, there is still a safety risk, because what might happen is, let's say the model fetches a URL, and then there's a prompt injection attack in the URL, and then the model writes malicious code to disk, and you don't realize it. Although, you know, there is code review as, like, a separate Kind of wear there as, as protection. But I think generally for file rights, the model might just do the wrong thing. That's the biggest thing. And what we find is that if the model is doing something wrong, it's better to identify that earlier and correct it earlier, and then you're gonna have a better time. If you wait for the model to just go down this like totally wrong path and then correct it 10 minutes later, you're gonna have a bad time. So it's better to usually identify failures early. But at the same time, there's some cases where you just want to let the model go. So for example, if cloud code is, uh, you know, it's writing tests for me, I'll just hit shift tab, enter auto accept mode, and just let it run the tests and iterate on the tests until they pass. Um, cause I know that's a pretty safe, safe thing to do. And then for some other tool…
AI assessment note: “prompt injection attack in the URL, and then the model writes malicious code to disk”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What are the best practices here? Because when it's non-interactive, it could run forever, and you're, you're not, you're not necessarily reviewing the output of everything, right? So I'm just kind of curious, how does, how is it different in non-interactive mode? What are the most important hyperparameters or arguments to set?
A Yeah, and for folks that haven't used this, so non-interactive mode is just cloud dash p, and then you pass in the prompt in quotes, and that's all it is. It's just the dash p flag. Generally, it's best for tasks that are read-only. That's the place where it works really well and you don't, you know, super have to think about permissions and running forever and things like that. Um, so for example, a linter that runs and doesn't fix any issues, or for example, we're working on a thing where we use quad in with dash P to generate the change log for quad. So every PR is just looking over the commit history and being like, okay, this makes it into the change log. This doesn't, um, because we know people have been requesting, uh, change logs, so we're just getting quad to build it. So generally, non-interactive mode, really good for read-only tasks. For tasks where you want to write, the thing we usually recommend is pass in a very specific set of permissions on the command line. So what you can do is pass in, ah, dash, dash, allowed tools, and then you can allow a specific tool. So for example, not just bash, but for example, git status, um, or git diff. So you just give it a, a set of tools that it can use or, you know, edit tool.
AI assessment note: “pass in a very specific set of permissions on the command line”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q much it costs to do these things, and they'll do a task and they're like, oh, that costs 20 cents. I can't believe I paid that much. How do you think, going back to, like, the product side too, is like, how much do you think of that being your responsibility to try and make it more efficient versus that's not really what we're trying to do with the tool?
A We really see quad code as, like, the tool that gives you the smartest abilities out of the model. Um, we do care about cost insofar as it's very correlated with latency, and we want to make sure that this tool is extremely snappy to use and extremely thorough in its work. We want to be very intentional about all the tokens that it produces. I think we can do more to, like, communicate the cost with users. Um, currently we're seeing costs around, like, Like, six dollars per day per active user, and so it's, like, it does come out to a bit higher, um, over the course of a month in Cursor, um, but I don't think it's, like, out of band, and that's, like, roughly how we're thinking about it.
AI assessment note: “we do care about cost insofar as it's very correlated with latency”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q YouTube, you can kind of see what the environment looks like. It's kind of like a cyberpunk industrial game. I don't know how else you guys would, would describe it. Let's start from the game hardness. Was it easy to plug a model into the game? Like, does it already have like an API? I know Minecraft has very good support for kind of like shell commands and all that.
A So It does to an extent. Exactly the same as Minecraft. Factorio exposes a low level Lua API to interact with the game. The, the way you mod Factorio traditionally is you, you know, you write some Lua scripts, you hook that into the game, and then when events happen, your script runs. But this doesn't really work for doing what we were interested in, which was kind of running this at scale and getting, you know, a cluster of servers running. So we had to develop a, an entirely different, uh, way to mod the game by exploiting The admin console of multiplayer Factorio servers. So, you know, you want to run a Factorio server with your friends, want to use the admin, and as an admin, you have a little console where you can ban players. You can, you can change the permissions of different players and just basically modify the state. And what we realized is we could hook into this admin console over the network using a protocol called Archon, uh, you know, carried by TCP. And this basically enables us to execute Actions in the game remotely and, and kind of shard them across dozens or hundreds of factorial servers, which is necessary if you want to do training or any of the other workloads you might need at scale. So going back to your question, you know, what's the harness like? It was very easy to get a low level harness working, you know, tell the character move one step to the le…
AI assessment note: “It does to an extent. Exactly the same as Minecraft. Factorio exposes a low level Lua API”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q So then that somehow in 2024 ish turned to E to B?
A 2023, I think March 23. And we were pretty burned out. Thomas and I, we were working from Prague, from the Czech Republic, from my apartment. I like nothing was really moving, no growth. And GPT 3.5 came out like really first model kind of good ish with Cogen. So we took a, like, let's, let's take a 10 days break, two weeks break from, from DevBook. And because, like, everyone was trying things with AI, it was very clear, like, this is something, ah, where the future might go. So we wanted to just, like, from out of curiosity, try things out. We wanted to build, like, a AI, like, DevIn, kind of, like, thing. The first idea we had, had was, like, let's automate our work, ah, because with every project we were starting, There was a set of tools you always want to, want to integrate in your backend, like Stripe, like for size business, like Stripe, analytics, Slack notifications, emails, uh, sending out emails. And so we gave the agent tools to run code and we needed some kind of sandbox. We were like, yeah, that's good coincidence. We have sandbox from, from DevBook. And, uh, we posted about it on Twitter. Basically the agent actually Pulled GitHub repository, wrote code, started the server, tested everything, and at that point, I think we deployed it to Railway. Railway had, like, the best DX, and it was easiest to just plug it, plug it in, into the agent. I tweeted about it, an…
AI assessment note: “2023, I think March 23. And we were pretty burned out.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Like when were the models ready for people to start using it? Just, you know, you demoed at one of our first events as well, and I think there's always this lag between the infrastructure that you build and like the capabilities of the model. When did you go from just code interpreter to start saying, okay, now it's time to do computer use, now it's time to do RFT?
A That was probably end of 24, uh, start of 25, so when you look even at our data, like use, 24 we are growing, we are growing good, but 25 is like up to the right, and so it feels like at 24 people are like figuring out these agents, um, and, and building them and trying them, and 25 is everyone moving them into production and finding more and more use cases. Around end of 24, start of 25, we started seeing things like using Sandbox for Uh, reinforcement learning type of use cases or using sandbox as a for computer use, which was very interesting when Anthropic launched their computer use. We had like a desktop version of a sandbox that was sitting like in our GitHub repository for six months. We were like, okay, this is probably interesting, but no model can actually use it. So, uh, when Anthropic announced it, it was like, oh, like we have something here. Uh, we can show you using, uh, with, with like Lavables and Blitz's type of products also, like, started using the sandbox for more than just, like, run code snippet, like data analysis, and then deep research agents. That, that has been something really big in the last few months.
AI assessment note: “That was probably end of 24, uh, start of 25”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Yeah. I think the other comparisons also would be, you said magic, like there's magic.dev and poolside, poolside also trying to go to market with, with extensions. But I guess those are training proprietary models, right? They're, they're investing a lot of money in that, and you've just decided to do no custom models here.
A Yes, I would say for now. So for us, there's always the question of for every feature, do we use an off-the-shelf model or an external model, or do we train our own? I think for now, for agents, uh, except for the code-based understanding, we, we, we outsource the models. However, we believe that the cost of these, or let's say, we believe that the usage of agents is going to explode. And as a result, the cost is going to explode as well. And so We feel that there is room here for training custom, custom models to help with that. But yes, for now, building a product and going to market quickly, this is what we've prioritized. And there's clearly a lot of, a lot of demand for these things out there. So we're pretty happy with that decision.
AI assessment note: “Yes, I would say for now. So for us, there's always the question”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q How do you think about that when running, building agent.ai even? It's like, you know, instead of just choosing one, I could like literally just run across all of them and figure out which one is going to work best.
A I'm a big believer. So, uh, under the covers, when you build an, um, because the primitives are so simple, you have some sort of inputs. We know that what the variables are. Every agent that's on agent.ai automatically has a REST API that's callable in exactly the way you would, uh, you'd expect. Automatically shows up in the MC, MCP server. So you're able to invoke it in whatever form you decide to. And so my expectation is that in this future state, whether it's a human hiring, uh, an agent to do a particular task or evaluating a set of five agents to do a particular task and picking the best one for their particular use case, we should be able to automate that. It's like, I just want to try it. Um, and there should be a policy that the publisher, builder of the agent has that says, okay, well, I'm going to let you call me 50 times, a hundred times before you have to pay or something like that. Uh, we should have, Effectively like an audit trail, like, okay, this agent has been called this many times. We also have a kind of human ratings and reviews right now, and we have tens of thousands of reviews of the existing agents on agent.ai average is like 4.1 out of five stars. And all those things are nice signals to be able to have, but the, the kind of callable, uh, verifiable kind of thing I think is super useful. Like if I can just call, uh, give me an API that says here are …
AI assessment note: “Every agent that's on agent.ai automatically has a REST API that's callable”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. Do you want to talk about the chat.com? Thing. I would love just the back story. It's like, did you just call up Sam one day and be like, I got the domain? Did they kind of get back to you knowing that you had it?
A It's a, it's a good story. Back, uh, in the original ChatGPT days, uh, the first thought I had in my head, which lots of people had in their head, is that OpenAI is going to build a platform and ChatGPT is actually just a demo app to show off the thing. And there's been precedence for tech companies that have had, uh, you know, demo apps to kind of help normies understand the underlying technology. And even after the kind of the boost or whatever. So my original thought was, well, someone should actually create like an actual real product. And so I'm like, and that product should be called chat.com because GPT is not a consumer friendly thing at all. Like that's an acronym, uh, not pretty, it doesn't roll off the tongue. And so like, I'll build chat GPT because that was just a demo app back then. So I, you know, got chat.com. And then as it turns out, chat GPT is like a real product. And I was at an event here in San Francisco that Sam spoke at where he launched, uh, Plugins, I think it was the announcement at that time. Yeah. And that's the thing is like, I had sort of suspected, it's like, okay, things sort of be like, there's no way that OpenAI is going to launch plugins for ChatGPT if they were not thinking of it as an actual platform. It's not just about the GPT APIs. This is like a real thing. I'm like, crap. Like this violates the first rule of Dharmesh, which is don't c…
AI assessment note: “Back, uh, in the original ChatGPT days, uh, the first thought I had”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Because like, That's, that's where, that's where I think, like, people fall apart, which is, like, not sticking enough to, like, ok, like, how do you use it? What does it do? What doesn't it do? Yeah, go ahead.
A So first, uh, couple of links. Of course, all the code is under cloudflare slash agents. We have a whole bunch of documentation under developers.cloudflare.com slash agents, and this we are updating like on a daily basis. Shout out to Matt, like he's being like grinding on this like a lot. And we have a starter kit at Under cloudflare slash agent starter, which is, as is the style in JavaScript development, there's a little one liner that you can take, run it in your terminal, and it spins up a project that has not just the agent framework, but a little front end for it that has react and a little chat agent starter that lets you just take it and like, you can basically start adding tools into it and working with it. So let's just go into the code. Let's see what it looks like. As I mentioned, so there's a bunch of configuration stuff because it's not a JavaScript project if there's not 12 configuration files in it.
AI assessment note: “we have a starter kit at Under cloudflare slash agent starter”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q go one by one and we'll leave the Agents SDK towards the end. So Responses API, I think the sort of primary concern that people have and something I think I voiced to you guys when, when I was talking with you in the planning process was, is chat completions going away? So I just wanted to let it, let you guys respond to the concerns that people might have.
A Chat completion is definitely, like, here to stay. You know, it's a bare metal API we've had for quite some time, lots of tools built around it, so we want to make sure that it's maintained and people can confidently keep on building on it. At the same time, it was kind of optimized for a different world, right? It was optimized for a pre-multi-modality world. We also optimized for, kind of, single turn, text prompt in, text response out, and now with these agentic workflows, we, we noticed that, like, Developers and companies want to build, um, longer horizon tasks, you know, like things that require multiple returns to get the task accomplished and computer use is one of those, for instance. And so that's why the responses API came to life to kind of support these new agentic workflows. But chat completion is definitely here to stay.
AI assessment note: “Chat completion is definitely, like, here to stay.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q use each type of API? So I know that in the past, the assistance was maybe more stateful, kind of like long running, many tool use kind of like file based things and the chat completions is more, Stateless, you know, kind of like traditional completion API. Is that still the mental model that people should have? Or like, should you by default always try and use the responses API?
A So, so the responses API is going to support everything that the chat, it's at launch going to support everything that chat completion supports. And then over time, it's going to support everything that assistance supports. So it's going to be a pretty good fit for anyone starting out with OpenAI. They should be able to like go to responses. Responses, by the way, also has a stateless mode. So You can pass in store false, and that'll make the whole API stateless, just like chat completions. We're really trying to, like, get this unification story in so that people don't have to juggle multiple endpoints. That being said, like, chat completions, just like, it's our most widely adopted API. It's so popular, so we're still going to, like, support it for years with, like, new models and features. But if you're a new user, you want to, or if you want, like, existing user, you want to tap into some of these, like, Built-in tools or something, you should feel, feel totally fine migrating to responses, and you'll have more capabilities and performance than, than check completions.
AI assessment note: “it's going to be a pretty good fit for anyone starting out with OpenAI.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Dota in the beginning, and then everybody just focused on language. And I think now the pendulum is shifting back to RL. Is there anything specific, um, in the last six to 12 months that kind of made you decide, okay, now is the time to do it? Or is it just a matter of, you know, the team coming together at the right time and the market being ready?
A It was really that, so Giannis and I led a lot of the work for post-training and kind of RL check for Gemini, and Giannis being my co-founder, and when we shipped Gemini One, we just realized that the models, like, models that were basically at GPT-IV level or above, were capable enough as kind of starting points to then post-train, or, I mean, it's, I wouldn't even call it post-training, just training. This reinforcement learning. So it's really that, I don't know if you've been able, you would have been able to do the same stuff with an earlier model, like with a GPT-III or a GPT-II, but GPT-IV had this kind of base of intelligence that you could actually go from there. And by the way, this actually happened in AlphaGo as well, and some of these systems previously, where before you did reinforcement learning, they first did imitation learning on human games. And it was important that the human playing ability was sufficiently high for you to be able to then bootstrap on top of that and then train with reinforcement learning. That is, if you would have trained AlphaGo on very weak human gameplay, then for that first project, reinforcement learning wouldn't have taken you as far. Of course, AlphaGo then figured out transition to AlphaZero, where it was trained without human data whatsoever, but I think that's where the analogy breaks. I think In the era of language models, I do…
AI assessment note: “when we shipped Gemini One, we just realized that the models... were capable enough”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And then on the environment, you briefly mentioned the computer is the idea that, you know, you mentioned code, but do you want to be very computer use focused? You know, some people will say the browser is kind of like the new OS anyway. Uh, I'm curious your thoughts on where you want the agent to, to live in.
A Yeah, it's, it's a really good question. I mean, if you start kind of an autonomous agent company today, you might pick from one of two categories, maybe some others, but I think broadly speaking, you could either say I'm going to build browser agents. Or kind of computer use kind of agents more broadly, or build coding agents. And I think that we're, you know, our belief as a company is that the correct wedge in, the correct starting point to this entire problem is decoding agent, because it's already, you know, software engineering is already what I would call kind of ergonomic for a language model. That is to say, You have to design the problem that you're striving for to be compatible with sort of, um, what is intuitive for the intelligence that's working on it. And for humans, what's intuitive to us is, you know, using a mouse and keyboard and kind of geospatial reasoning, because that's just how we evolve. Uh, but language models never evolved. They were trained on the internet. And because they were trained on the internet, their priors for what they, it's intuitive to them is completely different for what's intuitive to us. And so A language model, for example, has no prior for a mouse movement. It, you know, it really never seen that on the internet, but it has a really strong prior for code. So out of the two categories, say web browsing and coding, coding is the only…
AI assessment note: “our belief as a company is that the correct wedge in... is decoding agent”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Like I noticed that you use the Ray one model in there. What is it? Is it.
A Yeah. You spotted that automatically. One thing that we did is basically we started exploring all the different models, and then, and really quickly realized, like, we have such a unique use case, right? So we have quite a large set of, like, certain functions or tools we want to call, but then also a very specific need for that. So at some point we decided, like, hey, let's go down the route of, like, fine-tune a model, and so we looked into all the various models we had, um, and then we picked, at the moment, it's gbd-for-o and gbd-for-o-mini, Which we basically did a fine tune to really optimize for our use case, and then basically shipping that in the app as, as Ray one and Ray one mini. So they are highly optimized for our function calling and basically make it like one faster, but also the most important part, more accurate because like you want to make sure that those things happen as best as possible.
AI assessment note: “we picked, at the moment, it's gbd-for-o and gbd-for-o-mini, Which we basically did a fine tune”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q you to start the company. Then what was maybe the initial idea, Mace? Because, you know, if you think about, if somebody told you that was the hugging face founding story, people might believe it. It's kind of like a similar ethos behind it. How did you land on the product feature today? And like, maybe what were some of the ideas that you discarded that initially you thought about?
A So the first thing we built, it was fundamentally an, an API. So nowadays people would describe it as, like, agents, right? But anyone could write a Python script, they could submit it to the CHI backend, and we would then host this code and execute it. So that's like the developer side of the platform. On their Python script, the interface was essentially text in and text out. An example would be the very first bot that I created. I think it was a, like, a Reddit news bot. And so it would first, it would pull the popular news. Then it would prompt whatever, like I just used some external API for like BERT or GPT-II or like it was a very, very small thing. And then the user could talk to it. So you could say to the bot, hi bot, what's the news today? And it would say, this is the top stories, and you could chat with it. Now, four years later, that's like perplexity or something. That's like the, Right. But back then, the models were, first of all, like, really, really dumb. You know, they had an IQ of, like, a four-year-old. And users, there, there really wasn't any demand or any PMF for interacting with the, with the news. So then I was like, okay, um, clearly no PMF for that, so let's make another one. And I made a bot which was like, you could talk to it about a recipe. So you could say, I'm making eggs. Like, I've got eggs in my fridge. What should I cook? And it'll say, yo…
AI assessment note: “So the first thing we built, it was fundamentally an, an API.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And from the outside, these labs kind of look like huge organizations that have this like obscure ways to organize. How did you get, you joined Anthropic? Did you already know you were going to work on like SweetBench and some of the stuff you publish or you kind of join and then you figure out where you land? I think people are always here to learn more.
A Yeah. I've been very happy that Anthropic is very bottoms up and sort of very sort of receptive to whatever your interests are. Um, and so I joined sort of being very transparent. Like, Hey, I'm most excited about code generation and AI that can actually go out and sort of touch the world or sort of help people build things. And, you know, those weren't my initial, uh, initial projects. I also came in and said, Hey, I want to do the most valuable possible thing for this company and help Anthropic succeed. And, you know, like, let me find the balance of those. So I was working on lots of things at the beginning, um, you know, function calling tool use, uh, and then sort of, as it became more and more relevant, I was like, Oh, Hey, yeah, like let's, it's time to go work on encoding agents. And sort of started looking at SweetBench as sort of a really good benchmark, uh, for that.
AI assessment note: “those weren't my initial, uh, initial projects. I also came in and said”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q How do we fix that? Are you supposed to fix it at the model level? Like how do I know what prompt I'm supposed to use?
A Yeah. And I'll say this was a very small effect size. And so I think this is not, I think this isn't like worth obsessing over, but I would say that as people are building systems around agents, I think the more you can separate out the different kinds of work the agent needs to do, the better you can tailor a prompt for that task. And I think that also creates a lot of like For instance, if you were trying to make an agent that could both, you know, solve hard programming tasks and it could just like, you know, write quick test files for something that someone else had already made. The best way to do those two tasks might be very different prompts. I see a lot of people build systems where they first sort of have a classification and then route the problem to two different prompts. Um, and that's sort of a very effective thing because one, it makes the two different prompts Much simpler and smaller. And it means you can have someone work on one of the prompts without any risk of affecting the other tasks. So it creates like a nice separation of concerns.
AI assessment note: “first sort of have a classification and then route the problem to two different prompts”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q But it's not for your actual computer, right? Like the Docker instance is like runs in the Docker. It's not for...
A Yeah, it runs its own browser. I think, um, I mean, the main reason for that is one is sort of security. You know, we don't want You know, the model can do anything. Uh, so we wanted to give it a sandbox, not, not have people do their own computer, at least sort of for our default experience. We really care about providing a nice sort of making the default safe, I think is the, is the best way for us to do it. And I mean, very quickly people made modifications to let you run it on your own desktop. Uh, that's fine. Someone else can do that, but we don't want that to be the official anthropic thing to run. I would say also like from a product perspective right now, Because this is sort of still in beta. I think a lot of the most useful use cases are like a sandbox is actually what you want. You want something where, Hey, any, it can't mess up anything in here. It only has what I, what I gave it. Also, if it's using your computer, you know, you can't use your computer at the same time. I think you actually like want it to have its own screen. It's like you and a person pair programming, but only on one laptop versus you have two laptops.
AI assessment note: “Yeah, it runs its own browser. I think, um, I mean, the main reason”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q OpenAI part of the history. Exactly. So then you leave OpenAI in September, 20, 22. And I would say in Silicon Valley, the two hottest companies at the time were you and Langchain. What was that start like? And what did you decide to start with a more developer focus, kind of like a, AI engineer tool rather than going back and to do some more research on something else.
A Yeah. First, I'm not a trained researcher, so going through OpenAI was really kind of a, the PhD I always wanted to do. But research is hard. You're digging into a field all day long for weeks and weeks and weeks, and you find something, you get super excited for 12 seconds, and at the 13 seconds you're like, oh yeah, that was obvious. And you go back to digging. I'm not a trained Like formally trained researcher, and it wasn't kind of a necessarily an ambition of me of creating, of having a research career. And I felt the hardness of it. I enjoyed a lot of like that a ton, but at the time I decided that I wanted to go back to something more productive. And the other fun motivation was like, uh, I mean, if we believe in AGI and if we believe the timelines might not be too long, It's actually the last train leaving the station to start a company. After that, it's going to be computers all the way down. And so that was kind of the true motivation for like, uh, trying to go, uh, to go there. So that's kind of the core motivation at the beginning of personally. And the, uh, the motivation for starting a company was pretty simple. I had seen GPT-IV internally at the time. It was September, 20, 22. So it was pre-chat GPT, but GPT-IV was ready since, I mean, I'd been ready for a few months internally. I was like, okay, that's, that's obvious. The capabilities are there to create an in…
AI assessment note: “thesis was there's probably a lot to be done at the product level to unlock the usage.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q And what was the origin of it? How did you come up with the idea? Uh, how did you get people to buy in? And then maybe what were one or two of the pivotal moments early on that kind of made it the standard for, for these things?
A Yeah. Yeah. Chatbot Arena project was started last year in April, May, around that. Before that, we were basically experimenting in the lab how to fine-tune a chatbot open source based on the Lama-one model that had released. At that time, Lama-one was like a base model, and people didn't really know how to fine-tune it, so we were doing some explorations. We were inspired by Stanford's Alpaca project. So, we basically, yeah, grow a data set from the internet, which is called shared gpt data set, which is, like, a dialog data set between user and chat gpt conversation. And it turns out to be, like, pretty high-quality data, dialog data. So, we fine-tune on it, and then we try to, and release the model called wikunia. And people were very excited about it because it kind of like demonstrate open way model can reach this conversation capability similar to ChatGPT. And then we basically released the model ways and also do the demo website. The model. That, people were very excited about it, but during the development, the biggest challenge to us at the time was, like, how do we even evaluate it? How do we even argue this model we trend is better than others? And, like, what's the gap between this open source model and other proprietary offering? At that time, it was, like, GPT-FOR was just announced. It's, like, cloud one, right? What's the difference between them? And then after …
AI assessment note: “we quickly realized that people need a tool to compare between different models.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What's the initial product, just text to speech, or were you also doing kind of like a synthesizing of the content, refining it, or were you just helping people read through it?
A Before we did the IO announcement in 23, we'd already done a lot of studies. And one of the first things that I realized was the first thing anybody ever typed was summarize the thing, right? Summarize the document. And it was like half like a test and half just like, oh, I know the content. I want to see how well it does this. So as part of the first thing that we launched, um, it was called Project Tailwind back then. It was just Q and A. So you can chat with the doc just through text and it would automatically generate a summary as well. I'm not sure if we had it back then. I think we did. It would also generate the key topics in your document and it could support up to like 10 documents. So it wasn't just like a single doc.
AI assessment note: “It was just Q and A. So you can chat with the doc”