The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Sholto Douglas argument clarity score 4.5/5 from 13 exchanges on raw tape · average scores: directness 4.7 · coherence 4.8 · precision 4.4 · compression 4 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
13exchanges match
13on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So, um, so Google noticed you, and then what happened?

A Um, so Google noticed me, uh, I started at Google, I think like a month before ChatGPT or something like this, so it was actually, it was a fantastic time to start at Google, because the entire company was suddenly forced to react, um, instantaneously and, and, and compete with the Gemini program. Um, so, uh, it meant that there was this Gap of, I guess, like the typical, uh, command structures and everything were not well suited for that particular battle. You know, it wasn't a pre-existing org. You know, Gemini was sort of forged out of the foundation, out of the merging of, uh, of brain and deep mind. Um, it, it meant that There was just a huge gap in terms of agency, really, of figuring out what we needed to do, doing it as fast as possible, organizing people together to, to work on important things. Um, and so I ended up, uh, one, getting the chance to develop a lot of taste by working closely with people, um, in those like early months of Gemini. Um, But two, also quickly got the opportunity to step up, uh, and, and to lead various parts of this. So one example of this is we just didn't have an inference stack that was, uh, at all, like, sensible for the modern world of LLMs. Um, and so we had to notice that, design one from scratch. A lot of the things you now see, uh, in, you know, like, the sort of SGLangs and stuff of the world are, like, things that we, um, had to de…

AI assessment note: “quickly got the opportunity to step up, uh, and, and to lead various parts”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Is this your haiku somewhere? Maybe walk us through the differences between those models.

A Yeah, so we, Release models along three categories, uh, three tiers. So Opus, which is the smartest model, Sonnet, which is, uh, the mid-tier model, and Haiku, which is the, the fastest, cheapest model. Um, one of the interesting things about, uh, this most recent release is actually Sonnet is smarter than Opus. And this has happened before. In fact, this happens last year. It's a reflection of fast progress because, uh, it, it is cheaper to train, uh, you know, mid-tier models than large models. And so what happens is that you end up doing a lot of progress on smaller models. Eventually you need to choose when to scale up and, and sort of get the benefits of scale in a model. Often you make progress fast enough that your mid-tier model is, is, is super, is like great anyway. Um, and, and it's actually better than the large scale up model that you did previously. Uh, and I think you, I think this is also a little bit of a reflection of the reinforcement learning, um, paradigm where you can, you can take a model and you can train it, uh, and, you know, it is, uh, and extend it with reinforcement learning, basically. So that allows you to take a, a mid-tier model and make it as good as a larger-tier model of, of six months ago or three months ago.

AI assessment note: “Opus, which is the smartest model, Sonnet, which is, uh, the mid-tier model, and Haiku”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Yeah. Let's unpack the 30 hour aspect, which is, which is fascinating. So, uh, first of all, to, to just ground it, um, for people. So this is a computer use, uh, just coding, I think. Just, just, just, um, coding. So what does the agent do for 30 hours? It's just clicking on stuff?

A Yeah. Uh, it is there. It's, it's reading files, and it's writing code, um, and running tests. Uh, so what, in exactly the same way that a human would, um, it, uh, Basically, you can think of the model as running in a loop where it can constantly decide what to do. People often mention something called tool use, and tool use is the ability to, uh, well, I mean, it's in the name, but in this case, it can use things like tools, like read file, write file, et cetera, um, or run code in the terminal, uh, and it is sitting there in a terminal on a computer in a loop, just constantly looking at the current code, deciding, oh, well, It can't quite do this yet, so I'm going to work on that next. It's often making plans, um, particularly to run for, uh, you know, 30 hours. One of the things that we're pretty happy with about the recent launches, we've finally taught the models to, uh, to use what's called memory. Um, and so, and we've built that into the agentic, uh, harness. So it's able to create a markdown file of to-dos and things that it thinks are important to do, uh, check them off and work on them and check whether they've been completed. There's almost this, like, Self-verification loop. One of the things that people were worried about, um, with language models over, like, I think a year ago or so, was that they would fall off track. Like, they wouldn't be able to self-correct,…

AI assessment note: “It's reading files, and it's writing code, um, and running tests.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Okay, so maybe let's talk about, uh, progress at a more abstract level, uh, but grounded in, in, in, in 25. So, a big part of the discussion seems to have been the evolution from a focus on pre-training, Uh, to RL, which we touched upon a couple of times. Talk about, uh, the impact of RL, and why is RL such a big part of the conversation today?

A For those listening, a good way to, like, understand at a high level of pre-training in RL, pre-training is like skim reading every textbook in existence, and RL is like doing the work problems and getting feedback on whether you were wrong or right. And there are actually a lot of things that you can only learn via RL. And a good example of this is the skill to say, I don't know, in response to a question. Because in pre-training, remember, you're modeling the, you know, you're trying to predict what, what text is going to come next in the, you know, all of these, you know, textbooks, the entire internet in the world. Um, and so, The only reason you would say, I don't know, on, you know, as a pre-trained model, is if you think the character that you're modeling in the text would say, I don't know. Like, if it's a likely completion, right? Not whether you, in fact, don't know, but whether you think that the, like, the sort of player that you've pulled from this cast of characters that you could model, um, would say, I don't know. Whereas in reinforcement learning, you could, in theory, set up a battery of tests where there are things the model knows and things the model doesn't know, and you could reward it for correctly answering, uh, things it should know and penalize it for, uh, for, for falsely answering when it doesn't know. And what it will then learn to do is it will lea…

AI assessment note: “there are actually a lot of things that you can only learn via RL.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yep. Got it. Uh, so closing the loop on, on something that you mentioned earlier, uh, why is Anthropic so focused on coding?

A Yeah. Uh, we're really focused on coding for two reasons. The first one is that we think it is the, uh, how should I say? It's the thing that will allow us to Assist ourselves in AI research faster. So, um, there's this, this notion of like automating AI research, right? And, and that work. Um, we think that one of the most important signals of whether or not we are, uh, basically the speed of takeoff, the speed of progress is driven by how much AI is able to assist, uh, AI research. Um, and so pre-fetching this is really important, we think. Um, secondly, we think it's the nearest term tractable, uh, problem domain, um, to, In terms of economic impact. Um, for Anthropic to be a viable research program that can research, uh, the things that we think are important requires economic, uh, like return. Um, and coding is, uh, a huge market full of people who are really, really, really like keen early adopters who love trying and switching things, um, who are really excited to play with new tools. Um, it's, there's massive, massive demand. There's dramatically more demand for software in the world than there is You know, good software. Um, we've seen that in every previous, uh, iteration of, um, you know, compilers and, like, general, like, web abstractions and so forth. Like, there's just a booming demand for software. Um, and so, like, you know, uh, I mean, yeah. Basically, the mod…

AI assessment note: “we're really focused on coding for two reasons”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And maybe to continue making this super educational, um, how do, uh, test time compute and RL overlap?

A Yes. One way of thinking about this is test time compute is doing a lot of reasoning, and then RL is the feedback signal on whether or not that reasoning was right or wrong. Uh, and so, Uh, test time compute is a way of answering questions that are hard for you to answer. Let's say I, I ask you a question that you just know off the cuff of your, like, you know, off the back of your hand, basically. Um, it's not like from a field that you really know, or whatever, a heuristic that you've already done. You've already baked that in, like, your muscle memory, so to speak. Um, But for something which requires you to really think and really learn, like when you're first doing math, if I ask you a basic times table right now, you can say that off like this, but if you're a kid, you have to like do out the math and all this kind of stuff. You need to like do the reasoning chain to learn it, and then you get feedback on whether it's right or wrong. Uh, so test time compute lets you do harder problems than you can currently do, than you can like currently do off the cuff, and RL Then allows you to sort of distill that back into the model. It's almost like a ladder. You can like constantly do slightly harder problems because you're learning strategies to, to do harder and harder and harder problems.

AI assessment note: “test time compute is doing a lot of reasoning, and then RL is the feedback signal”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So you think people don't realize, uh, you know, it's, it's always interesting, right, because, uh, You know, reading, uh, stuff like online in the, the last, you know, three, four months. It's like this theme of the, you know, we've reached a plateau, but basically saying the opposite, right? Like we, we are in an exponential curve and many people don't realize that it's the case.

A Exactly. And I mean, people have said that we've, we're hitting a plateau every month for the last three years. Um, and if you look at what we've come over the last three years, it's incredible. Uh, I think that. One other thing that makes me think, God, we're not anywhere close to a plateau, is I look at how these models are produced, um, and every part of it could be improved so much, like, it is a primitive pipeline held together by duct tape, and the best efforts, and elbow grease, and late nights, and like, God, it's, actually, I remember, um, uh, I don't know if this is a good analogy or whatever, but I remember I went sailing with a couple of friends a few months ago, And the boat was so well designed. It was just like clearly the product of, you know, like millennia or like, you know, centuries of, of like accumulated human design and effort, and I was like, wow, like, This is, this is what it feels like to be in a, uh, sort of the, the accumulation of a lot of human effort, right? It's actually pretty hard to beat today's best sailboat designs. Um, but when I look at an LLM training pipeline, it is two and a half years of best effort, last minute, desperate effort. Um, and, and there's just so much room to go on every part of it.

AI assessment note: “God, we're not anywhere close to a plateau, is I look at how these models”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Which, uh, is, is one of those important words in, in 25. What does taste mean when it comes to AI research?

A Yeah. I had a really interesting discussion, ah, about this with a, with a biology friend. We were comparing tastes across, like, biological research and ML. Um, I think one of the most important things is mechanistically understanding, ah, exactly what you're trying to do, and, and having an important simplicity regularizer. When you think about taste in ML, it's often, ah, It's the crucial ingredient that allows you to decide what goes into your large, ah, like, training run when you have imperfect information, ah, because we can study very deeply what, ah, what is, like, what the impact of an architectural change is, right? But past a point, past a certain level of scale, you have to guess whether or not the, Impact of that change will compound with other ones, whether it will, you know, conflict, um, because you can't test like your full scale run, right? End times. You only have one shot at that. And so a lot of taste comes from, uh, being able to make good inferences about, do we think that this ultimately, uh, like will, will sort of return, uh, deliver increasing returns to scale. Um, it also comes down to, do I think this direction of research is worth pursuing? Because often our baselines in ML are so well tuned. That, uh, it's very hard to beat them, even with what is, like, theoretically a better method, because there are so many small tricks that are required to, t…

AI assessment note: “a lot of taste comes from, uh, being able to make good inferences”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What do you make of the, you know, the counter thesis of the, again, Rich Sutton or Yama Khan, um, that seem to be saying that, like, a different approach is needed, or RL only. What do you make of that debate?

A Yeah. Um, I think that it's true that our models don't learn anywhere near as efficiently as humans do, right? Uh, They, they take, you know, a thousand lifetimes to, to learn, but this is, I think fine because they can live those thousand lifetimes, whether in simulations or doing, you know, a job at a thousand firms and so on and so forth. Um, I think that the, maybe, maybe I would disentangle. There's two arguments. One is like architecturally the transformers are like insufficient. Um, I don't think that's true. I think we haven't yet really found anything that transformers haven't been able to model provided sufficient data and sufficient compute. Um, I think RL as an objective is pretty powerful one. Rich Sutton is actually a big fan of RL as an objective. He just thinks we're actually encoding too many priors in with pre-training and this kind of thing.

AI assessment note: “I don't think that's true. I think we haven't yet really found anything”

Answered raw tape D 5 · C 4 · P 5 · Cm 4 4.55

Q Um, but is, does that suggest that, uh, being great academically and being a great, uh, anthropic researcher are two different things, so you need slightly different qualities?

A I think they're very highly correlated, but I think The signals that are usually used to gate academia are, like, there are dramatically more people that satisfy the criteria of being really effective than there are that have the correct signals that would then, like, enable them to progress to the next stage in academic career. For example, uh, if you're here in the US, you end up doing, as an undergrad, research that can get you in Europe, so ICLR paper, uh, whereas in Australia, that just isn't the case, right? Uh, I remember Peter Avil actually once visited our lab in Australia, And ask people to put their hands up if they were going to Europe, and no one put their hands up, not even the PhD students. Um, so it means you don't have, again, that mentorship aspect that is so, that is so important, and so you don't get a chance to develop problem taste, uh, on the things that mattered, um, and therefore you don't have, like, the correct signals that, that indicate you would have high potential for, for academia. I actually think that, right now, a lot of the signals we look for aren't, you know, traditional, uh, PhD or anything like this. I mean, this is obviously, like, very useful, uh, but, The fastest route, or like, the most immediate one is whenever we see a really good blog post where people have, like, done incredible amount of work, um, in an independent fashion, it's …

AI assessment note: “I think they're very highly correlated, but I think The signals that are usually used”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q And, uh, to the taste, uh, discussion, uh, the art versus science part of, of this. So, uh, are you, Saying that, um, uh, at least in terms of, like, anticipating what, how the, the, the, the training run may go, it's more intuition than actual numbers?

A So you can, you can do actual numbers up to a point. Like, the way to sort of illustrate or think about this is you are testing a system at multiple levels of scale. And actually, the analog to biology was, you know, if you think about it, you might test a new therapeutic in a cell and in mice and in model organisms. But that's no guarantee that it will work in a human, right? So you test across multiple different scales and multiple different model organisms, and it seems to work in, you know, basic single-cell bacteria, seems to work in mice, maybe works in monkeys. That gives you a lot of indication it's gonna work in humans, but it's not a guarantee. Um, so that, at that point, you need to understand the underlying mechanisms of how does this thing work? Like, what receptors is it binding to, and, and so forth. In ML, it's exactly the same, right? You have your different model scales, and you figure out, well, okay, it's, it's delivering benefits of these model scales, and I think it should work, because, like, mechanistically, I, I understand what this is doing to the learning dynamics of the model. And then, uh, and then you can have confidence that it's gonna work. But if it's like, ah, it's a hack, we don't really understand how it works, and it's really complicated, and introduces all this stuff in the code, then.

AI assessment note: “you can do actual numbers up to a point.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Yes. So, yeah, just to close on that last theme, um, so awesome, all very exciting, uh, what do we all do? How do we prepare for this world that seems to be around the corner?

A Yeah, I think the most actionable piece of advice is keep planning for a world where you as an individual have more leverage, right? Right now, I can Use two coding agents to do the, like, twice the work that I could have done before. Um, if coding agents progress in the way I've been saying, in a year or two, you'll be able to, you'll be able to manage a team, basically, that works 24 seven, um, for you doing work. I think we should expect in the digital domain for individuals to get dramatically more leverage over the next couple of years. Um, I think then many, like, incredibly important problems we're going to track with, like, our world is so imperfect in so many ways. Um, people still live in dramatic poverty, you know, health and medicine is unsolved, housing is, you know, completely unsolved. Like, the world could be a million times better in so many different ways, and what I hope is that people take, you know, initially models giving us leverage over the digital world, um, and then hopefully models giving us leverage over the physical one through robotics, uh, to, To, like, dramatically improve it.

AI assessment note: “the most actionable piece of advice is keep planning for a world”

Redirected raw tape D 2 · C 5 · P 4 · Cm 4 3.70

Q Going back to the jump in, uh, performance from last year's models, or even this year's models, or even actually Sonnet to, like, 4.5, uh, again, to the, to the, uh, point about the, the pace of progress accelerating. What, what were some of the, some of the breakthroughs?

A That I can't really talk about. Yeah, I, I mean, I think, it's, like, important to recognize that it's, like, not one individual breakthrough, really. Um, It is the continuous application of lots of different things across the entire stack for many people. Um, and It's like mostly just a function of compute in, in many ways. Um, there are, there are obviously individual breakthroughs, but fundamentally progress has been pretty, pretty smooth. Like on the meter eval, if you look at progress over the last two years, you can plot it with a straight line, right? And so similar to, you know, Moore's law of the past and this kind of thing, um, even Moore's law is made up of lots of individual improvements. It's not any one critical breakthrough. It's more the accumulation of a huge amount of work, um, in an environment where there's a sort of exogenous force of, of compute pushing progress forward.

AI assessment note: “That I can't really talk about. Yeah, I, I mean, I think,”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.