Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q That was, uh, awesome talk about all the scaling laws, and recently, Anthropic just launched Cloud Four, which is just available. Curious, uh, how does it change what is possible as all these model releases keep compounding for the next 12 months?
A I think that, ah, we'll be in trouble if it's 12 months before, before an even better model comes out, but, ah, I guess, ah, a few things with, with Claude IV. I think that, with Claude III. VII's Sonnet, Uh, it was already really exciting to use 3.7 for coding, but I think something that everyone noticed was that 3.7 was a little bit too eager. Um, sometimes it just really wanted to make your tests pass, um, and it would do things that, that you don't really want. Uh, there are a lot of like try accepts, things like that. Um, so with cloud four, I think that we've been able to improve the model's ability to act as an agent Specifically for coding, but, but in a lot of other ways, for search, for all kinds of other applications, um, but also improve its supervision, the sort of oversight that I, I mentioned in my talk, so that it, ah, it follows your directions and hopefully improves in, in code quality. I think the other thing that we've worked on is improving its ability to, ah, save and store memories, and we hope to see people leveraging that, because Claude IV can blow through its context window with a very complex task, but can also, ah, store memories as files or records, retrieve them in order to sort of keep doing work across many, many, many context windows. But I guess finally, I think the picture that scaling was paint is one of incremental progress, and so I think …
AI assessment note: “improve the model's ability to act as an agent Specifically for coding”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q human approval before they would send the reply for a customer. But one thing that has changed just in the spring batch, I think a lot of the AI models are very capable to do tasks end-to-end, to your point of that, which is, ah, remarkable. Founders are selling now directly replacements of full workflows. How have you seen this translate to what you hope the audience here will build?
A I think there are a lot of possibilities. Basically, it's a question of What level of success or performance is, is acceptable. There are some tasks where getting it sort of 70% right is, is good enough, and others where you need 99.9% to, to deploy. I think that, honestly, I think it's probably a lot more fun to build for use cases where, ah, 70 80% is good enough, because then you can really get to the frontier of what AI is capable of, but I think that we're sort of pushing up the, The reliability as well. So I think that, ah, we will see more and more of these tasks. I think that, ah, right now, human AI collaboration is, is going to be the, sort of, most interesting place, because I think that for the most advanced tasks, you're really gonna need humans in the loop. But I do think in the longer term, there'll be more and more tasks that can be fully automated.
AI assessment note: “a lot more fun to build for use cases where, ah, 70 80% is good enough”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Now, other question is, you have an extensive training as a physicist, and you're one of the first to really observe this trend with scaling laws, and it probably comes from being a physicist and seeing all these exponentials that happen naturally in nature. How has that training come about with, ah, being able to perform like the best research in the world with, with, with AI?
A I think the thing that was useful from a physics point of view is looking for the biggest picture, most macro trends, and then trying to make them as precise as possible. So I remember meeting, like, kind of brilliant AI researchers who would say things like, learning is converging exponentially, and I would just ask really dumb questions like, are you sure it's an exponential? Could it just be a power law? Is it quadratic? Like, like exactly how is this thing converging? And it's a really dumb kind of simple question to ask, but basically I think there was a lot of fruit to be picked and, and probably still is in trying to make the big trends that you see as precise as possible because that, I don't know, it gives you a lot of tools. It allows you to ask like, What does it really mean to move the needle? I think with scaling laws, the, the holy grail is finding a better slope to the scaling law, because that means that as you put in more compute, you're going to get a bigger and bigger advantage over other AI developers. Um, but until you've sort of made precise what the trend is that you see, you sort of don't know exactly what it means to beat it and, and how much you can beat it by and how to know Systematically whether you're, you're, you're achieving that end. So I think those were kind of the tools that, that I think I used. It wasn't necessarily like literally applying …
AI assessment note: “useful from a physics point of view is looking for the biggest picture”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Which is very much, uh, the Jevons paradox. As intelligence becomes better and better, people are gonna want it more. Not that it's driving the cost down, which is this irony, right?
A Yeah, absolutely. I mean, I think that, uh, yeah, that's, that's certainly something that we've seen, that there are certain, uh, certain points where AI becomes accessible enough. That said, um, I think as AI systems become more and more capable, um, and can do more and more of the work that, that we do, it's going to be worth it to pay for, uh, frontier capability. So I think it's a question that I've, Always had and continue to have is kind of like, is all of the value at the frontier, or is there a lot of value with kind of cheaper systems that aren't quite as capable? And I think this sort of time horizon picture is maybe one way of thinking about this. I think that you can do a lot of very simple bite-sized tasks, but I think it's just much more convenient to be able to use an AI model that can do a very complex task end-to-end, Rather than requiring us as humans to sort of orchestrate a much dumber model to break the task down into very, very small slices and put them together. So I do kind of expect that a lot of the value is going to come from the most capable models, but I might be wrong. It might depend, and it might really depend on the capabilities of AI integrators to sort of leverage AI really efficiently.
AI assessment note: “it's going to be worth it to pay for, uh, frontier capability.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Are there specific areas that you think a lot more builders could go into and build with these new models? I mean, there's a lot that has been done, let's say, for coding tasks, but what are some tasks that have a lot more greenfield that are just getting unlocked right now with the current models?
A I come from a research background rather than, uh, rather than business, so I don't, I don't know that I have anything very, uh, very deep to say, but I think that, like, in general, any place where, um, it requires a lot of skill, um, and it's a task that mostly involves sort of sitting in front of a computer, interacting with data, I think finance, uh, people who use Excel spreadsheets a lot, um, I think I, I expect law, although maybe, maybe, maybe law, ah, is, is, is more regulated, requires more, ah, more, more expertise, um, as a stamp of approval, but I think all of these areas are probably green field. I think another that, that I sort of mentioned is, how do we integrate AI into existing businesses? I think that, like, when electricity came along, there was some long adoption cycle, and, The very first, simplest ways of, say, using electricity weren't necessarily the best. You wanted to not just replace a steam engine with an electric motor, you wanted to sort of remake the way that factories work. And I think that probably leveraging AI to integrate AI into parts of the economy, um, as quickly as possible, I expect there's just a lot of, a lot of leverage there.
AI assessment note: “I think finance, uh, people who use Excel spreadsheets a lot, um, I think”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q Now, one aspect about, uh, scaling laws, they've held for over five orders of magnitude, which is wild. This is a bit of a contrarian question, but what empirical sign would convince you that the curve are changing? That maybe we're getting off the curve.
A I think it's a really, I think it's a really hard question, right? Because I mostly use scaling laws to diagnose whether AI training is broken or not. So I think that, uh, Once you see something and you find it very, it's a very compelling trend, it becomes very, very interesting to examine where it's failing. But I think that my first inclination is to think if scaling laws are failing, it's because we've screwed up AI training in some way. Maybe we got, ah, we got the architecture of the neural network wrong, or there's some bottleneck in training that we don't see, or there's some problem with Precision and the algorithms that we're using. So I think it would take a lot to convince me at least that scaling was really no longer working at the level of the sort of these empirical laws because so many times in my experience of the last five years when it seemed like scaling was broken, it was because we were doing it wrong.
AI assessment note: “it would take a lot to convince me at least that scaling was really no longer working”