Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Right. And so the week we're talking, you at Box are releasing a number of different agents. Um, let me start this discussion by just asking you, what is an agent? Because it does seem like it's an overused term and, and even myself who I'm, I'm in this all the time. I don't fully have clarity on what that word actually means.
A Um, I, I think the, ah, I think we should anticipate that it's fully overused. It, it is now the new term of art for talking to a, an AI system that is doing work for you. So just, we will hear, this will be the main term that we use going forward as an industry. And not because it's a buzzword, but actually it's a, it's a useful term. It's a, it's a definable object that is doing automated work for you. That could be in some cases as simple as answering a question. Um, but I think most people in, in the tech industry would generally argue that it should be doing some degree of, of work and looping through the AI model multiple times, um, uh, to do that work, and so, uh, that could be everything from, you know, very clearly something like Claude Code, or Cursor has an agent, or Replit has an agent, where you give it a task like, build me a website that has these qualities, and it will go off and do, you know, weeks worth of human work, In 10 minutes, and that's an agent that is managing that whole process, looping through the model multiple times, keeping track of what it's doing, updating its memory in the process, and that's effectively an agent. So that's an agent in coding, and we're going to see that same kind of agent architecture emerge in law, in healthcare, in finance, in education, where you can deploy agents to go off and do work for you. And, um, and, and there'll b…
AI assessment note: “a definable object that is doing automated work for you”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q public sector, you saw a jump from 77 to 88% in accuracy for complex tasks. Healthcare saw a jump from 60 to 78%, and legal saw A jump from 57 to 69%, uh, uh, accuracy on complex tasks. Uh, that's pretty, pretty big. It seems like this model has, has almost been under hyped. Uh, can you talk a little bit about these, these jumps and what the significance is?
A I think, I think probably the, the main takeaway should be that, that the progress of these meaningful jumps that we've been seeing in AI coding Over the past couple of years where, you know, the model at best could do a couple lines of code You know, in a, in a kind of type ahead type format two, two and a half years ago in, in coding space. And now obviously people are giving the model a task of, you know, write me tens of thousands of lines of code for a full project. And, and we've just seen this incredible rate of progress and this March, uh, up toward, you know, more and more capability over time, uh, with, uh, within coding. I think that same trend is going to come to other Other now fields of knowledge work. And so, so this jump in sonnets model from four or five before six, I think represents an example of what happens when these models just get trained across more areas of knowledge work. What happens when they are getting better and better at reasoning capabilities that go beyond coding? What happens when they get better at using tools and deciding when to use tools? And that's what our complex work eval You know, is, is meant to represent is, is sort of how does it think through a problem? How does it decide it's got the right answer? How does it check its work? Um, and these models are getting much better at, at being able to deliver on that. So I think that'll be …
AI assessment note: “I think that same trend is going to come to other Other now fields of knowledge work.”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q here's my computer, have my files, take actions on my behalf. And, and honestly, they work better when you take the guardrails off and trust them to do things for you. Um, do you think we're like, again, for this product vision to work, that has to happen. Do you think we're in a place where it's feasible for people to give up that type of control to these bots?
A Well, so this is, this is where the diffusion, this general category is where the diffusion will be longer than, than where people in Silicon Valley think. So if you're in Silicon Valley and, you know, every tweet that you and I read, you know, that goes viral in, in the Valley is, is often It's coming from like a 10 person startup. They have, they have basically like, they started from a completely clean slate of, of the way that they work, that their environment, the tools they use, the data that they have, and they can just, they can build their organization around, around getting, uh, output from agents. And, uh, you go to the rest of the world, take a company that has, you know, 10,000 employees, been around for, you know, decades. Their data is in, Again, 2030, 50, a hundred different systems. The, uh, if you go and ask that company, um, where are your latest, you know, contracts for this client? It could be in five different places. If you go and say, where's the latest marketing campaign assets? It could be in 10 different places. If you say, where's the research for the new, um, uh, for that new breakthrough that you're working on, it could be in, you know, five different repositories. So the challenge is if you're in, if you now want to go deploy an AI agent in that environment, Uh, you can almost think about it like, like a new employee joining that company and that …
AI assessment note: “this general category is where the diffusion will be longer than”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q Is it actually, does it feel like real autonomous driving, or were you still, is there still fear for people's life when they're in there?
A Uh, well, those could be the same thing. So, um, so it, it might be that real autonomous driving, you still fear for everybody's life, because you're just like, I do not know how this works. Like, this is kind of alchemy. This is crazy, but, uh, it was, it was definitely, you know, very crazy. It was a very, you know, kind of relatively boring suburban, um, kind of trip, but, um, uh, but it was, uh, like there was just zero need to ever, ever interject. I mean, you have to, to kind of show that you're, you're still paying attention. Um, but, uh, but I mean, it just, it shows again, we're like the past, Year and certainly for the next couple of years, you get the sense that we're going to see hundreds of these, these like, like early previews about the future, um, which is just pretty exciting. Like, uh, I, I mean, I, I've just never seen a period where, you know, in any given week you could see two to three things, which are just like, obviously that's going to be the future. Maybe it doesn't work perfectly right now, but, but it's like, there's nothing that is, is stopping it from working perfectly in a world of more compute, And, and just more breakthroughs on, on, on the, uh, on the models themselves, and that is kind of where we're at right now.
AI assessment note: “there was just zero need to ever, ever interject”
Answered raw tape
D 4 · C 3 · P 3 · Cm 2 3.15
Q mean, if I didn't create it, the company on the receiving end of this, though, is anthropic. And, you know, you talked about these mythical Capabilities. They called the model mythos. They put in the documentation that like it broke out of its containment and wrote the engineer while he was having a sandwich in the park. Is it that surprising that this is one of the downstream impact? Yeah.
A But if you put that in your announcement blog post, you know, people might be able to kind of extrapolate and get pretty, pretty, pretty scared of things. I think it's interesting. So, um, you know, on the anthropic front, first of all, I have I have a huge amount of respect for the entire kind of stack of researchers and policy folks across AI. I happen to have disagreements with some of the, the categories, but, but I think there's a deep, let's say, if you were, if you imagined a continuum of the most, like, you know, if, if you, uh, uh, if you kind of had like, like the most, I, I, I mean, it's only in like a polite way. It will sound impolite, but like, I mean, like, like if you're the most doomer on one end of the spectrum and the most like, like accelerationist on the other end of the spectrum, here, here's kind of the, the views, the most doomer, Uh, possible is, was afraid of like GPT three and GPT three was going to like, you know, sort of accelerate and, and, you know, kind of achieve some kind of unstoppable continual improvement. Um, and, you know, the acceleration that says like, we need like fable 20 as soon as possible. Right. So that's, that's sort of the continuum. I'm probably like, I, you know, maybe two thirds up to the acceleration is kind of side of things. But if you were on the, on the Doomer and I, I, I'm trying to say the polite version of Doomer, lik…
AI assessment note: “if you put that in your announcement blog post... people might be able to kind of extrapolate”
Redirected raw tape
D 2 · C 4 · P 3 · Cm 3 3.00
Q But do you think this is evidence for or against?
A I think the smartest people on the planet have two totally different views, and so I am, uh, I'm not gonna get in, in the middle of that one. I mean, clearly you have people like Ilya where, you know, it's rumored that he's working on a different architecture or, and, and, and maybe a different path, um, and then obviously you have other people that are, are, you know, let's just throw more compute and data at the problem. I think you can start to sense actually as an industry that the, the, the AGI term has actually kind of gone into the backseat and obviously more of the conversation is around super intelligence. Um, and I think there's more and more comfort around this idea that actually the race really is just how do we build intelligence that far exceeds a human and what will the economic and, and, you know, kind of societal benefits be of just even accomplishing that. Which are massive. And, um, I have always sort of found the AGI thing to be, you know, particularly squishy as a concept. Um, I, I, in the B to B world, I, I deal way more with just like utilitarian concepts. And so super intelligence and this idea of we have AI that will far exceed a human, like that, that alone is enough of a breakthrough to be shooting for. And I think what you're seeing with scaling is we will be able to certainly accomplish our collective definition of super intelligence. With the curre…
AI assessment note: “I'm not gonna get in, in the middle of that one.”