Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q kind of like the multiplayer mode? I think subagents is like, Single players split up the context. Yeah. And the multiplayer coworker is like, my colleague has some file on their machine that I want to know about, or I want to know how their task is going to then update my thing. Like, is that interesting? Is that something that makes sense for you to build or for like?
A It's like super interesting to me. It, it almost goes back to like some of the scaffolding where I'm like, okay, are we going to be, end up, are we, will we end up building scaffolding that will just go away? And like a question I have here is, At what point do we just assign these things like their own Gmail account? We'll just give them their own like Slack handle, and then they will just like use the same tools we humans use to interact with each other. You mentioned our finance people. They've been working pretty hard on very good office integrations. And I think for a while we've been like, we built so much tech around Claude leaving useful comments inside a Google Doc, and now it just does it, just like leaves a comment in your Google Doc and that's how you interact with it. Maybe like the similar thing where I still have open questions around what is the best interaction mode? Is it for us to be something super custom for co-work agents to talk to each other? Or is it, okay, let's just jump straight to the finish line and say, well, we're just going to give this thing, if you use Slack at work, we're just going to give this thing a Slack handle, and that's going to be the way it's like multiplayer capable.
AI assessment note: “It's like super interesting to me. It, it almost goes back to like”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q it's like fundamentally cloud code. We don't want to touch it. There's the cloud app. There's cloud in Chrome. I think you guys do something different in planning, but, uh, I've been talking with Tariq who is on the cloud code team and you guys are, he's like, no, we just exposed planning. Maybe you can clarify what are the major pieces that people should be aware goes into co-work?
A Like, okay, I think you basically have them. So, um, you can take planning more or less out. I think that's a few things that are really valuable in Kovac. Um, the virtual machine is probably the most powerful thing. So we currently run like a, we currently run like a lightweight VM and we put clock code into the VM and we do that for, for, um, a number of reasons. Safety and security is a big one, but even if you, even if you ignore for a second safety and security and you're just like, okay, Yolo, I want this thing to do whatever. It is quite powerful to give cloud its own computer. That is like generally a good idea. And in terms of architecture and UX and everything else that we've been working on Anthropic, it often is quite useful for you to like anthropomorphize, um, cloud aggressively and just be like, this is a person. What would you do if you give, if you had a person, right? And the analogy I've given my dad this morning, who is still like quite insistent on using chat, even for like coding things is if you were a developer, And your employer told you that you don't need a computer. They're just going to like send you emails with the code and you send emails with code back. Like that maybe worked for Petrars in the back, but that is not very effective. Um, so what we can do with the VM is because it's a, it's a Linux system. Cloud code has more or less free reign to …
AI assessment note: “The virtual machine is probably the most powerful thing.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Cool. I guess a couple parting questions. Uh, what's the future of Clockwork?
A I think we're still, we're still such early days. We're going to keep shipping things that we're going to keep shipping things that, um, we're going to keep iterating on this thing like pretty quickly, but which I mean, you can sort of continue to expect that every single week, there's going to be like a small new feature, if not a big new feature. Um, I'm going to continue probably to double down on your computer and like making you effective on your computer and making cloud effective on your computer. Um, we're starting to grapple as we talked about today, grapple more with the question of like, what does it mean? What does your computer mean? Does it have to be the one in front of you or like a VM on your computer or like a computer somewhere else? And then the third thing that I'm quite excited about is we're continuing to go up this hill climbing on slowly taking users who are used to asking questions and getting an answer to slowly teaching them to like step more and more away and that claw take over like bigger and bigger tasks and work both in time as well as in like scope. And I think you can probably see most of our investments on our feature releases to, like, work on both of those things. Like, the ability to do more on your computer, and then the ability to do it more independently for longer.
AI assessment note: “the ability to do more on your computer, and then the ability to do it more independently”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q curious, um, how much of the failure modes are the model intelligence versus like the usage of the end tool to put the intelligence in? Like the world planning is like a good example, right? It's like one thing is to come up with a plan. The other thing is like make a nice spreadsheet that kind of runs you through the plan. Like how have you seen that evolve?
A The thing that I grapple with a lot is that whatever scaffolding you come up with, I think we still have a bit of sort of like model overhang where the model is dramatically more capable than users end up using it for. And I think part of that is that we're just not getting the model, all the tools to do all the things that's theory capable of, right? That's like one thing. Um, however, whenever you do about the scaffolding, I'm sort of wondering at what point, at what point will that scaffolding go away? And like, How much you invest in figuring out what the right scaffolding is. It's kind of up to, it's a little bit of a bet, right? And one thing that I, as an engineer quite enjoy is that like working in Anthropik and working at a frontier lab, I maybe have a little bit more insight into what's coming, coming down the truth in terms of like, what's the next model? What is the model capable of? What is good at? What is it bad at? And I'm, I'm increasingly wondering is the right thing for us to like really invest too much in sort of these like scaffolding corrections where the model might otherwise Not misbehave, but just not do the thing that you want. Yeah. Or is it to just like give it as many capabilities as possible, try to make those safe so that the worst case scenario is like not as bad as it might be otherwise, and then just simply wait a second for the next model drop…
AI assessment note: “we still have a bit of sort of like model overhang where the model is dramatically more capable”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q you have skills, you have like, obviously the models, which is a big part. All these things kind of come together. Do you feel like that's a valid way to think about it where people should invest even more in kind of like these primitives? To rebuild on, or are you like recreating a lot of it each time because like things change and it's easier to rewrite than reuse?
A You know, I think, I think you're right. I think you're right that the holistic platform is really useful. And this is maybe a whole, like a somewhat contrarian view to a lot of people in AI. I actually don't think that the future is going to be hyper personalized software down to the point where everyone is running their own version. Like, I actually think it's going to be quite hard for one of us to have our own internal chat tool. And like, if I want to talk to you, like, How is that going to work? Right. In the, in the context of call work and how we build it, I think it's a bit of a combination. Like what the, the execution that gets cheap is not necessarily rebuilding all the primitives. I think a priori, there's also not a lot of value in it. So for instance, my team did not think about rebuilding cloud code. We like very much started with the, with the core thesis of this should be cloud code. And then we'll like build things on top of it. The part of the execution that gets a little cheaper, it's like, how do you take all of these Lego pieces And put them together in a way that makes sense for users. It is like actually valuable. You have so many different approaches now in terms of what kind of, what kind of things do you actually elevate to a primitive? Do you strongly believe that all your products should be built by just combining primitives that the public also ha…
AI assessment note: “the execution that gets cheap is not necessarily rebuilding all the primitives”