Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Where do skills fit into the picture of co-work?
A So skills are essentially just markdown files that explain to the model how to do things. And I'm always surprised at how well this works. If you treat model, the model Claude in this, in this case, like a coworker, you get very, very far. My recommendation to like everyone I always talk to is just treat Claude the way you would treat a coworker. So a skill is fundamentally just a text file. And in the text file, you explain how to do a certain thing. My default example is always say booking a flight. At Anthropic, we have a specific particular vendor that helps us with our travel booking, so you can't just go to Google Flights, you need to go to this, like, particular vendor portal, and then we have various travel policies, and the same way I would explain this to a co-worker, I can explain it to the model. I'll just make a file that is like, here's how you book flights. You go to this website, and on this website, please consider the following things, and then maybe you also sprinkle in, like, a few personal things, right? Like, in my case, avoid red-eye flights, but also, I do actually enjoy my weekend quite a bit, so, like, Try to book a flight. If I have to fly to New York from San Francisco, try to like take the four p.m. flight. That's my favorite flight. And you put all of those things in the text file and the model then is extremely capable of understanding the instruc…
AI assessment note: “skills are essentially just markdown files that explain to the model how to do things”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q it's been an absolutely epic time at Anthropic and very hard to start this conversation with anything but, uh, the announcement which came out, uh, just yesterday as we were recording this of, uh, Project Glasswing and then, uh, Claude Mythos preview, which you tweeted about and you said it's Pretty hard to overstate what a step function change this model has been inside Anthropic. Can you elaborate on that?
A Yeah, sure. Mythos is a unreleased frontier model. It's, it's a general purpose model that was trained not specifically for cybersecurity or specifically for coding or specifically for software, but, uh, we have discovered what we believe to be outsized capabilities specifically in the aspect of cybersecurity, and we believe that it has Far reaching implications for the safety of software and infrastructure. I think there's two things I'm alluding to, uh, in my tweet. We've obviously used the model internally for a while now. As a software engineer, I think many of us have gone through this exercise of the last couple of years of like our first initial contact with AI was like, you know, probably not that impressive. The first time I touched AI was like sometime in 2013. This was before we had large language models. I was at Microsoft at the time. We had something called Project Oxford where we had an Ngram model. You would give us a token. You would say something like world and the model would return world wide web. And that was sort of the, I want to say the frontier of what language models were capable of doing. And I think a lot of us in the public over the last couple of years had these moments of being like, oh, this model is more capable. I can do more things than I may be expected. Mythos preview is a model that for us as engineers internally, Feels like a dramatic step…
AI assessment note: “Mythos preview is a model that for us as engineers internally, Feels like a dramatic step up”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q famously, uh, was, was, uh, coded in 10 days or so. At least that's, that's, that's, that's, that's the lore of it. Actually, let's, let's spend a minute on this if, uh, if the industry lore is, is not entirely, um, uh, correct. I guess what, what happened? Tell us that story of the, of the 10 days and the core work beyond entirely, uh, built by a cloud code.
A Yeah, I can kind of see, I can kind of see why that caught on. In software, nothing is ever built from scratch, right? And I think the, the exact quote that I gave that people used was that my team sprinted on this for, I think, the last 10 days or so, which is accurate. That is, that is the case. My team got together 10 days before release, and I was like, alright, we should probably release something. What did we release? What does it look like? What is it named? What can it do? However, however, as anyone who's ever built any software can attest to, it's not like you start from scratch with, like, ones and zeros, right? You, like, make use of a lot of libraries. You make use of, like, The research you've done in the past, in particular in Anthropic, the, the core problem that I tried to solve for, which is how do you make it easier to bring the power of cloud code to non-coding work, right? Like general knowledge work. A lot of very smart people have thought about that at length. And, um, it would be inaccurate to say that Anthropic has not thought about this problem. And it would also be inaccurate to say that I feel like slowly came into this cold without benefiting from all that work.
AI assessment note: “my team sprinted on this for, I think, the last 10 days or so, which is accurate”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q And where does memory live for co-work to remember you and remember your task? Is that the model? Is that in the harness?
A It's in the harness, actually, and it's, like, often surprising to people when I talk to them how we, how we've implemented memory, because I think it maybe points at the simplicity underneath all of those models. Memory is just text files. It's really just the, the model being instructed, hey, if you feel like anything was pertinent that you might want to remember in the future, just write it down. And then we help the model a little bit with, like, organizing its memory so you can, you can set up projects that have isolated memory versus, like, your overall memory. But the, the underlying technology that sort of is bolted on on top of the model is sometimes surprising to people that it's not, you know, like a, some complex, fancy database technology.
AI assessment note: “It's in the harness, actually”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Would you say that, uh, UX is as important to the success of an AI agent as the technology itself? Like how you take users on a journey, uh, so that they are empowered? And, uh, if so, what, what are some other lessons learned building AI agents from a UX standpoint?
A It's a really good question because I actually think, I actually think that's true. I do think the UX matters quite a bit. Right? Like, even if you go back to our, one of our most popular products, Cloud Code, um, the very genesis of it was, what if Cloud, but instead of, like, in the cloud, it's running on your computer in your terminal. That, that is almost entirely UX. It's the same model. It's the same core capabilities. Um, it's really all around what is the user experience and, like, how do you interact with the model, right? But it's fundamentally the same model, and that's really where a lot of the, a lot of the benefits came from. And I think similarly today, um, the AI products I see resonate with people the most are rarely the ones that deliver the most raw potential, the most raw power. Um, and I, I would actually go one step further and say this is probably true, not just with AI, but maybe with software overall, right? Like, um, I'm going to blindly assume that plenty of startups out there offer email with more features than Gmail. There's plenty of companies that try to, like, sort of jump ahead by offering a larger amount of features or more buttons or, like, more capabilities. Um, I often think a lot about the, um, The silly times of mobile phones right before the smartphone was invented, right? All the things that people bolted onto phones. We had like phones …
AI assessment note: “I actually think that's true. I do think the UX matters quite a bit.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q You mentioned local, and I know you have a strong thesis about local AI. Do you want to get into this? Why does Cowork need to live on your laptop as opposed to the cloud?
A The two biggest things that Cowork gives you today are access to your local computer. And also access to your local files. Why does that not work in the cloud, right? Like, I think a good example for me is always maybe using your Chrome. Cloud, if you give it access, and again, only if you give it access, can use your Chrome, um, which is a pretty powerful tool for Cloud to, like, interact with the rest of the world, right? Like, be it responding to emails, summarizing your emails, or, like, maybe interacting with a tool that only you have at your company. Um, I, I often play this through for people who might think, why can't we just do that in the cloud, right? Like, The first case is your sessions, and it's quite useful for cloud to have access to the websites that you care about with your accounts, right? Like Gmail is not all that useful to my agent. Gmail with my login information is quite useful. The second case is that, um, and this is usually a debate I get more in with other software engineers. As to a software engineer, this is an implementation detail, right? We could find some kind of way to take your local Chrome, zip it all up, put it in the cloud, Ask you for your passwords, do all kinds of things. Um, there, there's two oppositions I have to that. The first one is probably sort of on the basis of safety and security. I, I don't think we should teach people that …
AI assessment note: “The two biggest things that Cowork gives you today are access to your local computer”
Answered raw tape
D 5 · C 4 · P 4 · Cm 3 4.15
Q Any other lessons come to mind? So if I'm, uh, listening to this and I'm building an AI agent of, of some sort, um, about the, the process, like building that harness and specializing it, and it could be guardrails, it could be like industry specialization, any, any, uh, thing that people, uh, could learn?
A I would probably recommend first not to actually build your, not to build too much of your own infrastructure and use a product that we've launched today called Cloud Managed Agents that make this particular case very useful. You know what? I'm going to give you, I'm going to give you both. I'm both going to give you the advice and the reason for building custom agents and a lot of harnesses and trying to make a company on top of that. And then I'm also going to give you the case for. The case against is That as the models get more and more capable, what I'm noticing inside my products and inside my work is that we're sort of like pulling back the edge cases we account for.
AI assessment note: “not to build too much of your own infrastructure and use a product”
Partly raw tape
D 3 · C 4 · P 4 · Cm 4 3.70
Q So, so just to play it back, so you're saying that you'll, you'll actually create 10 products or 10 versions of the product actually running, and then you'll have people at an anthropic test and, and, and sort of guide which one you should eventually pick?
A We probably have easily a hundred different prototypes of like various applications inside the company right now. None of them have necessarily yet hit the confidence of like, this is good enough to show to a user, but the amount of prototypes you can build internally very, very quickly completely War of anything I've done in the past because of the cost of execution, right? Like in the past, the thing that would always hold you back is, you know, for me as the engineering leader, um, if you had a good idea, you would come to me and I would tell you, oh, we can work on this next month. It's going to take us three weeks. Until then, like, go and talk to the customer, validate your ideas, and now you can come to me and say, oh, I have an idea, and I'm like, cool, give me 10 minutes, I'll send you something. And that is, that is just, it's like going from the painting to, to the photograph, you know.
AI assessment note: “We probably have easily a hundred different prototypes of like various applications inside the company”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q Just looking back on the journey, which at the end of the day is, uh, five months, uh, what four months old journey. It's insane. Uh, the impact that, that you've had in such a short period of time. What was the hardest part?
A I'm thinking about your question to the lens of like, what is the hardest to replicate, right? Like if you, if you told me, okay, now do it again, do it with another product. Like what would be the most difficult to replicate? I think there's probably something about a point in time. And I mentioned that co-works sort of came on the heels of us, like, keeping our ear to the ground and saying, oh, there's something here, there's latent demand. Latent demand is a gift. I don't think it's, I don't think you can, you can go and try to look for it, you can try to find it, but it's very hard to create out of nothing. Recreating that would probably be the hardest thing. Now, I do think software has always had ample latent demand. Like, if you, if you looked for it, you could always find it quite a bit. That's, that's certainly one thing that is That I think is hard to, like, replicate. In terms of actually building co-work, I would not say that anything was particularly hard. I think the things that are hard about building good products remain hard, right? Like you, there's sort of, like, the perils of success. Like, what do you do if, right, like, you open up a cafe, and instead of 10 people, twenty million people show up, what do you do? Um, that's, that's, that, that was probably at times, sometimes hard for us, and like, Remains a challenge, the overwhelming demand for anthropics …
AI assessment note: “In terms of actually building co-work, I would not say that anything was particularly hard.”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q So it does feel like a major discontinuity moment, uh, potentially, right? I mean, And then hearing the words terrifying is, uh, not necessarily reassuring.
A I mean, I think, I think Anthropic has long held the position that AI can be extremely powerful, very beneficial, but that, that there are risks that we ought to take seriously. Right. And I think this is one of the areas where we, for the first time see this, I want to say like applied in practice, which is like, Quite interesting to watch, right? Like, you know, have this, you now have this model that is very capable of breaking into software systems. What does that mean? What do we do with it? How do we handle this responsibly? And it's, it's Not like to add anthropics horn too much, but for me as an individual, it's like a bit of a point of pride. I'm like very proud to see the company handle this very responsibly. And I think a lot of my colleagues share similar appreciation. Um, you've alluded to the fact a little bit that I, we've had this model before, right? It's not like we immediately found a model that was very powerful. And I think there's, there's an alternative universe in which maybe a company with a less steady hand would have Race to get it onto the market as quickly as possible, put a very expensive price tag on it, and just like reap the benefits.
AI assessment note: “we now have this model that is very capable of breaking into software systems”
Partly raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q as in you're not gonna take files that you shouldn't have access to. There's also trust in Okay, coworker, I'm, I'm trusting you to run certain tasks which are going to be increasingly important to me and my work life in a way that's going to make me great and not embarrass me. What have you learned as a head of product about building that level of trust with people?
A Yeah. Yeah. That's a good point. That's a good point. I think there's something interesting if you build AI products in 20, 26, which is that most of the buttons you add and most of the product services you build are probably more for the human than they are for the model. And this is an interesting shift in how we build technology, right? Like in the past, we've usually built buttons for the benefit of the computer and the human was just there to provide information so the computer could do things. Now we're actually doing it the other way around. Um, I'll give you one, one quick example. We have recently launched a feature called Dispatch, which allows you from your phone to talk to Claude on your computer. As a very conscious choice, we decided not to add too many buttons. So, one of the pieces of feedback I got on social media the most, easily got 50 messages every single day, from people asking me, hey, it would be cool if Dispatch could access my local files. That would be really nice. Like, can you, can you find a way so that it can attach a folder? I mention this because Claude can access all your files and folders. The way this currently works is that you ask Claude, hey, can you also see my downloads folder? Claude will say, yeah, I can see it. Do you give me permission to interact with your downloads folder? And once granted, it would go. So we're debating, do we add…
AI assessment note: “less about Claude proving itself to the human and more so slowly educating”
Answered raw tape
D 4 · C 3 · P 3 · Cm 3 3.30
Q And to double click on that, um, what, what does that mean? You mentioned, uh, like understanding, what was the term you used a minute ago? Understanding your, uh, industry and your users. And now you're mentioning the, the human aspect. So is that a question of UX to the earlier discussion? How does that manifest?
A Like, I think, I think successful software developers, 20 years ago, were very good at understanding computers, right? Like in order to build successful software, you need to be very good at computer. You were a computer expert. And I think the people who will build successful software going forward will increasingly understand humans and users very well. And I think this has been a gradient. This has already happened somewhat, right? Like building software 10 years ago was already much easier than 30 years ago. And I think AI is another step function change. When it comes to the market, I am not an economist. I'm a software engineer. I've never fully understood what the markets do, um, and I would recommend to other software engineers not to, like, base too much of what they do on what the markets do. That is my personal recommendation. But I, I really do think, like, to answer your specific question, what is left to do, I think there's, there's mountains upon mountains of, like, things we can automate for people, work we can make easier for people, problems we can solve. I think as long as humans have questions and problems, like, the software will be a reasonable answer.
AI assessment note: “people who will build successful software going forward will increasingly understand humans and users very well”
Redirected raw tape
D 2 · C 4 · P 3 · Cm 2 2.85
Q you could talk about, um, speaking for ourselves, One question we're curious about is, uh, whether you're going to, uh, enable regulated industries to, uh, have a better, easier access to co-work, uh, because as a venture capital firm, um, uh, we don't have access to co-works. I have access to co-work in my personal life, uh, but, uh, not, not at work. Is there a, uh, roadmap for that?
A What I'm gonna say is that you're not the only one who's asking for co-work for the particular regulated industry. Um, it's something we hear quite a bit, and whenever users ask for something, um, we listen very carefully. Uh, right, that's fundamentally our job. Um, I can't, I can't particularly, like, comment on anything that we're currently working on, but I can sort of mention, like, the general concept of things that I'm so excited about, and the general concept of things that I'm very, very excited about still in Is really the idea of helping people organize their work in a way that, like, makes most use of the capabilities in AI. Um, and if people are sort of, like, listening to them, like, what does that mean? What is he talking about? Once upon a time, I, I spent five years working at a company called Slack, and Slack at the time was, we've certainly felt like we were having some companies revolutionize the way they work, but we were certainly not the first chat app. And we're also not the first company to tell you that, like, your company will be more effective if you don't have all of these information silos. But very similarly there, a huge part of the thing that we sold people was not just a chat application. It was like this different way of working, like a more transparent, more open way of working. And for AI, there's a similar change in this tool is most effect…
AI assessment note: “I can't particularly, like, comment on anything that we're currently working on”
Not addressed raw tape
D 1 · C 3 · P 3 · Cm 3 2.40
Q broad group of professionals, you know, smart people that are good at their jobs, trying to be good at their jobs, and some of them will be doing revenue ops, some of them will be doing marketing, some of them will be lawyers, some of them will be accountants. What does taste mean in a context where you have such a broad audience, and how do you test for it?
A Yeah, I think, I think a lot about, I've been mentioning it so much already in our conversation, I feel like almost silly about it, but I think a lot about the phone and how all of us start with the same phone, but like no two phones are the same. The exact apps you have installed probably makes your phone like unique among all the phones on the planet. Um, it's almost like a fingerprint, same with my phone. We all start with a device that probably looks very similar to the other devices, but then the way it integrates itself into our lives is not always good, not always bad. Um, but certainly very unique and certainly very personalized. And I think for co-work, our approach is similar that we want something that generalizes extremely well, that we can apply to your life across a broad range of applications. Um, and maybe just speaking from my personal life, currently in the process of moving and moving my family into, into a different house. Um, and as many people, certainly, certainly the ones who are also listening in America know that involves about 500 pages with a lot of words that I barely understand. Cowork here is extremely helpful, but it's also extremely helpful in, like, healthcare scenarios. I just had a daughter this year, and working, working through all of that paperwork has been super helpful for me, too. But these are two widely different things, right? Like, …
AI assessment note: “I think a lot about the phone and how all of us start with the same phone”