The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Cat Wu no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q his phone. Just like, I don't even know what the number is anymore. I think people don't give you enough credit for the success that Claude Code has had and co-work and all the things you all are building. Help us understand your role on the team, how you work with Boris, how you split responsibilities, just like what does the PMRO look like on, on the Claude Code team?

A I feel very lucky to work with Boris. He's been an amazing thought partner. He's our tech lead. He's very much the product visionary. And he is great at setting like, this is what the product needs to be. In like three months, six months from now. This is like what the AGI-pilled version of the product is. And a lot of my role is figuring out, okay, what is the path from where we are today to like that vision three to six months from now. And I, I spend more of my time on the cross-functional. So making sure that our marketing team, sales team, finance, capacity, et cetera, are like bought in on the plan and that we're all rowing the same direction and that Once the feature is ready, that there aren't any blockers to shipping it. I think in many ways it works well because we kind of like mind meld, but it is actually like remarkably blurry of a line. Like, I think we're like, 80% mind meld, and then there's like this 20% of things that like, maybe I care a lot more about than Boris, so like I'll drive those, and then like 20% where he cares a lot more than me, and he just like drives those.

AI assessment note: “a lot of my role is figuring out, okay, what is the path”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q But that was, that was wild. Okay. Uh, favorite product you recently discovered that you really love?

A The product that is like most changed my life outside of Claude products is probably Waymo. Like, I'm a diehard Waymo user, um, use it twice a day, get to and from work. So the two things that I really like about it are, one, I don't feel bad if a Waymo is waiting for me, and so I feel like, I feel less pressure to be right at the curbside the moment it arrives. And the second thing is, I feel like it lets me be a bit more productive. Um, when, when I'm in the car with another human, I, I typically try not to like do any work calls. I, I feel a little rude if I'm like on my laptop the whole time, but one thing I really appreciate about the Waymo is I can call into a work call. I'm not worried about someone overhearing me. I'm not worried about, hey, is this like rude? Am I talking too loud? Do I need to tell, ask someone to like change the music? And so this has been like, I feel like this has given me back like 30 minutes every day.

AI assessment note: “The product that is like most changed my life outside of Claude products is probably Waymo.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q every day there was like a major feature or product. So one question people had online is, uh, you guys just launched this, uh, not launched, but built this incredible model mythos that is still in preview. Cause it's so powerful. People are a little afraid of what it can do. Have you guys been using this? Is this part of the reason you've been able to move so fast?

A We've been moving pretty fast for Several quarters now, so I think it, it's not fully Mythos. Um, Mythos is an incredibly powerful model. We do use the models internally, and I think this has increased our rate of shipping a little bit, but I don't think it explains the bulk, bulk of the increase. I think a lot of it is the process and the expectation on the team. So we're very low on process. We want to remove every single barrier to shipping things. We want to make sure every single person on the team Feels empowered to take their idea from just an idea to like out in the world in less than a week, sometimes even in a day. Cool.

AI assessment note: “We do use the models internally, and I think this has increased our rate”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q essentially Picking the things to work on, knowing where the market's going and figuring out where, what to prioritize essentially. And then it's knowing if the thing you've built is good and right and getting it out there in some early version at least. Does that sound right? Is there anything else of just like where human brains will continue to be useful for at least the next few months?

A I think humans still provide a level of common sense that the models don't. And there's like a thousand moving pieces to any product launch. Some of them are very small, but there's always a lot that could potentially go wrong. I think the model doesn't always have a great sense of who all the stakeholders are, how they relate to each other, what their preferences are, what are the right venues to communicate with them, to keep them on board. I think a lot of this, like more tacit common sense, like EQ kind of knowledge is still very valuable. Of course we want the models to get better at this, and I think they will be, but right now I think there's still gaps.

AI assessment note: “I think humans still provide a level of common sense that the models don't.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Along those lines, something I'm always curious about is what's kind of in your, in your stack of tools as a PM and anthropic, obviously cloud code and co-work and all the anthropic tools. What else are you using? What other slack you mentioned? Is there anything else?

A So my stack is pretty heavily cloud code, co-work and slack. Anthropic largely runs on Slack. Um, I feel like it's like the core OS of our company and day to day, like a lot of, I would say maybe 30% of my time is pushing the boundaries of what cowork and quad code can do so that I have a very strong sense of what we're not good at and I spent a lot of time talking with the model to understand why it makes mistakes that it does. We actually have a lot of internal tools that we make. Like, I think one of the things that quad code has really unlocked for our entire company is it really lowers the barrier to making any custom app that you want. And so we we've seen this like surge in personalized work software that people are building for like custom use cases instead of. Um, using tools that don't perfectly fit the use case.

AI assessment note: “So my stack is pretty heavily cloud code, co-work and slack.”

Answered raw tape D 3 · C 5 · P 4 · Cm 4 4.00

Q We've covered evals a bunch. There's this trend of just like, that is the future of product management is writing evals because it, and essentially it's what does success look like? Okay, cool. Let me actually concretely define it and then we'll know. How much of your time are you spending writing evals, would you say?

A I think the importance of evals varies a bit based on the feature that you're working on and, or like what the problem you're trying to solve is. So there are a lot of folks on our team who do spend a lot of time working on evals. We have a small pod of folks who collaborate very closely with research to more precisely understand our cloud code behaviors and what the Largest areas of improvement are and trying to measure those pretty concretely. I personally jump into evals when there's a feature that I think needs a bit more product definition. And often the output of this is, okay, here are like five evals that I made. Um, this is how you run them. These are the ones that succeed and these are the ones that don't. And this is like the prompt that I've used to increase the success rate. It varies a lot, though, based on the exact feature. Uh, not every feature needs it, but I think features such as memory benefit a lot from it.

AI assessment note: “I personally jump into evals when there's a feature that I think needs a bit more”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.