Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Um, what's your, uh, Vibevel? So as soon as you get a new model, what's like the first two, three things that you do with it?
A It's such a good question. I have three, uh, and, and the, the team got really tired of me by the end of this, uh, Sonnet four foot five process. Cause like every, literally every snapshot, you know, I would drop everything I was doing and run these. So one is, um, the virtual boy, I had a, you know, I actually never owned a virtual boy, but the local video game store in Brazil when I was growing up had one. And then if you remember, this is like a doomed Nintendo console that had these like a stereoscopic wireframe Red and black, three D graphics is a very early product. It totally failed. Um, but I always, I like to have, uh, in cloud AI, like create me a like virtual boy style, three D shooter game. And it's really funny to see the, the, all the checkpoints from like early Sonnet 4.5, where I was like, I don't know, guys, this is like, it's pretty early. I know there's a lot of RL left, but it's not looking that good to, you know, about a week ago, I was like, okay, great. This is like officially good. It's like better than Opus at this. It's like, It generated this, like, great split-screen stereoscopic thing, three-dimensional, like, thing. So anyway, you could really see it evolve. So that's, that's one I always, uh, shoot for. Um, within Cloud Code, there's a particular, uh, sort of change to our code base that I, I, I, like, once did in Cloud Code. I was like, oh, this …
AI assessment note: “I have three, uh, and, and the, the team got really tired of me”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q the final one with Sonnet 4.5, it has the chat history right on the side, versus on the actual Cloud AI, you have to click through to go to the history. Like, do you use these models to, like, think about all the different permutations of, like, how to build these UIs and products, or do you feel like the models are still, you know, pretty median results so far?
A One of the better projects that we did here was get a lot of our kind of products into artifacts, or at least into a harness that we could iterate on them really quickly. So one really fun use, um, Nate Parrott is one of the designers on our team. He, he got Claude to be able to prototype Claude code UIs, even though it's a terminal UI, Claude can sort of imagine what that's like. And it's very valuable to say like, all right, well, what it would look like if we revamp settings in this way and not have to go and necessarily code the end. So I think they can be very useful in sort of exploring that space of What, how could this product evolve? What does it mean to do this, uh, differently? And they can also actually just implement the changes as well, but even sort of from a prototype phase, it's valuable to have it sort of iterate quickly over ideas for non-engineers on the team.
AI assessment note: “I think they can be very useful in sort of exploring that space”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q have a code diff, you can give it to the model and the model reads the code and understands what's being changed. It's kind of hard for the model to understand the aesthetic change in a way. Like, how do you think about explaining to the model these things? Like, have you come up with like a good model of like taste semantics to put it in the latent space?
A I actually think Figma did a really good job with Make in that they, um, you can tell that there's much more of a bridge between the sort of underlying design model and what the LLMs are doing in kind of coordination between Sonnet, and that's, I think that's part of it, which is, um, Giving the model a sense of sort of the design building blocks is important rather than just the finished product. And I think you can reason about that as well. But the second part, and I think we need to make some strides here too, going forward is the models don't see as well as they could. Um, they see, okay, you know, you ask them analyze a complex photo and they're able to do it, but I want them to be as persnickety as a like really good visual designer. Like, no, that looks, the baseline looks a little off, you know, or this needs to be good. And I think That's going to come from sort of additional vision capabilities that, you know, we'll work on, but I think that's going to be a really key piece of, of closing that loop so that, you know, model, and I've seen people do this with cloud code and like take an MCP with Playwright, for example, in a browser and basically do the loop of, right, you generated the UI, now look at what you did, and is it right? And can you iterate on that? And I think the, is it right, still needs some, some visual help before it's, it's fully done.
AI assessment note: “Giving the model a sense of sort of the design building blocks is important”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q telecom at the three agentic tool use top bench, uh, categories that you mentioned. How much of alignment is there between that and what you think is important or like, for example, financial services, you know, there's like the financial analysis agent at the bottom that, you know, that's another category that popped up. Like, is that usually directionally correct? What you benchmark against or like spaces you're going into.
A It's interesting and like I'm a click removed from this, but what I've observed like working with our research team is that the, um, the benchmarks are, and evals in general are a helpful parameter of how the models are doing relative to the industry, but it's important to really ground sort of in your really hard customer problems almost more so than that. I think that's the case in, in, um, uh, even within coding, like there's been deltas, you know, Maybe two weeks before the final snapshot of the final snapshot where yes, it improved on, on sweet bench, but even more so it went from, ah, it's mostly reliable to, yeah, it's great. I'm using it every day in cloud code. And I can't tell you that there was a crossing line between like 76 and 76.5 and sweet bench. And actually it was interesting is even when it was already outperforming Opus, for example, on sweet bench, people still didn't feel it was better, but then it continued to train and it was like now better than Opus and people don't want to switch back. So There is this, Jared calls it an isaqua, right? Like there's, there's something that can happen in the models where they become more useful for, for particular, uh, pieces. So in terms of verticals that we look at, you know, we have a couple of customers that really love pushing the model in different ways, right? Either they have like a super complex, you know, agen…
AI assessment note: “benchmarks are, and evals in general are a helpful parameter... but it's important to really ground”