Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q get my guests to do a better job of is just brag. Could you brag a bit? Yeah. Just like some really impressive project that you accomplished, just, just the opens people's minds. Like, let's get, let's get specific without maybe naming the exact client, unless you can. Um, and then also like, what's the highest hourly rate that one of the engineers has made since you're technically on path?
A Yeah. So I'll answer the last one, uh, or the second one first. We will probably have more than one engineer make million dollars cash next year based on this model. And that is just with story point compensation. It's very likely that we will have more than a handful of folks make more than a million dollars next year. Um, The answer to the first question, like, for example, one project we built, so we work with this company that's a, they, they build, they work, they partner with retailers to basically make cameras in the business more valuable, and the way that they do that is they deploy what was historically like a gen four raspberry pi to the stores, and they would, they would run like one model on that device. We basically took some off-the-shelf models and trained some models ourselves, and then quantized them down so they could actually run on that Um, on that for, but also on jets and nanos, and we got them to all run in parallel. So now basically what these models allow you to do is as a store, you can get a heat map. You can see where the lines and the cues are forming in your store. You can even get pictures of shelves and understand what needs to be stocked. And you can do things like theft detection because we have body analysis and we can understand things like things are crossing arms, right? And This took our team two weeks to put together early prototypes, an…
AI assessment note: “We will probably have more than one engineer make million dollars cash next year”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Yeah. That's amazing. Okay. So, uh, like quick question on just the, uh, the stack that you guys have landed on, like, is there a house stack, what are you guys finding in terms of like the various coding agents and all that?
A Yeah. Um, we do work in a number of different stacks, a number of different languages and stuff, but We feel pretty strongly in like high structure allows for agents to work autonomously for longer. And so our default stack is TypeScript front end TypeScript back end with a shared file where, or a shared folder where all of our shared types and schemas and things like that live and typically react front end, or even something as simple as like express on the back end. Like we don't really care about the frameworks. It's more just like TypeScript allows us to have that flexibility to, um, like the flexibility of JavaScript, but the, but the constraints of TypeScript. And then those error messages allow the cloud code or cursor agents or whatever to iterate on themselves and, and run things, see the errors and, and continue in terms of the actual like AI engineering stack and what coding agents and things like that we're using. I always tell clients this, like our team doesn't have a favorite coding agent of the year or of the month or even of the week. Like if I go over there to our team right now and I ask them what model is performing the best for coding? Right now. They'll say today at four 42, we're noticing that Claude code is actually performing better because of X, Y, Z reason. But yesterday codex was outperforming Claude code on object on, on activities like X, Y, and Z.…
AI assessment note: “our default stack is TypeScript front end TypeScript back end with a shared file”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I think like, I think that the classic thing is, well, what is a unit of output of a software engineer? Is it a PR? Is it a story point? It's extremely unclear and it's very basically unsolved. Like, I mean, don't tell me you've solved it. You know, like what's, Maybe you have, I don't know, but I'm, I'm default skeptical on the, well, what gets measured gets gained.
A Yeah. We do use story points, but you're right that it's, it's easy to game it, right? Like, like if we were to hire somebody who just, like, if you think about a technical system, right? A smart hacker will find ways to exploit it. And the easy way to exploit the story point system is to deflate the concept of the story point and decide that, okay, Any line of code, like lines of code are going to be directly proportional and equal to story points. Well, then of course you've hacked the system, right? But your clients will churn and you'll probably get let go of, and it just won't work long-term. And so what we found is that hiring two, what we found is that this problem gets solved in the hiring process and it gets solved by hiring people who fall into two buckets. One is people who are selfish, but they're long-term selfish. Everybody's selfish, but we need to look for people who are long-term selfish, people who understand that these incentives are longer than just today's story points. They're forever, right? And we need to think about how do we maintain the client relationship. And that means that we're going to give them very robust story points so that we can maintain the relationship and continue to make money. But the other is that we hire people who just like writing code and like working with really smart people and, and they're not sharp elbowed and they just want …
AI assessment note: “We do use story points, but you're right that it's, it's easy to game it”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q I mean, well, so yeah, but you're gonna, it's very anecdotal, right? Like, don't you need more comprehensive evals? Because otherwise it's like you are just behaving or believing things based on the luck of the draw.
A I think at this state, did a samurai have a measurably better sword than the person to their left or right? No, right? At a certain point, I think a, a, a, a warrior's weapon becomes something of a feel. And I think that at this point, a lot of these, like, The coding agents are so good. Like, yes, you can have evals that, that provably show that one is better than the other. But for a lot of these things, it really is feel it's like, Hey, this agent actually, like it just, I can work better with it on a warm blooded level, or it writes code more like I like to, or whatever. And at least that's what we've noticed.
AI assessment note: “yes, you can have evals... But for a lot of these things, it really is feel”