The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Ben Hillock no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q on the, on the post training side. So that kind of changes how you interact. Um, What, when did you first make the distinction? You know, is there kind of like some day that you were using the model that you were like, okay, like, I just need to completely change the mental model to think about it? Or was it just kind of like, you know, a gradual progression?

A There actually was a day and it's, uh, it's kind of funny. Uh, I think that, um, I, I was just trying to do something over and over again, and it's obviously super long response times. Uh, separately, uh, there was, I was having a bug with the ChatTript app on my phone where, like, it wasn't, uh, like, every time I, like, went to a different app, it would, like, stop responding the, the O-one response, which was, like, kind of silly. So I was hitting all these bugs, and so the only way to get around this bug was to just, like, not have to iterate with it, right? Like, essentially, like, uh, I, I was, like, I can't keep doing these back and forth because everyone takes, like, You know, five, or like, you know, five minutes to respond, but then every time I would accidentally, like, do something, it was, you know, adding more time. So then I actually took the time, like, not thinking about it from like a, oh, you know, like model perspective, whatever, but just because I was tired of waiting so long between, uh, iterations. And, and then the, the answer was really, really good. Um, so that's when, that's when it clicked for me.

AI assessment note: “There actually was a day and it's, uh, it's kind of funny.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q covered this, but, uh, coding slides. Dan, you, you had a video, which I think you, you showed off, uh, some tips on how you, how you were doing stuff. And then like, uh, Ben, I felt like a lot of your examples were driven from copy, placing code into chat to me. Is that how you code? You don't use cursor, like, you know, like, like a normal person.

A No, I go back and forth. So honestly, I do use cursor. Um, I actually just recently switched to cursor, believe it or not. I, I've actually been sharing some of my thoughts, like, uh, switched probably, uh, like a month ago or so. So like after, after the sort of initial high for context, like I had, um, when we first started our company, we had like a super early co-gen plugin in VS code called sidekick. So we had like shut it down publicly, but we kept like adding stuff to it. And like, honestly, it's really good. So we, we kept using it for a long time. I think for specifically, uh, a lot of what I've been using it for is, like, a SQL queries and, like, generating, like, a lot of, like, backend and architecture stuff, so that, that's where I, I do use ChatsDBT. I have found, like, differences in, so for example, I actually generally get better results out of Claude. It's in, you know, just Claude.com, uh, versus, uh, Cursor a lot of times, uh, which is interesting, like, I love being able to apply code, et cetera, But as far as just like, a lot of times I find it getting stuck in sort of a loop that I don't see in quad.com, and so I haven't seen that for, I haven't tried it too much for O.N. specifically, uh, but, uh, that was just sort of an experience I had, uh, early using Cursor, so.

AI assessment note: “No, I go back and forth. So honestly, I do use cursor.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q Are there for either of you, any use cases that you were not able to do at all in previous models that you didn't even try at first in a one, and then you found out it was actually good. Any lessons there?

A I think this is, this is interesting. Like, uh, one, one thing you'll see a lot, unlike even reactions to the post is like people looking for specific examples where, you know, one is a lot better. I think Dan actually did a great job of like pulling out some of those examples. I think it's really hard because a lot of it is just like deeply specific to like, like, I think that it generally understands the project I'm working on better, for example. Um, and like the intricacies of how the code base is laid out. Uh, there's like, we use click house a ton. Um, I think specifically it seems to like confuse like quick, uh, click, click house, uh, SQL, um, syntax less, but these are very hard things to like articulate to people in an example.

AI assessment note: “specifically it seems to like confuse like quick, uh, click, click house, uh, SQL”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q Are there bad return formats that you tried and kind of moved off of? Or, you know, like the hiking example is one. Are there any other good or bad examples that come to mind for both of you?

A I mean, interestingly, like, I think, you know, the hiking is somewhat like a convoluted example, because again, I think it is sort of hard to have an example that actually, like, is sort of universally applicable to people, which is, like, what my goal for that, um, like, I think that if I, if I was actually looking for hiking things, well, I actually did try that prompt, and I, I think that something with, like, access to the internet actually will be better in those cases, which makes sense. Think that I talked about it a little bit there. I think that I've had a very hard time getting it to actually write stuff. I know that I've heard of people using it for writing where it's like processing diffs, more like providing critiques or feedback, but at least for myself, I haven't found a good way to get it to write in a specific tone or like match my tone, for example, or really match anyone's tone. But if there is a way to do that, I would, I would love to know.

AI assessment note: “I've had a very hard time getting it to actually write stuff”

Answered raw tape D 3 · C 3 · P 3 · Cm 2 2.85

Q you think about describing, like, how much of an important should how you want it be? I mean, going back to the tone, right? It's almost like people are so focused on, oh, I need it to be my tone, but then maybe if I can write a more compelling Copy, and it's not in your tone, then that's just as good, right? So how do you think about that?

A That's interesting. I think the tone one for specifically, I think, I mean, even taking a step back, I think the sort of tough part here is that, um, Even with LLMs in general, I think that, uh, you know, even since like, you know, GP four, GP four five, like there hasn't been an instruction manual. Um, I think that, you know, even opening eye doesn't really have an instruction manual. Like they sort of like this thing is, is born and they're trying to figure it out in the same ways that we're trying to figure it out. I think that what makes a one even trickier than other models is that, um, there, there is actually an asymmetric miss to how well open AI understands the model and how well we, for example, the fact that like reasoning tokens are hidden, right? So there's all this stuff happening that we actually just can't see, right? Even as developers, which is super interesting. Um, and then also there's this whole component of like how it was strained that we're also not privy to, which is like, I think very different. It seems like than how other sort of like what the industry standard is. So yeah, just, just prefacing that, I think that the tone thing is somewhat a side effect of the fact that it has all these reasoning tokens that is in a very specific like tone. It feels like very academic tone. So I think that's where that comes from. I do think that the describing what…

AI assessment note: “I do think that the describing what you want part is extremely, extremely important.”

Partly raw tape D 2 · C 3 · P 3 · Cm 3 2.70

Q should improve it, especially when you have kind of like this long tail of use case where maybe people are trying to jam new models into old use cases. Are there any heuristics that you think about also? Like I should not be using a one for this, or are you just trying to upgrade most of the tasks and workflows to, to be a one compatible, so to speak?

A Yeah, that's really interesting. I think one really fascinating thing for us where, you know, for folks that don't know, we're, we're essentially like a century for AI products where we help people catch issues in production. And, uh, I think that models are actually getting harder to use, which is like very interesting place for us to be where, you know, we're, we're finding that like, oh, one is the most capable model. I think that, uh, opening eye has made. And it's also, I think the hardest, uh, to use. And so, being able to understand, like, how users are using it wrong, being able to catch those issues, the ones that you're not seeing, actually, I think is becoming more and more important in the network. I think that, um, as far as, like, how applications should be thinking about using it, I think what I'm excited about is it being used in, I think that we, we could see real, like, background level intelligence. I think this has been a promise, like, uh, An example of what I mean by background level intelligence is like, you can imagine that you have something looking through your GitHub repo and it is like finding, you know, oh, you should actually be refactoring a certain part of your project a certain way. I think that that's something that like, you know, you could have the quality of answers continues to scale with compute time essentially, or like tokens outputted t…

AI assessment note: “as far as, like, how applications should be thinking about using it”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.