The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Sarah Guo no published score: only 4 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 4 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
4exchanges match
4on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q So is the question around, um, How do you leverage alternative sources of data?

A Yeah, the, the question is, um, I, I think there is, like, uh, I, I don't want to, like, over analogize to robotics, right? But within robotics, you have learning from world models, you have learning from sim, you have learning from embodied data that, uh, of different types, right? Um, imitation, then you have RL. I, I think it's, like, much less clear that You can use RL for a lot of robotics today, especially some of the harder, like, manipulation problems, and I'm curious, just given, you know, your team has this enormous strength in RL's, like, a starting premise, how you look at other types of data to create the, you know, coding agent experiences you want.

AI assessment note: “Yeah, the, the question is... how you look at other types of data”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q now. Overall, though, the key The key thing that we have to achieve is to show value, show ROI, right? Like in the case of, of Gucci and other retailers, are we reducing average handle time? Are we driving sales conversion uplift over the baseline? And so long as we're doing that, I think customers are willing to pay. It can't just be an added cost without a clear benefit.

A Yeah, maybe two reasons I'm still pretty optimistic about this question is you, you certainly have a more complicated unit economics equation than like. It's a web app and like we're using database services that are like very efficient today. Uh, but in so many applications, you have net new capabilities or like massive productivity gains, right? So wherever you see orders of magnitude improvement in terms of value you can give to the customer. And then on the cog side, you know, being able to do the same task with these AI models decreases over time monotonically, right? As. Uh, we improve at every level of the stack for AI. And so I definitely think it's a complicated question, as you said, in terms of presenting that answer, getting the answer, like right to a really fair trade for customers. But I'm, I'm, I think there's a lot of opportunity.

AI assessment note: “in so many applications, you have net new capabilities or like massive productivity gains”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q language models? Um, effectively, the blueprint for building super intelligent systems was developed. It happened with, um, the Atari games, AlphaGo, um, you know, then Dota V and AlphaStar were near super intelligent systems, and if OpenAI and DeepMind had sunk more compute into them, they would have definitely become super intelligent. It's just that at that point, it didn't really make sense. Economically, like, why would you do that?

A Then this is a definitional issue, because I, I was gonna ask, like, help me understand your view of, like, I don't, like, one of the big criticisms of RL overall has been lack of generalization. Um, that's been just kind of a general question for this direction. I do have friends at every large research lab that's somewhat, you know, some, I mean, tell me if you hear, uh, something of a different tenor or just believe differently. They believe we're going to have systems that are much more capable than humans and many types of knowledge work, but they believe less in generalization. And so in a resigned way, they're also, as you're saying, like, I guess we're just going to bring all of it under distribution one way or another. But that means like, you know, it's a little bit different than my, my view of like, it's, um, at some point you're, you're just, you know, you have enough capability that the rest you get for free. Right. The rest is sort of useful capability you get for free.

AI assessment note: “Then this is a definitional issue, because I, I was gonna ask”

Not addressed raw tape D 1 · C 3 · P 3 · Cm 3 2.40

Q if you're doing a task for a really long time, you're going to run out of context, so what's an efficient way of, like, dealing with that? Um, allowing the model to continue to do its thing. And then, yeah, the, just the task of making, making the data and making the tools. I mean, I've, I've said this already a few times, but that, that's a lot of work.

A I was just looking at my history of queries. My user request is like, I want to see what things I asked of deep research versus other models in particular in my memory. But it has ranged from like, obviously, you know, if I'm trying to get up to speed on a market for a company I'm looking at or on a technical topic or, Travel planning. It's a big one. Also, I, I have looked for things that are taste-related, so I'll be like, okay, I, um, like, you know, this set of books for these reasons. I want you to, you know, actually just giving me a long-form summary of a bunch of other things you think I should read and explain why. I realize I don't have a super clear mental model of, like, when deep research should be better than O-three. What instinct can you give me here?

AI assessment note: “I was just looking at my history of queries.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.