The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mukund Sridhar no published score: only 4 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 4 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
4exchanges match
4on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What's the initial ranking of the websites? So when you first started it, there were 36. How do you decide where to start? Since it sounds like, you know, the initial websites kind of carry a lot of weight too, because then they inform the following.

A Yes. So what happens in the initial turns, again, this is not like a, it's not something we enforce. It's mostly the model making these choices. But typically we see the model exploring all the different aspects in the, in the research plan that was presented. So we kind of get like a breadth first idea of what are the different topics to explore. And in terms of which ones to double click on, I think it really comes down to every time you search, the model gets some idea of what the page is. And then depending on what pieces of it, sometimes there's inconsistency, sometimes there's just like partial information. Those are the ones that double clicks on. And, uh, yeah, you can continually like iteratively search and, and, and browse until it feels like it's done.

AI assessment note: “it's not something we enforce. It's mostly the model making these choices.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I think we all, we all get that. I was more curious, like from a product perspective, when you decide to do rag versus shit like this, you didn't need to, You know, do you get better performance just putting everything in context or?

A The tricky thing for RAG, it really works well because a lot of these things are doing like cosine distance, like a dot product kind of a thing, and that kind of gets challenging when your query side has multiple different attributes. Uh, the dot product doesn't really work as well. I would say, at least for me, that's, that's my guiding principle on Uh, when to avoid drag, that's one. The second one is, I think, every generation of these models are, uh, like, the initial generations, even though they offered, like, long context, their performance as the context kept growing was, you would see some kind of a decline, but I think, uh, as the newer generation models came out, uh, they were really good even if you kept filling in the context in being able to piece out, uh, like, these really fine-term information. So I think these two, at least for me, are like guiding principles on Movento.

AI assessment note: “at least for me, that's, that's my guiding principle on Uh, when to avoid drag”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And so from a user perspective, is it better to just start a new research instead of like extending the context?

A Yeah. I think that's a good question. I think if it's a related topic, I think there's benefit to continue with this thread, uh, because you could The model, since it has this in memory, could figure out, oh, I've found this niche thing about, I don't know, milk regulation in this case. In the US, let me check if your, you know, follow-up country or place also has something like that. So these kind of things you might have not caught if you started a new thread. So I think it really depends on, on the use case. If there's a natural progression and you feel like this is like part of one cohesive kind of a project, you should just continue using it. My follow-up tone is going to be like, oh, I'm just going to look for summer camps or something. Then, yeah, I don't think it should make a difference, but we haven't really, uh, you know, pushed that to see, uh, and, and, and tested that, that aspect of it for us. Most of our tests are like more natural transitions.

AI assessment note: “if it's a related topic, I think there's benefit to continue with this thread”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Obviously they're going to like do a normalization number thing, but users are always going to want one, right?

A We've discussed this a bit. Like if I wear my pure user hat, I don't want to say anything. Like I come with a query, you figure it out. Like sometimes I feel like there will be based on the query. Like for example, right? If I'm, uh, asking about, Hey, how does rising rates from the fed How's old income for a middle class? And, and how, how has it traditionally happened? These kind of things you want to be very accurate, uh, and you want to be very precise on historical trends of this, and so on and so on. Whereas there is a little bit more leeway when you're saying, hey, I'm trying to find businesses near me to go celebrate my birthday or something like that. So, in an ideal world, we kind of figured that trade-off based on, uh, the conversation history and the topic. I don't think we're there yet, uh, as a research community, and it's an interesting challenge by itself.

AI assessment note: “if I wear my pure user hat, I don't want to say anything.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.