The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Amjad Masad no published score: only 1 usable exchange on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
1exchanges match
1on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Are there two, or are you, are they both running or only, you're only running? So you're not using any compaction from Claude?

A No, no, no, no. Just ours is better. And I think the thing about the primitives that the labs build, They tend to be very generic, and so I think every app will need to, uh, you could start with the generic firmatives, but for it to perform really well on your benchmark, on your users, every use case is going to be different, and so we care a lot about, like, you know, certain pieces of data around the architecture of the app. There are certain things that absolutely need to remember. Like, if it forgot that it has a database, like, that's catastrophic, right? So there are things that we really care about preserving, but then there are a lot of things that are like bugs and things that you solved. Actually, if you keep them in context, Asian performs worse because it is confused by, by the history of it. So you need to be very prudent about knowing what to delete and knowing what to, what to keep in the background. We have sort of a graph-like structure of understanding the different memories, and we do compaction on, um, actually there's like multiple layers. There's compaction, but there's also writing to long-term memory, and long-term memory today is marked down files. So repli.md, I'm sure you look at it sometimes. Repli.md is an example of that, but there are other files that the agent can write to, and you can also prompt it to write its own long-term memory. You can say…

AI assessment note: “No, no, no, no. Just ours is better.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.