The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Ankur Goyal no published score: only 1 usable exchange on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
1exchanges match
1on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q that you don't know that you don't know that is going to catch you off guard. Like some competitor is going to come in and eat your lunch or whatever, right? So to me, that is the, I think the debate of like why people are moving from offline to online or online to vibechecks is because it costs a lot of effort to do offline and it probably lasts

A Maybe six months if you're lucky. Well, I think, I think you're conflating two different things. One thing is sitting in a room and enumerating a set of scenarios or tests that you think represent a workload. And then the other thing is choosing to run those tests outside of production. And both of those things are activities that can happen offline. But I don't think anyone who's legit who's doing offline evals is doing the former anymore. Like, In fact, when we started BrainTrust and we talked to customers, we're, um, sort of very excited about BrainTrust helping them create golden data sets. And now I actually think that term is kind of a dirty term. People don't really want to create golden data sets. It's, I think it's, it's, it's often a wasted effort to, to the point that you're making. I think the best teams view offline evals as a mechanism of reconciling what they see in production with real users who are using the product with Tests and iteration that they can do offline. Like the best teams that we work with on a daily or even more frequently sometimes basis are discovering use cases from logs online and then pulling them into their environments and then playing with them. And I think the difference between doing offline evals and online evals or offline evals and just pure vibe checks is I think the pure vibe check or pure online version of this is you Observe some…

AI assessment note: “Well, I think, I think you're conflating two different things.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.