Methodology

What powers the Ledger, what the numbers mean, and why you can check every one of them.

How this was built

The 20VC Ledger is Starzero's video intelligence pipeline pointed at one show's entire archive: 64 episodes across 11 years, 2,544 statements. Every one is playable at its exact moment and attributed to a real person, and the checkable ones are checked against cited public sources. An agentic research harness assembled it in days, and it re-runs itself when new episodes drop.

Why you can't build this from scraped captions

The obvious shortcut is "pull the YouTube transcripts and prompt a model." It fails within a day. We measured why:

The harness

On that substrate runs an agentic research harness. It reads every scene, extracts what was actually claimed, checks every quote verbatim against the tape, resolves who said it, and checks every checkable claim against grounded web sources that are cited on the claim page. The assessment looks at the outside world rather than the transcript. Resolution criteria are written down before the assessment runs; low confidence leaves a claim open, never a settled result; assessment history is never overwritten. Then the part we're proudest of: it audits itself adversarially. A stronger model re-examines its assessments under instructions to refute them, and retractions happen in the open, never silently.

One standing rule governs every metric we publish: 132 guests appear in both produced audio and raw recordings, and any score must give the same person the same result on either tape. A metric that fails that test is measuring the microphone chain, and it doesn't ship. The stack that built this Ledger is the same one Starzero runs for enterprise video archives. 20VC is one channel; the pipeline doesn't care whose feed it is.

Reading the assessments

The argument clarity score, and how speaking style is treated

The headline number on every person's page is their argument clarity score, out of 5: do they actually answer what they're asked? Every question → answer exchange is assessed with names hidden, so prominence cannot tilt the score, across four dimensions (directness, coherence, precision, compression, 1 to 5 each). It is explicitly instructed to ignore disfluencies and read meaning alone. It also classifies each exchange answered / partly / redirected / not addressed, with a required evidence quote. That classification is what Straight Answers reports, always shown with its sample size. The host is not exempt: when a guest turns the tables and asks him a question, his answer goes through the same rubric. Corpus wide, 88% of questions get answered, 6% partly, 6% redirected, 1% not addressed.

It is a score against a rubric. Nobody is "worst" by construction. The corpus spans roughly 3.4 to 4.6, which is what you'd expect from a room this good: everyone competent, some exceptional. About 20 exchanges are assessed per person, on raw unedited episodes only, because the edit removes more than "um"s. It removes tangents, so even meaning level clarity inherits the edit. Small samples are shrunk toward the cohort mean rather than allowed to spike. The rubric passed its own audits: evasive answers do not drag the other dimensions down with them.

When a full score is impossible, we estimate, and we say so. A full score needs 8 or more assessed exchanges on raw tape. People below that get a coarse estimate instead, marked with ≈, drawn on a blue bar rather than gold, and rounded to the nearest half point: we pool every assessed exchange that exists for them, raw or produced, and shrink hard toward the cohort mean. Produced feed audio scores about 0.3 higher (the Edit Room measures this on people who appear on both tapes), which is why estimates never get finer than half a point. An estimate needs at least 3 exchanges; below that the page shows a dash and the hover text says exactly what is missing.

Speaking style is a different thing, so we only measure it. Person pages show the raw numbers (um and uh per 1k words, articulation speed, and on some cohorts false starts and pause placement) next to the corpus median, and that's all. No percentile, no badge, no ranking. The numbers come from listening to the audio itself: a speech model transcribes every episode verbatim, hesitations kept, its words are aligned onto our timed and diarised stream, and a filled pause is attributed to a speaker only where that alignment is unambiguous. Speech-to-text transcripts silently drop about a quarter of real hesitations, so counting them on the transcript would under-report everyone; listening is the honest way. Articulation speed comes from the word timings, pauses excluded. One caveat survives: most produced podcast audio has some stumbles edited out before any listener hears them. We proved this with 132 guests who appear on both raw and produced tape, same person, visibly different edit. So a figure measured on raw-level tape is shown plainly, and a figure measured only on an edited feed is marked as a floor: the person said at least that much, the editor cut the rest. And fluency is not thought: across this corpus, speaking smoothness and argument quality are nearly uncorrelated (ρ≈0.2). A careful fund manager choosing words slowly and a polished talker saying nothing both exist in this dataset. The argument clarity score tells them apart; a fluency rank never could.

Interview dynamics: the power balance, chapter by chapter

Every episode page carries a map of how that specific conversation went. The chapters come from the episode's own scene boundaries, detected upstream from topic shifts rather than invented by the scorer, so segmentation is reproducible and comparable between episodes. Each chapter is scored 0 to 10 on four independent scales:

Each score ships with the reasoning that justified it. Hover any point on the chart, or open the segment table underneath. A number you can't audit is a number you shouldn't trust. Monologue chapters force the host side scales to 0 so the model can't invent host activity, and the prompt explicitly instructs it not to inflate disagreement: most conversations are respectful, and anything scoring 7+ needs quotable evidence. Four episode wide standouts are extracted as playable clips. The speaking balance strip beneath the chart is pure arithmetic on the diarized turns. No model is involved.

Honest limits. This is rhetorical analysis of a transcript, so tone of voice is invisible: someone snarling a polite sentence, or needling with a smile, won't register. Re-running an episode moves individual chapter scores by a point or two. The shape of the curves is stable, single points are noisy, so don't over-read one chapter. On panels with several guests, "disagreement toward the host" necessarily blurs across guests. This replaced an earlier pass that scored each interview holistically in one shot; across all 1,474 episodes the two agree well, which is the reassuring result you want from two independent methods. Where they differ, we keep the chapter based number: averaging a dozen audited segments is more stable than one impression of a whole episode. Every episode level dynamics figure elsewhere on the site is now derived from these chapter scores rather than measured separately, so the site never shows two different numbers for the same thing.

Why quotes sometimes contain typos, and how assessments handle them

Quotes on this site are verbatim from the automated transcript, deliberately. We show you the tape as the machine heard it, without cleanup, because the source is the point. Speech to text mishears things, especially names: in one episode "YouChat" was transcribed as "uChap". So can you trust the analysis? Yes, because no assessment is ever decided by the raw transcript alone. The pipeline reads the full conversation in context: that garbled quote shipped as the claim "YouChat (by You.com) was the first language model with a search backend, citations, and web links". The product name was recovered from context even though the transcript never once spelled it right. A second model re-checks that every quote is real and correctly attributed, and every assessment is made against grounded web sources that are cited on the claim page, external evidence rather than the transcript. Assessments use context and outside sources, so a transcription error rarely changes one; it can. When a transcript is too broken to support a claim's meaning, the claim is dropped in verification rather than assessed.

We audit our assessments. Here is the score

This site was built by AI, and AI makes mistakes, so we hunt them on the record. Every audit uses a stronger model than the one being audited, with the full scene transcript and explicit instructions to refute our own work:

Honest limitations

Independent project. Not affiliated with 20VC or Harry Stebbings. Built on the Starzero video intelligence pipeline and an agentic research harness; every assessment cites grounded web sources on its claim page. On-page answers come from an AI analyst that searches the ledger itself and cites every source it uses. The graph's layout is computed from claim embeddings. Corrections: open an issue.

Made with StarZero

Turn any episode into a week of clips.

This entire site, over 60 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.