Methodology
What powers the Ledger, what the numbers mean, and why you can check every one of them.
How this was built
The 20VC Ledger is Starzero's video intelligence pipeline pointed at one show's entire archive: 1,508 episodes across 11 years, 37,600 statements. Every one is playable at its exact moment and attributed to a real person, and the checkable ones are checked against cited public sources. An agentic research harness assembled it in days, and it re-runs itself when new episodes drop.
Why you can't build this from scraped captions
The obvious shortcut is "pull the YouTube transcripts and prompt a model." It fails within a day. We measured why:
- Captions lie by omission. Starzero's transcription preserves every "um", false start and self correction. That is how we discovered that most produced podcast audio has the stumbles edited out before any listener hears them. We can tell you, month by month, how heavily the show's editors worked; The Edit Room lays out the whole forensic case. A transcript that was cleaned before you got it can never tell you how anyone actually speaks, and it hides the editing itself.
- Captions don't know who is talking. Accountability starts with attribution, and attribution is brutally hard: the host's own name appears 44 different ways across ten years of recordings, and guests get introduced as just "Rory". Starzero's diarization plus an identity layer built for scale is why every claim here is pinned on the right human. A wrong attribution is treated as the one unforgivable failure, gated by standing audits before anything ships.
- Captions don't know where the conversation turns. Every statistic on this site is computed on Starzero's semantic scenes, the actual topic boundaries of each conversation, instead of arbitrary time chunks. That is the difference between "assessed in context" and quote soup, and it is why processing eleven years of tape costs a fraction of what it costs anyone working from raw transcripts.
The harness
On that substrate runs an agentic research harness. It reads every scene, extracts what was actually claimed, checks every quote verbatim against the tape, resolves who said it, and checks every checkable claim against grounded web sources that are cited on the claim page. The assessment looks at the outside world rather than the transcript. Resolution criteria are written down before the assessment runs; low confidence leaves a claim open, never a settled result; assessment history is never overwritten. Then the part we're proudest of: it audits itself adversarially. A stronger model re-examines its assessments under instructions to refute them, and retractions happen in the open, never silently.
One standing rule governs every metric we publish: 132 guests appear in both produced audio and raw recordings, and any score must give the same person the same result on either tape. A metric that fails that test is measuring the microphone chain, and it doesn't ship. The stack that built this Ledger is the same one Starzero runs for enterprise video archives. 20VC is one channel; the pipeline doesn't care whose feed it is.
Reading the assessments
- Held up / Didn't hold up: predictions only. They record how the call stands against public evidence by the prediction date. We assess the call rather than the person; hyperbole isn't perjury.
- Supported / Contradicted: checkable assertions, measured against cited public sources as the record stands. Contradicted means the sources say otherwise, not that we looked and found nothing.
- Partly held up / Partly supported: materially right and wrong in parts. The reasoning explains the split.
- Open: the evidence doesn't exist yet. A review date is set.
- Not publicly verifiable: private information. No public evidence class exists.
The argument clarity score, and how speaking style is treated
The headline number on every person's page is their argument clarity score, out of 5: do they actually answer what they're asked? Every question → answer exchange is assessed with names hidden, so prominence cannot tilt the score, across four dimensions (directness, coherence, precision, compression, 1 to 5 each). It is explicitly instructed to ignore disfluencies and read meaning alone. It also classifies each exchange answered / partly / redirected / not addressed, with a required evidence quote. That classification is what Straight Answers reports, always shown with its sample size. The host is not exempt: when a guest turns the tables and asks him a question, his answer goes through the same rubric. Corpus wide, 88% of questions get answered, 6% partly, 6% redirected, 1% not addressed.
It is a score against a rubric. Nobody is "worst" by construction. The corpus spans roughly 3.4 to 4.6, which is what you'd expect from a room this good: everyone competent, some exceptional. About 20 exchanges are assessed per person, on raw unedited episodes only, because the edit removes more than "um"s. It removes tangents, so even meaning level clarity inherits the edit. Small samples are shrunk toward the cohort mean rather than allowed to spike. The rubric passed its own audits: evasive answers do not drag the other dimensions down with them.
When a full score is impossible, we estimate, and we say so. A full score needs 8 or more assessed exchanges on raw tape. People below that get a coarse estimate instead, marked with ≈, drawn on a blue bar rather than gold, and rounded to the nearest half point: we pool every assessed exchange that exists for them, raw or produced, and shrink hard toward the cohort mean. Produced feed audio scores about 0.3 higher (the Edit Room measures this on people who appear on both tapes), which is why estimates never get finer than half a point. An estimate needs at least 3 exchanges; below that the page shows a dash and the hover text says exactly what is missing.
Speaking style is a different thing, so we only measure it. Person pages show the raw numbers (um and uh per 1k words, articulation speed, and on some cohorts false starts and pause placement) next to the corpus median, and that's all. No percentile, no badge, no ranking. The numbers come from listening to the audio itself: a speech model transcribes every episode verbatim, hesitations kept, its words are aligned onto our timed and diarised stream, and a filled pause is attributed to a speaker only where that alignment is unambiguous. Speech-to-text transcripts silently drop about a quarter of real hesitations, so counting them on the transcript would under-report everyone; listening is the honest way. Articulation speed comes from the word timings, pauses excluded. One caveat survives: most produced podcast audio has some stumbles edited out before any listener hears them. We proved this with 132 guests who appear on both raw and produced tape, same person, visibly different edit. So a figure measured on raw-level tape is shown plainly, and a figure measured only on an edited feed is marked as a floor: the person said at least that much, the editor cut the rest. And fluency is not thought: across this corpus, speaking smoothness and argument quality are nearly uncorrelated (ρ≈0.2). A careful fund manager choosing words slowly and a polished talker saying nothing both exist in this dataset. The argument clarity score tells them apart; a fluency rank never could.
Interview dynamics: the power balance, chapter by chapter
Every episode page carries a map of how that specific conversation went. The chapters come from the episode's own scene boundaries, detected upstream from topic shifts rather than invented by the scorer, so segmentation is reproducible and comparable between episodes. Each chapter is scored 0 to 10 on four independent scales:
- Harry as informed peer: is the host drilling with facts, sources and counterexamples, or just listening? 0 means purely receiving.
- Guest teaching: is the guest explaining or correcting something the host didn't have? This is deliberately not the opposite of the first scale: in a genuine peer exchange both run moderate, and in small talk both sit near zero.
- Guest disagreement: rejecting premises, contrarian framing, refusing the question's assumptions.
- Harry pushing back: refusing a framing, restating a question, pressing after a swerve. A guest can challenge hard while the host accepts it, and the host can push a perfectly agreeable guest, so these two are scored separately too.
Each score ships with the reasoning that justified it. Hover any point on the chart, or open the segment table underneath. A number you can't audit is a number you shouldn't trust. Monologue chapters force the host side scales to 0 so the model can't invent host activity, and the prompt explicitly instructs it not to inflate disagreement: most conversations are respectful, and anything scoring 7+ needs quotable evidence. Four episode wide standouts are extracted as playable clips. The speaking balance strip beneath the chart is pure arithmetic on the diarized turns. No model is involved.
Honest limits. This is rhetorical analysis of a transcript, so tone of voice is invisible: someone snarling a polite sentence, or needling with a smile, won't register. Re-running an episode moves individual chapter scores by a point or two. The shape of the curves is stable, single points are noisy, so don't over-read one chapter. On panels with several guests, "disagreement toward the host" necessarily blurs across guests. This replaced an earlier pass that scored each interview holistically in one shot; across all 1,474 episodes the two agree well, which is the reassuring result you want from two independent methods. Where they differ, we keep the chapter based number: averaging a dozen audited segments is more stable than one impression of a whole episode. Every episode level dynamics figure elsewhere on the site is now derived from these chapter scores rather than measured separately, so the site never shows two different numbers for the same thing.
Why quotes sometimes contain typos, and how assessments handle them
Quotes on this site are verbatim from the automated transcript, deliberately. We show you the tape as the machine heard it, without cleanup, because the source is the point. Speech to text mishears things, especially names: in one episode "YouChat" was transcribed as "uChap". So can you trust the analysis? Yes, because no assessment is ever decided by the raw transcript alone. The pipeline reads the full conversation in context: that garbled quote shipped as the claim "YouChat (by You.com) was the first language model with a search backend, citations, and web links". The product name was recovered from context even though the transcript never once spelled it right. A second model re-checks that every quote is real and correctly attributed, and every assessment is made against grounded web sources that are cited on the claim page, external evidence rather than the transcript. Assessments use context and outside sources, so a transcription error rarely changes one; it can. When a transcript is too broken to support a claim's meaning, the claim is dropped in verification rather than assessed.
We audit our assessments. Here is the score
This site was built by AI, and AI makes mistakes, so we hunt them on the record. Every audit uses a stronger model than the one being audited, with the full scene transcript and explicit instructions to refute our own work:
- Audit 1, contradicted assessments on the weakest records (July 2026): all 12 contradicted assessments held by the 10 lowest scoring people. Result: 2 of 12 unsound (17%), both retracted.
- Audit 2, every negative assessment on the ledger (July 2026): all 502 contradicted and partly supported assessments. Result: 98 unsound (19.5%). 39 assessed a stricter claim than the speaker made, 28 penalised conversational hyperbole, 17 were transcript mishears, 10 weren't real assertions, 4 other. All 98 retracted and reopened. The resolution history stays visible, and retractions are permanent in the export pipeline.
- Audit 3, faithfulness of every visible claim: running continuously. Severity 3 problems (should not ship) are hidden rather than patched quietly.
- Your reports count: every claim page has a "spot an error?" button, and reports feed the next audit pass. Nothing is auto changed by a report, and nothing is deleted by an audit. Assessments are only ever retracted in the open.
Honest limitations
- Extraction and assessment use AI models; both can err. Every claim links to the exact video moment. The tape is the ground truth, and it's one click away.
- Quotes are verbatim from automated speech to text and can contain mishearings (see above). The playable clip is always the authority.
- Certainty scores measure how the claim was phrased, not what the speaker privately believed.
- Speakers control what they say on a podcast; a full support rate here measures public statements only.
Independent project. Not affiliated with 20VC or Harry Stebbings. Built on the Starzero video intelligence pipeline and an agentic research harness; every assessment cites grounded web sources on its claim page. On-page answers come from an AI analyst that searches the ledger itself and cites every source it uses. The graph's layout is computed from claim embeddings. Corrections: open an issue.