The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 8 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Insight
Chi: Bundling model evaluation with data consulting produces pay-to-win benchmarks
“If you look at auditing as an industry, you end up with issues like Enron, where if you have the same group who's responsible for doing the audit, as well as also consulting and supporting the company, you have a mixed incentive structure, and then it just bec…”
Ryan Chi Sep 9, 2026 ▶ 9:25 Inside the Race to Measure Frontier Intelligence
Insight
Chi: Fragmented sovereign AI infrastructure is extremely capital inefficient
“If I was taking a God's eye view, it would be extremely inefficient to build all of these data centers and replicate this data engineering process and train these very large models when in fact you could probably consolidate a lot of these efforts but it seems…”
Ryan Chi Sep 9, 2026 ▶ 34:32 Inside the Race to Measure Frontier Intelligence
Insight
Chi: Legible evaluation methods are a primary driver of model capability
“What was very clear to me was the very tight relationship between what it takes to build new systems for generation, and actually new mechanisms for evaluation. In fact, in order to get one, you often need to get better at the other. And actually one of the bi…”
Ryan Chi Sep 9, 2026 ▶ 1:25 Inside the Race to Measure Frontier Intelligence
Insight
Chi: Trillion-Dollar Industries Require Independent Testing and Auditing Groups
“Every time a new trillion dollar industry emerges there, there's a need for this independent testing group”
Ryan Chi Sep 9, 2026 ▶ 3:46 Inside the Race to Measure Frontier Intelligence
Insight
Ryan Chi: Direct recursive self-improvement testing is too slow and expensive
“In an ideal world, what you want to do is actually take a frontier model and have it train the next version of itself and see where the delta comes from. But obviously that's very expensive and slow.”
Ryan Chi Sep 9, 2026 ▶ 11:08 Inside the Race to Measure Frontier Intelligence
Insight
Chi: Retiring AI benchmarks is necessary to reflect current real-world knowledge
“There's another component of retiring benchmarks, which I think is, is underappreciated which is that benchmark should also be reflective of the current state of the world.”
Ryan Chi Sep 9, 2026 ▶ 12:49 Inside the Race to Measure Frontier Intelligence
Insight
Complex AI evals require smaller sample sizes and broader criteria
“Evaluations as they become more complex, Have a fewer sample size, but a larger set of criteria or expectations of them.”
Ryan Chi Sep 9, 2026 ▶ 14:40 Inside the Race to Measure Frontier Intelligence
Insight
Building evals is the hardest part of model routing
“Really the hardest part of routing is building the evals and trying to determine in what places a set of intelligences should be used for a particular application.”
Ryan Chi Sep 9, 2026 ▶ 16:04 Inside the Race to Measure Frontier Intelligence
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.