The Ledger, every show

Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.

shows every show 44 of 44
every show
clear all ✕
a16z Assertion Not checkable as stated
Chi: Meta Llama 4 Underperformed on Private Benchmarks Despite Public Scores
“One of the early indications of that you saw was when Meta released Lama four that was a bit of a disaster, and interestingly, what we saw is that on our held out private benchmarks, the model is actually underperforming, but on all of the major public benchma…”
Ryan Chi Sep 9, 2026 ▶ 2:38 Inside the Race to Measure Frontier Intelligence
a16z Assertion Not checkable as stated
OpenRouter functions predominantly as a gateway rather than an automated router
“Open route is a bit of a misnomer in that most of their usage comes from being a model gateway. And so it's actually up to their users to decide which models they want to use when.”
Ryan Chi Sep 9, 2026 ▶ 15:53 Inside the Race to Measure Frontier Intelligence
a16z Assertion Not checkable as stated
Chi: Anthropic operates on narrow margins due to high serving costs
“Anthropic is running on pretty narrow margins to support this. And they have, you know, massive cost to serve these models.”
Ryan Chi Sep 9, 2026 ▶ 17:53 Inside the Race to Measure Frontier Intelligence
a16z Prediction Not checkable as stated
Chi: Enterprise AI token spend may start to eclipse salary spend
“Token spend may start to eclipse salary spend.”
Ryan Chi Sep 9, 2026 ▶ 18:13 Inside the Race to Measure Frontier Intelligence
a16z Assertion Not checkable as stated
Chi: Claude Sonnet Often Costs More Than Opus Due to Token Appetite
“We're actually seeing in a lot of cases, Sonnet is more expensive than Opus because it is so token hungry.”
Ryan Chi Sep 9, 2026 ▶ 21:35 Inside the Race to Measure Frontier Intelligence
a16z Assertion Supported
Chi: Models tested for cybersecurity risks are actively reward hacking
“I think there's places where you see that born out now where models that are being tested for one cybersecurity risk are actually reward hacking and figuring out other ways to get around it.”
Ryan Chi Sep 9, 2026 ▶ 28:41 Inside the Race to Measure Frontier Intelligence
a16z Assertion Not checkable as stated
Chi: Fortune 10 firm's daily Claude Code limit shifted peak work hours
“I have a small anecdote related to this actually, you know, was meeting with a company and the fortune 10 and they, the way that they've adopted cloud code has been with roughly a hundred dollar a day budget for their engineers. And so what I was hearing is th…”
Ryan Chi Sep 9, 2026 ▶ 16:51 Inside the Race to Measure Frontier Intelligence
a16z Assertion Not checkable as stated
Chi: Vals consumed $1.5M in model tokens in one month, 10x salaries
“In that month we spent roughly 1.5 million dollars worth of tokens. This is free, by the way. I, no, I don't want to but it was actually 10 X more we were spending in tokens than employee salary for that month.”
Ryan Chi Sep 9, 2026 ▶ 23:15 Inside the Race to Measure Frontier Intelligence
a16z Prediction Not checkable as stated
Chi: Legible enterprise evals will be AI adoption's biggest long-term bottleneck
“And I think long-term that will be actually the biggest bottleneck, our ability to take companies and their evals and make them legible because that's how we'll figure out what signal we hill climb on and where we actually adopt.”
Ryan Chi Sep 9, 2026 ▶ 6:41 Inside the Race to Measure Frontier Intelligence
a16z Prediction Not checkable as stated
Chi: High-Performing Coding Agents Will Also Automate Excel and PowerPoint Tasks
“I think coding is a sign for what's to come in every domain. And a lot of the primitives established there are carrying over to other places. You know, if you have a very good coding agent chances are you have a model that can also make PowerPoint slides or DC…”
Ryan Chi Sep 9, 2026 ▶ 21:56 Inside the Race to Measure Frontier Intelligence
a16z Assertion Not checkable as stated
Chi: Vals runs massively distributed evals at maximum model rate limits
“Now we built up a team, but we've also really invested heavily in infrastructure. And so we're able to run evaluations in a massively distributed way running effectively the maximum possible rate limits with every model we get access to.”
Ryan Chi Sep 9, 2026 ▶ 4:58 Inside the Race to Measure Frontier Intelligence
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.