Sep 9, 2026 · 39m · a16z

Inside the Race to Measure Frontier Intelligence

Ryan Chi · 21m spoken Jennifer Li · 7m spoken Ben Horowitz · 6m spoken Erik Torenberg · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

On The a16z Show, Vals AI co-founder Ryan Chi and investor Ben Horowitz discuss the vital necessity of independent AI evaluation, addressing the flaws of self-reported benchmarks, soaring enterprise token economics, and geopolitical risk governance. They argue that rigorous, conflict-free third-party auditing is essential to establish transparent software markets, optimize enterprise deployment, and guide frontier safety regulation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 3.1% of the talking time here. How this is scored →

The host as informed peer 1.4 Guest teaching 3.2 Guest disagreement 0.8 The host pushing back 0.2
05100:0010:0020:0030:000:55–4:15 · The host as informed peer 3/10 The Inception of Vals and Flaws in Public Benchmarks Erik steps in to probe for historical precedents like rating agencies or audit firms. Ryan educates on the necessity of independent evals by citing Meta Llama 4 failing held-out tests despite acing public benchmarks.4:15–6:54 · The host as informed peer 3/10 The Six-Hour Pre-Release Crunch and Automating Evals with Steve Erik poses an informed question detailing the operational reality of running tens of billions of tokens before launches. Ryan details their transition from all-nighters to their automated internal system Steve.6:54–9:48 · The host as informed peer 0/10 The MPAA Analogy and Avoiding Enron-Style Conflicts of Interest Erik remains silent as Jennifer and Ben explore MPAA analogies and Ryan explains refusing to sell training data to avoid Enron-style auditing conflicts.9:48–13:23 · The host as informed peer 0/10 Popular Benchmarks and Measuring Recursive Self-Improvement Erik does not speak in this segment while Jennifer explores Vals' benchmark catalog. Ryan explains their recursive self-improvement index and why saturated benchmarks must be deprecated.13:23–16:19 · The host as informed peer 0/10 Evaluating AI Agents: Infrastructure, Long Horizons, and Complex Criteria Erik remains silent while Jennifer and Ben query Ryan on agentic evals. Ryan gently reframes Ben's conception of OpenRouter, clarifying that routing fundamentally requires evaluation rubrics.16:19–18:40 · The host as informed peer 2/10 The Existential Enterprise Dilemma: When Token Spend Eclipses Salaries Erik prompts Ryan on why evaluations are existential for enterprises. Ryan delivers a striking breakdown of a Fortune 10 firm where token rate limits warped daily work hours and token spend threatened to eclipse salaries.18:40–22:40 · The host as informed peer 2/10 Introducing ValSmith: Private Repository Benchmarking and Enterprise Workflows Erik asks how code evaluation frameworks generalize into wider knowledge work. Ryan points out that software engineering benchmarks serve as leading indicators for financial modeling and slide generation.22:40–25:04 · The host as informed peer 0/10 Vals' Internal Token-Maxing Experiment: Burning $1.5M in Model Usage Erik is silent during this segment while Jennifer probes Ryan about their internal token-burning trial. Ryan reveals their team burned 1.5 million dollars in tokens in a month, exceeding payroll tenfold.25:04–28:58 · The host as informed peer 4/10 AI Policy, Compute Thresholds, and Frontier Risk Governance Erik demonstrates strong contextual knowledge regarding the lag between regulatory pace and AI capability velocity, challenging Ryan on who should set governance standards.28:58–33:30 · The host as informed peer 2/10 Division of Labor: Government Rule-Making vs. Private Auditing Erik follows up to ask specifically how policymakers and third-party evaluators should interact. Ryan outlines their regular briefing channels with legislative and executive branches.33:30–37:14 · The host as informed peer 0/10 Geopolitics and Sovereign AI: The 'Trust but Verify' Nuclear Analogy Erik does not speak while Jennifer and Ben discuss geopolitical competition. Ryan expresses contrarian skepticism over sovereign AI investments and invokes the Reagan 'trust but verify' nuclear treaty model.37:14–38:53 · The host as informed peer 1/10 The Future of Frontier Evals: Grid-Scale Infrastructure and Pure Alignment Jennifer prompts Ryan on frontier evaluation horizons and Erik wraps up the interview. Ryan outlines expanding evaluations into grid-level infrastructure environments.0:55–4:15 · Guest teaching 4/10 The Inception of Vals and Flaws in Public Benchmarks Erik steps in to probe for historical precedents like rating agencies or audit firms. Ryan educates on the necessity of independent evals by citing Meta Llama 4 failing held-out tests despite acing public benchmarks.4:15–6:54 · Guest teaching 3/10 The Six-Hour Pre-Release Crunch and Automating Evals with Steve Erik poses an informed question detailing the operational reality of running tens of billions of tokens before launches. Ryan details their transition from all-nighters to their automated internal system Steve.6:54–9:48 · Guest teaching 3/10 The MPAA Analogy and Avoiding Enron-Style Conflicts of Interest Erik remains silent as Jennifer and Ben explore MPAA analogies and Ryan explains refusing to sell training data to avoid Enron-style auditing conflicts.9:48–13:23 · Guest teaching 3/10 Popular Benchmarks and Measuring Recursive Self-Improvement Erik does not speak in this segment while Jennifer explores Vals' benchmark catalog. Ryan explains their recursive self-improvement index and why saturated benchmarks must be deprecated.13:23–16:19 · Guest teaching 3/10 Evaluating AI Agents: Infrastructure, Long Horizons, and Complex Criteria Erik remains silent while Jennifer and Ben query Ryan on agentic evals. Ryan gently reframes Ben's conception of OpenRouter, clarifying that routing fundamentally requires evaluation rubrics.16:19–18:40 · Guest teaching 5/10 The Existential Enterprise Dilemma: When Token Spend Eclipses Salaries Erik prompts Ryan on why evaluations are existential for enterprises. Ryan delivers a striking breakdown of a Fortune 10 firm where token rate limits warped daily work hours and token spend threatened to eclipse salaries.18:40–22:40 · Guest teaching 3/10 Introducing ValSmith: Private Repository Benchmarking and Enterprise Workflows Erik asks how code evaluation frameworks generalize into wider knowledge work. Ryan points out that software engineering benchmarks serve as leading indicators for financial modeling and slide generation.22:40–25:04 · Guest teaching 4/10 Vals' Internal Token-Maxing Experiment: Burning $1.5M in Model Usage Erik is silent during this segment while Jennifer probes Ryan about their internal token-burning trial. Ryan reveals their team burned 1.5 million dollars in tokens in a month, exceeding payroll tenfold.25:04–28:58 · Guest teaching 3/10 AI Policy, Compute Thresholds, and Frontier Risk Governance Erik demonstrates strong contextual knowledge regarding the lag between regulatory pace and AI capability velocity, challenging Ryan on who should set governance standards.28:58–33:30 · Guest teaching 2/10 Division of Labor: Government Rule-Making vs. Private Auditing Erik follows up to ask specifically how policymakers and third-party evaluators should interact. Ryan outlines their regular briefing channels with legislative and executive branches.33:30–37:14 · Guest teaching 3/10 Geopolitics and Sovereign AI: The 'Trust but Verify' Nuclear Analogy Erik does not speak while Jennifer and Ben discuss geopolitical competition. Ryan expresses contrarian skepticism over sovereign AI investments and invokes the Reagan 'trust but verify' nuclear treaty model.37:14–38:53 · Guest teaching 2/10 The Future of Frontier Evals: Grid-Scale Infrastructure and Pure Alignment Jennifer prompts Ryan on frontier evaluation horizons and Erik wraps up the interview. Ryan outlines expanding evaluations into grid-level infrastructure environments.0:55–4:15 · Guest disagreement 1/10 The Inception of Vals and Flaws in Public Benchmarks Erik steps in to probe for historical precedents like rating agencies or audit firms. Ryan educates on the necessity of independent evals by citing Meta Llama 4 failing held-out tests despite acing public benchmarks.4:15–6:54 · Guest disagreement 1/10 The Six-Hour Pre-Release Crunch and Automating Evals with Steve Erik poses an informed question detailing the operational reality of running tens of billions of tokens before launches. Ryan details their transition from all-nighters to their automated internal system Steve.6:54–9:48 · Guest disagreement 1/10 The MPAA Analogy and Avoiding Enron-Style Conflicts of Interest Erik remains silent as Jennifer and Ben explore MPAA analogies and Ryan explains refusing to sell training data to avoid Enron-style auditing conflicts.9:48–13:23 · Guest disagreement 0/10 Popular Benchmarks and Measuring Recursive Self-Improvement Erik does not speak in this segment while Jennifer explores Vals' benchmark catalog. Ryan explains their recursive self-improvement index and why saturated benchmarks must be deprecated.13:23–16:19 · Guest disagreement 2/10 Evaluating AI Agents: Infrastructure, Long Horizons, and Complex Criteria Erik remains silent while Jennifer and Ben query Ryan on agentic evals. Ryan gently reframes Ben's conception of OpenRouter, clarifying that routing fundamentally requires evaluation rubrics.16:19–18:40 · Guest disagreement 0/10 The Existential Enterprise Dilemma: When Token Spend Eclipses Salaries Erik prompts Ryan on why evaluations are existential for enterprises. Ryan delivers a striking breakdown of a Fortune 10 firm where token rate limits warped daily work hours and token spend threatened to eclipse salaries.18:40–22:40 · Guest disagreement 1/10 Introducing ValSmith: Private Repository Benchmarking and Enterprise Workflows Erik asks how code evaluation frameworks generalize into wider knowledge work. Ryan points out that software engineering benchmarks serve as leading indicators for financial modeling and slide generation.22:40–25:04 · Guest disagreement 0/10 Vals' Internal Token-Maxing Experiment: Burning $1.5M in Model Usage Erik is silent during this segment while Jennifer probes Ryan about their internal token-burning trial. Ryan reveals their team burned 1.5 million dollars in tokens in a month, exceeding payroll tenfold.25:04–28:58 · Guest disagreement 1/10 AI Policy, Compute Thresholds, and Frontier Risk Governance Erik demonstrates strong contextual knowledge regarding the lag between regulatory pace and AI capability velocity, challenging Ryan on who should set governance standards.28:58–33:30 · Guest disagreement 0/10 Division of Labor: Government Rule-Making vs. Private Auditing Erik follows up to ask specifically how policymakers and third-party evaluators should interact. Ryan outlines their regular briefing channels with legislative and executive branches.33:30–37:14 · Guest disagreement 2/10 Geopolitics and Sovereign AI: The 'Trust but Verify' Nuclear Analogy Erik does not speak while Jennifer and Ben discuss geopolitical competition. Ryan expresses contrarian skepticism over sovereign AI investments and invokes the Reagan 'trust but verify' nuclear treaty model.37:14–38:53 · Guest disagreement 0/10 The Future of Frontier Evals: Grid-Scale Infrastructure and Pure Alignment Jennifer prompts Ryan on frontier evaluation horizons and Erik wraps up the interview. Ryan outlines expanding evaluations into grid-level infrastructure environments.0:55–4:15 · The host pushing back 1/10 The Inception of Vals and Flaws in Public Benchmarks Erik steps in to probe for historical precedents like rating agencies or audit firms. Ryan educates on the necessity of independent evals by citing Meta Llama 4 failing held-out tests despite acing public benchmarks.4:15–6:54 · The host pushing back 0/10 The Six-Hour Pre-Release Crunch and Automating Evals with Steve Erik poses an informed question detailing the operational reality of running tens of billions of tokens before launches. Ryan details their transition from all-nighters to their automated internal system Steve.6:54–9:48 · The host pushing back 0/10 The MPAA Analogy and Avoiding Enron-Style Conflicts of Interest Erik remains silent as Jennifer and Ben explore MPAA analogies and Ryan explains refusing to sell training data to avoid Enron-style auditing conflicts.9:48–13:23 · The host pushing back 0/10 Popular Benchmarks and Measuring Recursive Self-Improvement Erik does not speak in this segment while Jennifer explores Vals' benchmark catalog. Ryan explains their recursive self-improvement index and why saturated benchmarks must be deprecated.13:23–16:19 · The host pushing back 0/10 Evaluating AI Agents: Infrastructure, Long Horizons, and Complex Criteria Erik remains silent while Jennifer and Ben query Ryan on agentic evals. Ryan gently reframes Ben's conception of OpenRouter, clarifying that routing fundamentally requires evaluation rubrics.16:19–18:40 · The host pushing back 0/10 The Existential Enterprise Dilemma: When Token Spend Eclipses Salaries Erik prompts Ryan on why evaluations are existential for enterprises. Ryan delivers a striking breakdown of a Fortune 10 firm where token rate limits warped daily work hours and token spend threatened to eclipse salaries.18:40–22:40 · The host pushing back 0/10 Introducing ValSmith: Private Repository Benchmarking and Enterprise Workflows Erik asks how code evaluation frameworks generalize into wider knowledge work. Ryan points out that software engineering benchmarks serve as leading indicators for financial modeling and slide generation.22:40–25:04 · The host pushing back 0/10 Vals' Internal Token-Maxing Experiment: Burning $1.5M in Model Usage Erik is silent during this segment while Jennifer probes Ryan about their internal token-burning trial. Ryan reveals their team burned 1.5 million dollars in tokens in a month, exceeding payroll tenfold.25:04–28:58 · The host pushing back 1/10 AI Policy, Compute Thresholds, and Frontier Risk Governance Erik demonstrates strong contextual knowledge regarding the lag between regulatory pace and AI capability velocity, challenging Ryan on who should set governance standards.28:58–33:30 · The host pushing back 0/10 Division of Labor: Government Rule-Making vs. Private Auditing Erik follows up to ask specifically how policymakers and third-party evaluators should interact. Ryan outlines their regular briefing channels with legislative and executive branches.33:30–37:14 · The host pushing back 0/10 Geopolitics and Sovereign AI: The 'Trust but Verify' Nuclear Analogy Erik does not speak while Jennifer and Ben discuss geopolitical competition. Ryan expresses contrarian skepticism over sovereign AI investments and invokes the Reagan 'trust but verify' nuclear treaty model.37:14–38:53 · The host pushing back 0/10 The Future of Frontier Evals: Grid-Scale Infrastructure and Pure Alignment Jennifer prompts Ryan on frontier evaluation horizons and Erik wraps up the interview. Ryan outlines expanding evaluations into grid-level infrastructure environments.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 12.4% · guest 87.6%3:00 · the host 12.4% · guest 87.6%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 4.2% · guest 95.8%15:00 · the host 4.2% · guest 95.8%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 3.7% · guest 96.3%21:00 · the host 3.7% · guest 96.3%24:00 · the host 15% · guest 85%24:00 · the host 15% · guest 85%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 4.3% · guest 95.7%30:00 · the host 4.3% · guest 95.7%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0.9% · guest 99.1%36:00 · the host 0.9% · guest 99.1%
Sharpest disagreement ▶ 15:53 Ryan reframes Ben's OpenRouter assumption

Ryan directly pushes back against Ben's suggestion that routers evaluate models in real time, explaining that OpenRouter is primarily a gateway and routing requires bespoke evals.

Hardest push from the host ▶ 25:06 Erik presses on governance standard ownership

Erik directly challenges the viability of standard-setting given the regulatory lag, pressing whether labs, private auditors, customers, or government should hold authority.

Biggest teaching moment ▶ 16:40 Ryan illustrates token budget distortions in Fortune 10 firms

Ryan educates the room with empirical observations of enterprise workflows being distorted by artificial 4 PM token resets and corporate spend nearing employee salary parity.

The host holds their own ▶ 3:40 Erik connects third-party AI evals to historical audit and credit agencies

Erik demonstrates strong structural knowledge by immediately drawing analogies to historical capital market certification institutions like rating agencies and audit firms.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
The Inception of Vals and Flaws in Public Benchmarks 3411 Erik steps in to probe for historical precedents like rating agencies or audit firms. Ryan educates on the necessity of independent evals by citing Meta Llama 4 failing held-out tests despite acing public benchmarks.
The Six-Hour Pre-Release Crunch and Automating Evals with Steve 3310 Erik poses an informed question detailing the operational reality of running tens of billions of tokens before launches. Ryan details their transition from all-nighters to their automated internal system Steve.
The MPAA Analogy and Avoiding Enron-Style Conflicts of Interest 0310 Erik remains silent as Jennifer and Ben explore MPAA analogies and Ryan explains refusing to sell training data to avoid Enron-style auditing conflicts.
Popular Benchmarks and Measuring Recursive Self-Improvement 0300 Erik does not speak in this segment while Jennifer explores Vals' benchmark catalog. Ryan explains their recursive self-improvement index and why saturated benchmarks must be deprecated.
Evaluating AI Agents: Infrastructure, Long Horizons, and Complex Criteria 0320 Erik remains silent while Jennifer and Ben query Ryan on agentic evals. Ryan gently reframes Ben's conception of OpenRouter, clarifying that routing fundamentally requires evaluation rubrics.
The Existential Enterprise Dilemma: When Token Spend Eclipses Salaries 2500 Erik prompts Ryan on why evaluations are existential for enterprises. Ryan delivers a striking breakdown of a Fortune 10 firm where token rate limits warped daily work hours and token spend threatened to eclipse salaries.
Introducing ValSmith: Private Repository Benchmarking and Enterprise Workflows 2310 Erik asks how code evaluation frameworks generalize into wider knowledge work. Ryan points out that software engineering benchmarks serve as leading indicators for financial modeling and slide generation.
Vals' Internal Token-Maxing Experiment: Burning $1.5M in Model Usage 0400 Erik is silent during this segment while Jennifer probes Ryan about their internal token-burning trial. Ryan reveals their team burned 1.5 million dollars in tokens in a month, exceeding payroll tenfold.
AI Policy, Compute Thresholds, and Frontier Risk Governance 4311 Erik demonstrates strong contextual knowledge regarding the lag between regulatory pace and AI capability velocity, challenging Ryan on who should set governance standards.
Division of Labor: Government Rule-Making vs. Private Auditing 2200 Erik follows up to ask specifically how policymakers and third-party evaluators should interact. Ryan outlines their regular briefing channels with legislative and executive branches.
Geopolitics and Sovereign AI: The 'Trust but Verify' Nuclear Analogy 0320 Erik does not speak while Jennifer and Ben discuss geopolitical competition. Ryan expresses contrarian skepticism over sovereign AI investments and invokes the Reagan 'trust but verify' nuclear treaty model.
The Future of Frontier Evals: Grid-Scale Infrastructure and Pure Alignment 1200 Jennifer prompts Ryan on frontier evaluation horizons and Erik wraps up the interview. Ryan outlines expanding evaluations into grid-level infrastructure environments.

Statements from this episode (28)

Insight
Chi: Legible evaluation methods are a primary driver of model capability
“What was very clear to me was the very tight relationship between what it takes to build new systems for generation, and actually new mechanisms for evaluation. In fact, in order to get one, you often need to get better at the other. And actually one of the bi…”
Ryan Chi Sep 9, 2026 ▶ 1:25
Assertion Not checkable as stated
Chi: Meta Llama 4 Underperformed on Private Benchmarks Despite Public Scores
“One of the early indications of that you saw was when Meta released Lama four that was a bit of a disaster, and interestingly, what we saw is that on our held out private benchmarks, the model is actually underperforming, but on all of the major public benchma…”
Ryan Chi Sep 9, 2026 ▶ 2:38
Insight
Chi: Trillion-Dollar Industries Require Independent Testing and Auditing Groups
“Every time a new trillion dollar industry emerges there, there's a need for this independent testing group”
Ryan Chi Sep 9, 2026 ▶ 3:46
Assertion Not checkable as stated
Chi: Vals runs massively distributed evals at maximum model rate limits
“Now we built up a team, but we've also really invested heavily in infrastructure. And so we're able to run evaluations in a massively distributed way running effectively the maximum possible rate limits with every model we get access to.”
Ryan Chi Sep 9, 2026 ▶ 4:58
Disclosure
Chi: Vals automates human evaluation work with internal system Steve
“We also have this internal system called Steve. Steve the Economic Vals employee. And so that, that's been a mechanism by which we're able to actually take more of the human work over time and put it into Steve.”
Ryan Chi Sep 9, 2026 ▶ 5:11
Prediction Not checkable as stated
Chi: Legible enterprise evals will be AI adoption's biggest long-term bottleneck
“And I think long-term that will be actually the biggest bottleneck, our ability to take companies and their evals and make them legible because that's how we'll figure out what signal we hill climb on and where we actually adopt.”
Ryan Chi Sep 9, 2026 ▶ 6:41
Disclosure
Chi: Vals AI committed to never sell training data to labs
“At VALS, one very early decision we made was the decision to never sell training data to labs. It's often a place that we're pushed. When we start working with a new lab to actually source and sell for them a bunch of training data.”
Ryan Chi Sep 9, 2026 ▶ 8:58
Opinion
Chi: AI data vendors create gimmick benchmarks to sell data
“And actually a lot of that industry has now Built these gimmick style benchmarks as a mechanism to sell their data. And so that, that's become kind of their go-to-market as well.”
Ryan Chi Sep 9, 2026 ▶ 9:15
Insight
Chi: Bundling model evaluation with data consulting produces pay-to-win benchmarks
“If you look at auditing as an industry, you end up with issues like Enron, where if you have the same group who's responsible for doing the audit, as well as also consulting and supporting the company, you have a mixed incentive structure, and then it just bec…”
Ryan Chi Sep 9, 2026 ▶ 9:25
Opinion
Ryan Chi: AI industry lacks shared framework for recursive self-improvement
“There isn't a shared language to talk about the RSI potential of models, and so we created this as an apples to apples way to actually benchmark across the models.”
Ryan Chi Sep 9, 2026 ▶ 10:45
Insight
Ryan Chi: Direct recursive self-improvement testing is too slow and expensive
“In an ideal world, what you want to do is actually take a frontier model and have it train the next version of itself and see where the delta comes from. But obviously that's very expensive and slow.”
Ryan Chi Sep 9, 2026 ▶ 11:08
Insight
Chi: Retiring AI benchmarks is necessary to reflect current real-world knowledge
“There's another component of retiring benchmarks, which I think is, is underappreciated which is that benchmark should also be reflective of the current state of the world.”
Ryan Chi Sep 9, 2026 ▶ 12:49
Disclosure
Vals AI is testing models on tasks running across hours to weeks
“Now we're testing models and their ability to run over hours, days, sometimes weeks. And so the infrastructure needs to be very stable to support evaluation over time.”
Ryan Chi Sep 9, 2026 ▶ 14:22
Insight
Complex AI evals require smaller sample sizes and broader criteria
“Evaluations as they become more complex, Have a fewer sample size, but a larger set of criteria or expectations of them.”
Ryan Chi Sep 9, 2026 ▶ 14:40
Assertion Not checkable as stated
OpenRouter functions predominantly as a gateway rather than an automated router
“Open route is a bit of a misnomer in that most of their usage comes from being a model gateway. And so it's actually up to their users to decide which models they want to use when.”
Ryan Chi Sep 9, 2026 ▶ 15:53
Insight
Building evals is the hardest part of model routing
“Really the hardest part of routing is building the evals and trying to determine in what places a set of intelligences should be used for a particular application.”
Ryan Chi Sep 9, 2026 ▶ 16:04
Assertion Not checkable as stated
Chi: Fortune 10 firm's daily Claude Code limit shifted peak work hours
“I have a small anecdote related to this actually, you know, was meeting with a company and the fortune 10 and they, the way that they've adopted cloud code has been with roughly a hundred dollar a day budget for their engineers. And so what I was hearing is th…”
Ryan Chi Sep 9, 2026 ▶ 16:51
Assertion Not checkable as stated
Chi: Anthropic operates on narrow margins due to high serving costs
“Anthropic is running on pretty narrow margins to support this. And they have, you know, massive cost to serve these models.”
Ryan Chi Sep 9, 2026 ▶ 17:53
Prediction Not checkable as stated
Chi: Enterprise AI token spend may start to eclipse salary spend
“Token spend may start to eclipse salary spend.”
Ryan Chi Sep 9, 2026 ▶ 18:13
Assertion Not checkable as stated
Chi: Claude Sonnet Often Costs More Than Opus Due to Token Appetite
“We're actually seeing in a lot of cases, Sonnet is more expensive than Opus because it is so token hungry.”
Ryan Chi Sep 9, 2026 ▶ 21:35
Prediction Not checkable as stated
Chi: High-Performing Coding Agents Will Also Automate Excel and PowerPoint Tasks
“I think coding is a sign for what's to come in every domain. And a lot of the primitives established there are carrying over to other places. You know, if you have a very good coding agent chances are you have a model that can also make PowerPoint slides or DC…”
Ryan Chi Sep 9, 2026 ▶ 21:56
Assertion Not checkable as stated
Chi: Vals consumed $1.5M in model tokens in one month, 10x salaries
“In that month we spent roughly 1.5 million dollars worth of tokens. This is free, by the way. I, no, I don't want to but it was actually 10 X more we were spending in tokens than employee salary for that month.”
Ryan Chi Sep 9, 2026 ▶ 23:15
Opinion
Chi: AI policy discussions have been too abstract to define regulation
“I think the main issue though is that policy conversations as they've happened over the last couple of years have been very abstract. And there, there's been no material grounding to figure out what policy should cover.”
Ryan Chi Sep 9, 2026 ▶ 25:40
Assertion Supported
Chi: Models tested for cybersecurity risks are actively reward hacking
“I think there's places where you see that born out now where models that are being tested for one cybersecurity risk are actually reward hacking and figuring out other ways to get around it.”
Ryan Chi Sep 9, 2026 ▶ 28:41
Insight
Horowitz: Governments should enforce rules while private firms audit AI capabilities
“And I think the government is particularly Ill suited to do the latter, particularly over time. It's just not a good government function, but they're very good at setting the rules because they can enforce the rules. So I think that that's kind of a combinatio…”
Ben Horowitz Sep 9, 2026 ▶ 30:30
Assertion Supported
Li: Chinese open-source AI models cannot freely discuss the CPC or history
“I use a lot of like Chinese open source models too. Like you still cannot let them, you know, just go freely talk about CPC and all the history there because you know, what happens in China.”
Jennifer Li Sep 9, 2026 ▶ 33:57
Insight
Chi: Fragmented sovereign AI infrastructure is extremely capital inefficient
“If I was taking a God's eye view, it would be extremely inefficient to build all of these data centers and replicate this data engineering process and train these very large models when in fact you could probably consolidate a lot of these efforts but it seems…”
Ryan Chi Sep 9, 2026 ▶ 34:32
Opinion
Chi: AI cybersecurity risks are primarily in infrastructure, not code
“But actually a lot of the biggest concern or risk is in the infrastructure level. And so these are not things that are expressed in code, but take simulating larger environments of enterprise cloud infrastructure, or even grid infrastructure, for us to be able…”
Ryan Chi Sep 9, 2026 ▶ 38:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.