Sep 9, 2026 · 39m · a16z
Inside the Race to Measure Frontier Intelligence
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
On The a16z Show, Vals AI co-founder Ryan Chi and investor Ben Horowitz discuss the vital necessity of independent AI evaluation, addressing the flaws of self-reported benchmarks, soaring enterprise token economics, and geopolitical risk governance. They argue that rigorous, conflict-free third-party auditing is essential to establish transparent software markets, optimize enterprise deployment, and guide frontier safety regulation.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 3.1% of the talking time here. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Ryan directly pushes back against Ben's suggestion that routers evaluate models in real time, explaining that OpenRouter is primarily a gateway and routing requires bespoke evals.
Hardest push from the host ▶ 25:06 Erik presses on governance standard ownershipErik directly challenges the viability of standard-setting given the regulatory lag, pressing whether labs, private auditors, customers, or government should hold authority.
Biggest teaching moment ▶ 16:40 Ryan illustrates token budget distortions in Fortune 10 firmsRyan educates the room with empirical observations of enterprise workflows being distorted by artificial 4 PM token resets and corporate spend nearing employee salary parity.
The host holds their own ▶ 3:40 Erik connects third-party AI evals to historical audit and credit agenciesErik demonstrates strong structural knowledge by immediately drawing analogies to historical capital market certification institutions like rating agencies and audit firms.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| The Inception of Vals and Flaws in Public Benchmarks | 3 | 4 | 1 | 1 | Erik steps in to probe for historical precedents like rating agencies or audit firms. Ryan educates on the necessity of independent evals by citing Meta Llama 4 failing held-out tests despite acing public benchmarks. | |
| The Six-Hour Pre-Release Crunch and Automating Evals with Steve | 3 | 3 | 1 | 0 | Erik poses an informed question detailing the operational reality of running tens of billions of tokens before launches. Ryan details their transition from all-nighters to their automated internal system Steve. | |
| The MPAA Analogy and Avoiding Enron-Style Conflicts of Interest | 0 | 3 | 1 | 0 | Erik remains silent as Jennifer and Ben explore MPAA analogies and Ryan explains refusing to sell training data to avoid Enron-style auditing conflicts. | |
| Popular Benchmarks and Measuring Recursive Self-Improvement | 0 | 3 | 0 | 0 | Erik does not speak in this segment while Jennifer explores Vals' benchmark catalog. Ryan explains their recursive self-improvement index and why saturated benchmarks must be deprecated. | |
| Evaluating AI Agents: Infrastructure, Long Horizons, and Complex Criteria | 0 | 3 | 2 | 0 | Erik remains silent while Jennifer and Ben query Ryan on agentic evals. Ryan gently reframes Ben's conception of OpenRouter, clarifying that routing fundamentally requires evaluation rubrics. | |
| The Existential Enterprise Dilemma: When Token Spend Eclipses Salaries | 2 | 5 | 0 | 0 | Erik prompts Ryan on why evaluations are existential for enterprises. Ryan delivers a striking breakdown of a Fortune 10 firm where token rate limits warped daily work hours and token spend threatened to eclipse salaries. | |
| Introducing ValSmith: Private Repository Benchmarking and Enterprise Workflows | 2 | 3 | 1 | 0 | Erik asks how code evaluation frameworks generalize into wider knowledge work. Ryan points out that software engineering benchmarks serve as leading indicators for financial modeling and slide generation. | |
| Vals' Internal Token-Maxing Experiment: Burning $1.5M in Model Usage | 0 | 4 | 0 | 0 | Erik is silent during this segment while Jennifer probes Ryan about their internal token-burning trial. Ryan reveals their team burned 1.5 million dollars in tokens in a month, exceeding payroll tenfold. | |
| AI Policy, Compute Thresholds, and Frontier Risk Governance | 4 | 3 | 1 | 1 | Erik demonstrates strong contextual knowledge regarding the lag between regulatory pace and AI capability velocity, challenging Ryan on who should set governance standards. | |
| Division of Labor: Government Rule-Making vs. Private Auditing | 2 | 2 | 0 | 0 | Erik follows up to ask specifically how policymakers and third-party evaluators should interact. Ryan outlines their regular briefing channels with legislative and executive branches. | |
| Geopolitics and Sovereign AI: The 'Trust but Verify' Nuclear Analogy | 0 | 3 | 2 | 0 | Erik does not speak while Jennifer and Ben discuss geopolitical competition. Ryan expresses contrarian skepticism over sovereign AI investments and invokes the Reagan 'trust but verify' nuclear treaty model. | |
| The Future of Frontier Evals: Grid-Scale Infrastructure and Pure Alignment | 1 | 2 | 0 | 0 | Jennifer prompts Ryan on frontier evaluation horizons and Erik wraps up the interview. Ryan outlines expanding evaluations into grid-level infrastructure environments. |