Vals.ai

other on 1 show · 1 statements across 1 episodes · said 13 times in 1 episodes since 2026

the a16z Podcast 13

Mentions by year, every show

tap a year for its mentions
00811512026episodesmentions
0112026episodes it came up in
007.50.51512026episodesmentions per episode

the a16z Podcast 13

2026 13 mentions in 1 episode

every mention on every show, scene by scene, with the transcript →

5 statements about Vals.ai, every show

a16z Opinion
Chi: AI cybersecurity risks are primarily in infrastructure, not code
“But actually a lot of the biggest concern or risk is in the infrastructure level. And so these are not things that are expressed in code, but take simulating larger environments of enterprise cloud infrastructure, or even grid infrastructure, for us to be able…”
Ryan Chi Sep 9, 2026 ▶ 38:15 Inside the Race to Measure Frontier Intelligence
a16z Assertion Supported
Chi: Models tested for cybersecurity risks are actively reward hacking
“I think there's places where you see that born out now where models that are being tested for one cybersecurity risk are actually reward hacking and figuring out other ways to get around it.”
Ryan Chi Sep 9, 2026 ▶ 28:41 Inside the Race to Measure Frontier Intelligence
a16z Opinion
Chi: AI policy discussions have been too abstract to define regulation
“I think the main issue though is that policy conversations as they've happened over the last couple of years have been very abstract. And there, there's been no material grounding to figure out what policy should cover.”
Ryan Chi Sep 9, 2026 ▶ 25:40 Inside the Race to Measure Frontier Intelligence
a16z Disclosure
Vals AI is testing models on tasks running across hours to weeks
“Now we're testing models and their ability to run over hours, days, sometimes weeks. And so the infrastructure needs to be very stable to support evaluation over time.”
Ryan Chi Sep 9, 2026 ▶ 14:22 Inside the Race to Measure Frontier Intelligence
a16z Assertion Not checkable as stated
Chi: Meta Llama 4 Underperformed on Private Benchmarks Despite Public Scores
“One of the early indications of that you saw was when Meta released Lama four that was a bit of a disaster, and interestingly, what we saw is that on our held out private benchmarks, the model is actually underperforming, but on all of the major public benchma…”
Ryan Chi Sep 9, 2026 ▶ 2:38 Inside the Race to Measure Frontier Intelligence

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.