Assertion Not checkable as stated
Chi: Meta Llama 4 Underperformed on Private Benchmarks Despite Public Scores
“One of the early indications of that you saw was when Meta released Lama four that was a bit of a disaster, and interestingly, what we saw is that on our held out private benchmarks, the model is actually underperforming, but on all of the major public benchma…”
Opinion
Chi: AI data vendors create gimmick benchmarks to sell data
“And actually a lot of that industry has now Built these gimmick style benchmarks as a mechanism to sell their data. And so that, that's become kind of their go-to-market as well.”
Insight
Chi: Bundling model evaluation with data consulting produces pay-to-win benchmarks
“If you look at auditing as an industry, you end up with issues like Enron, where if you have the same group who's responsible for doing the audit, as well as also consulting and supporting the company, you have a mixed incentive structure, and then it just bec…”
Assertion Not checkable as stated
OpenRouter functions predominantly as a gateway rather than an automated router
“Open route is a bit of a misnomer in that most of their usage comes from being a model gateway. And so it's actually up to their users to decide which models they want to use when.”
Assertion Not checkable as stated
Chi: Anthropic operates on narrow margins due to high serving costs
“Anthropic is running on pretty narrow margins to support this. And they have, you know, massive cost to serve these models.”
Prediction Not checkable as stated
Chi: Enterprise AI token spend may start to eclipse salary spend
“Token spend may start to eclipse salary spend.”
Assertion Not checkable as stated
Chi: Claude Sonnet Often Costs More Than Opus Due to Token Appetite
“We're actually seeing in a lot of cases, Sonnet is more expensive than Opus because it is so token hungry.”
Assertion Supported
Chi: Models tested for cybersecurity risks are actively reward hacking
“I think there's places where you see that born out now where models that are being tested for one cybersecurity risk are actually reward hacking and figuring out other ways to get around it.”
Assertion Not checkable as stated
Chi: Fortune 10 firm's daily Claude Code limit shifted peak work hours
“I have a small anecdote related to this actually, you know, was meeting with a company and the fortune 10 and they, the way that they've adopted cloud code has been with roughly a hundred dollar a day budget for their engineers. And so what I was hearing is th…”
Assertion Not checkable as stated
Chi: Vals consumed $1.5M in model tokens in one month, 10x salaries
“In that month we spent roughly 1.5 million dollars worth of tokens. This is free, by the way. I, no, I don't want to but it was actually 10 X more we were spending in tokens than employee salary for that month.”
Insight
Chi: Fragmented sovereign AI infrastructure is extremely capital inefficient
“If I was taking a God's eye view, it would be extremely inefficient to build all of these data centers and replicate this data engineering process and train these very large models when in fact you could probably consolidate a lot of these efforts but it seems…”
Insight
Chi: Legible evaluation methods are a primary driver of model capability
“What was very clear to me was the very tight relationship between what it takes to build new systems for generation, and actually new mechanisms for evaluation. In fact, in order to get one, you often need to get better at the other. And actually one of the bi…”
Insight
Chi: Trillion-Dollar Industries Require Independent Testing and Auditing Groups
“Every time a new trillion dollar industry emerges there, there's a need for this independent testing group”
Prediction Not checkable as stated
Chi: Legible enterprise evals will be AI adoption's biggest long-term bottleneck
“And I think long-term that will be actually the biggest bottleneck, our ability to take companies and their evals and make them legible because that's how we'll figure out what signal we hill climb on and where we actually adopt.”
Opinion
Ryan Chi: AI industry lacks shared framework for recursive self-improvement
“There isn't a shared language to talk about the RSI potential of models, and so we created this as an apples to apples way to actually benchmark across the models.”
Insight
Ryan Chi: Direct recursive self-improvement testing is too slow and expensive
“In an ideal world, what you want to do is actually take a frontier model and have it train the next version of itself and see where the delta comes from. But obviously that's very expensive and slow.”
Insight
Chi: Retiring AI benchmarks is necessary to reflect current real-world knowledge
“There's another component of retiring benchmarks, which I think is, is underappreciated which is that benchmark should also be reflective of the current state of the world.”
Insight
Complex AI evals require smaller sample sizes and broader criteria
“Evaluations as they become more complex, Have a fewer sample size, but a larger set of criteria or expectations of them.”
Insight
Building evals is the hardest part of model routing
“Really the hardest part of routing is building the evals and trying to determine in what places a set of intelligences should be used for a particular application.”
Prediction Not checkable as stated
Chi: High-Performing Coding Agents Will Also Automate Excel and PowerPoint Tasks
“I think coding is a sign for what's to come in every domain. And a lot of the primitives established there are carrying over to other places. You know, if you have a very good coding agent chances are you have a model that can also make PowerPoint slides or DC…”
Opinion
Chi: AI policy discussions have been too abstract to define regulation
“I think the main issue though is that policy conversations as they've happened over the last couple of years have been very abstract. And there, there's been no material grounding to figure out what policy should cover.”
Opinion
Chi: AI cybersecurity risks are primarily in infrastructure, not code
“But actually a lot of the biggest concern or risk is in the infrastructure level. And so these are not things that are expressed in code, but take simulating larger environments of enterprise cloud infrastructure, or even grid infrastructure, for us to be able…”
Disclosure
Chi: Vals automates human evaluation work with internal system Steve
“We also have this internal system called Steve. Steve the Economic Vals employee. And so that, that's been a mechanism by which we're able to actually take more of the human work over time and put it into Steve.”
Disclosure
Chi: Vals AI committed to never sell training data to labs
“At VALS, one very early decision we made was the decision to never sell training data to labs. It's often a place that we're pushed. When we start working with a new lab to actually source and sell for them a bunch of training data.”