AI coding benchmarks
2 statements across 2 episodes · 0 bullish · 1 bearish · 2 people on the record · first statement Aug 5, 2025 by Dax Reed · across every show →
Everything said about AI coding benchmarks, oldest first
Aug 5, 2025 negative
Dax Reed says almost no AI coding benchmarks resemble real-world engineering tasks.
“Yeah, so we looked at all the different benchmarks we could find, and it's almost comical how none of them look like my day-to-day work. Most of these evals are like, You know, solve this maze. And I'm just like, I'm never solving a maze. Like it's never anyth…”
Feb 23, 2026 neutral
Watkins: OpenAI will probably not release proprietary AI research coding benchmarks
“Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary. So if we make evals for that, like, we're probably not gonna release them. And it's harder for people in the field to make evals that kind of measure, like, is this a realist…”