AI coding benchmarks

2 statements across 2 episodes · 0 bullish · 1 bearish · 2 people on the record · first statement Aug 5, 2025 by Dax Reed · across every show →

Everything said about AI coding benchmarks, oldest first

Aug 5, 2025 negative
Opinion
Dax Reed says almost no AI coding benchmarks resemble real-world engineering tasks.
“Yeah, so we looked at all the different benchmarks we could find, and it's almost comical how none of them look like my day-to-day work. Most of these evals are like, You know, solve this maze. And I'm just like, I'm never solving a maze. Like it's never anyth…”
Dax Reed Aug 5, 2025 ▶ 19:43 ⚡️OpenCode: Claude Code but Open Source, with Any Model, and frontier TUI - with Dax Reed (@thdxr)
Feb 23, 2026 neutral
Disclosure
Watkins: OpenAI will probably not release proprietary AI research coding benchmarks
“Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary. So if we make evals for that, like, we're probably not gonna release them. And it's harder for people in the field to make evals that kind of measure, like, is this a realist…”
Olivia Watkins Feb 23, 2026 ▶ 20:04 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.