Mia Glaese (VP of Research at OpenAI) explains why SWE-bench Verified has reached saturation and is no longer an effective benchmark for frontier coding capabilities.
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Mia Glaese
Insight
Glaese: Open-source benchmarks cannot use canary strings to avoid contamination
“There's like multiple avenues, but like the problems are sourced from open source repos. So it's not just like when we usually publish evaluations, we publish evaluations, and then we add canary strings to ensure that, you know, they are easily filtered out at…”
Mia GlaeseFeb 23, 2026▶ 4:52The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.