“Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary. So if we make evals for that, like, we're probably not gonna release them. And it's harder for people in the field to make evals that kind of measure, like, is this a realistic research coding?”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Olivia Watkins
Opinion
Watkins: SWE-bench Verified is saturated, contaminated, and should be retired
“SweetBenchVerified has been one of the Northstar coding benchmarks that the field has looked at to measure coding progress. But recently we've seen that progress has kind of stalled and this, we realized that this is because the eval is effectively saturated a…”
Olivia WatkinsFeb 23, 2026▶ 1:08The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
AssertionSupported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia WatkinsFeb 23, 2026▶ 7:26The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
AssertionNot checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Olivia WatkinsFeb 23, 2026▶ 11:54The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
AssertionSupported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia WatkinsFeb 23, 2026▶ 10:59The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Olivia WatkinsFeb 23, 2026▶ 2:56The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.