Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
Olivia Watkins · The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals · Feb 23, 2026 · at 7:26
Olivia Watkins of OpenAI's Frontier Evals team discusses audit findings analyzing why advanced models failed certain SWE-bench Verified tasks.
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wasn't specified in the problem description, so it wasn't fair to expect that model to make that particular design choice.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →