Glaese: Open-source benchmarks cannot use canary strings to avoid contamination
Mia Glaese · The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals · Feb 23, 2026 · at 4:52
Mia Glaese, VP of Research at OpenAI, discusses how coding benchmarks derived from public GitHub repositories inherently suffer from training data contamination.
“There's like multiple avenues, but like the problems are sourced from open source repos. So it's not just like when we usually publish evaluations, we publish evaluations, and then we add canary strings to ensure that, you know, they are easily filtered out at training time. Obviously, if you use sort of like Data from, like, open market. You don't have, actually, like, a canary string.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →