Insight certainty 4/5 debate potential 2/5

Glaese: Open-source benchmarks cannot use canary strings to avoid contamination

Mia Glaese · The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals · Feb 23, 2026 · at 4:52

Mia Glaese, VP of Research at OpenAI, discusses how coding benchmarks derived from public GitHub repositories inherently suffer from training data contamination.

0:00 / 0:27exact quote · 27.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“There's like multiple avenues, but like the problems are sourced from open source repos. So it's not just like when we usually publish evaluations, we publish evaluations, and then we add canary strings to ensure that, you know, they are easily filtered out at training time. Obviously, if you use sort of like Data from, like, open market. You don't have, actually, like, a canary string.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Mia Glaese

Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Mia Glaese Feb 23, 2026 ▶ 14:34 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.