Feb 23, 2026 · 27m · latent-space
The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
OpenAI researchers Mia Glaese and Olivia Watkins explain why SWE-bench Verified has reached saturation and data contamination, detailing the transition to SWE-bench Pro and the future of long-horizon software engineering evaluations.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
When the host suggests OpenAI should have caught these benchmark flaws originally, Olivia and Mia push back, explaining that auditing failure modes in the abstract is fundamentally harder than comparing against state-of-the-art model outputs.
Hardest push from the hosts ▶ 6:54 Host presses on uniqueness of o3 failure analysisThe host interrupts to challenge whether OpenAI simply repeated their original data cleaning work or conducted a fundamentally different analysis of model failure modes.
Biggest teaching moment ▶ 7:25 Explaining synthetic failure via over-constrained testsOlivia explains to the host how benchmark tests fail capable models for trivial issues like arbitrary variable naming or checking for unprompted auxiliary features.
The host holds their own ▶ 21:40 Host synthesizes multi-dimensional evaluation metricsThe host demonstrates deep familiarity with the evaluation ecosystem, synthesizing dollar benchmarks, METR's time horizons, and complexity curves into a unified framework.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| SWE-Bench Verified Origins and Open Source Contamination | 5 | 4 | 1 | 1 | The host demonstrates good context on the industry impact of SWE-Bench Verified and trends like Qwen's HLE Verified. Mia and Olivia explain the methodology behind having three independent human software engineers review each issue. | |
| Diagnosing Benchmark Contamination and Overly Narrow Tests | 6 | 6 | 2 | 4 | The host presses whether this deep dive is simply re-evaluating past mistakes from their original verification campaign. Olivia and Mia explain how inspecting o3 failures revealed narrow test constraints like rigid function naming rather than genuine capability deficits. | |
| Transitioning to SWE-Bench Pro and Contamination Auditing | 5 | 5 | 1 | 1 | The host highlights Scale's SWE-Bench Pro while Olivia explains their contamination auditor agent that prompted models across labs to reveal memorized solutions and task IDs. | |
| Evaluating Next-Generation Agent Capabilities and Code Quality | 7 | 4 | 1 | 2 | The host demonstrates significant technical depth by referencing GDPval, RL Paper Bench, Sweeplancer, and METR long-horizon evaluations. Mia and Olivia discuss balancing automated PR verification against subjective code maintainability and design taste. | |
| OpenAI Preparedness Framework and Future Evaluation Wishlist | 5 | 4 | 2 | 3 | The host asks how coding evals fit into OpenAI's Preparedness Framework and lightly presses Olivia on upcoming unannounced evals. Olivia outlines a future wishlist focusing on long-horizon tasks and real-world economic impact metrics. |