Feb 23, 2026 · 27m · latent-space

The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals

Olivia Watkins · 8m spoken Mia Glaese · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

OpenAI researchers Mia Glaese and Olivia Watkins explain why SWE-bench Verified has reached saturation and data contamination, detailing the transition to SWE-bench Pro and the future of long-horizon software engineering evaluations.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.6 Guest teaching 4.6 Guest disagreement 1.4 The hosts pushing back 2.2
05100:0010:0020:001:51–5:28 · The hosts as informed peer 5/10 SWE-Bench Verified Origins and Open Source Contamination The host demonstrates good context on the industry impact of SWE-Bench Verified and trends like Qwen's HLE Verified. Mia and Olivia explain the methodology behind having three independent human software engineers review each issue.5:29–10:45 · The hosts as informed peer 6/10 Diagnosing Benchmark Contamination and Overly Narrow Tests The host presses whether this deep dive is simply re-evaluating past mistakes from their original verification campaign. Olivia and Mia explain how inspecting o3 failures revealed narrow test constraints like rigid function naming rather than genuine capability deficits.10:45–12:58 · The hosts as informed peer 5/10 Transitioning to SWE-Bench Pro and Contamination Auditing The host highlights Scale's SWE-Bench Pro while Olivia explains their contamination auditor agent that prompted models across labs to reveal memorized solutions and task IDs.12:58–22:31 · The hosts as informed peer 7/10 Evaluating Next-Generation Agent Capabilities and Code Quality The host demonstrates significant technical depth by referencing GDPval, RL Paper Bench, Sweeplancer, and METR long-horizon evaluations. Mia and Olivia discuss balancing automated PR verification against subjective code maintainability and design taste.22:31–26:56 · The hosts as informed peer 5/10 OpenAI Preparedness Framework and Future Evaluation Wishlist The host asks how coding evals fit into OpenAI's Preparedness Framework and lightly presses Olivia on upcoming unannounced evals. Olivia outlines a future wishlist focusing on long-horizon tasks and real-world economic impact metrics.1:51–5:28 · Guest teaching 4/10 SWE-Bench Verified Origins and Open Source Contamination The host demonstrates good context on the industry impact of SWE-Bench Verified and trends like Qwen's HLE Verified. Mia and Olivia explain the methodology behind having three independent human software engineers review each issue.5:29–10:45 · Guest teaching 6/10 Diagnosing Benchmark Contamination and Overly Narrow Tests The host presses whether this deep dive is simply re-evaluating past mistakes from their original verification campaign. Olivia and Mia explain how inspecting o3 failures revealed narrow test constraints like rigid function naming rather than genuine capability deficits.10:45–12:58 · Guest teaching 5/10 Transitioning to SWE-Bench Pro and Contamination Auditing The host highlights Scale's SWE-Bench Pro while Olivia explains their contamination auditor agent that prompted models across labs to reveal memorized solutions and task IDs.12:58–22:31 · Guest teaching 4/10 Evaluating Next-Generation Agent Capabilities and Code Quality The host demonstrates significant technical depth by referencing GDPval, RL Paper Bench, Sweeplancer, and METR long-horizon evaluations. Mia and Olivia discuss balancing automated PR verification against subjective code maintainability and design taste.22:31–26:56 · Guest teaching 4/10 OpenAI Preparedness Framework and Future Evaluation Wishlist The host asks how coding evals fit into OpenAI's Preparedness Framework and lightly presses Olivia on upcoming unannounced evals. Olivia outlines a future wishlist focusing on long-horizon tasks and real-world economic impact metrics.1:51–5:28 · Guest disagreement 1/10 SWE-Bench Verified Origins and Open Source Contamination The host demonstrates good context on the industry impact of SWE-Bench Verified and trends like Qwen's HLE Verified. Mia and Olivia explain the methodology behind having three independent human software engineers review each issue.5:29–10:45 · Guest disagreement 2/10 Diagnosing Benchmark Contamination and Overly Narrow Tests The host presses whether this deep dive is simply re-evaluating past mistakes from their original verification campaign. Olivia and Mia explain how inspecting o3 failures revealed narrow test constraints like rigid function naming rather than genuine capability deficits.10:45–12:58 · Guest disagreement 1/10 Transitioning to SWE-Bench Pro and Contamination Auditing The host highlights Scale's SWE-Bench Pro while Olivia explains their contamination auditor agent that prompted models across labs to reveal memorized solutions and task IDs.12:58–22:31 · Guest disagreement 1/10 Evaluating Next-Generation Agent Capabilities and Code Quality The host demonstrates significant technical depth by referencing GDPval, RL Paper Bench, Sweeplancer, and METR long-horizon evaluations. Mia and Olivia discuss balancing automated PR verification against subjective code maintainability and design taste.22:31–26:56 · Guest disagreement 2/10 OpenAI Preparedness Framework and Future Evaluation Wishlist The host asks how coding evals fit into OpenAI's Preparedness Framework and lightly presses Olivia on upcoming unannounced evals. Olivia outlines a future wishlist focusing on long-horizon tasks and real-world economic impact metrics.1:51–5:28 · The hosts pushing back 1/10 SWE-Bench Verified Origins and Open Source Contamination The host demonstrates good context on the industry impact of SWE-Bench Verified and trends like Qwen's HLE Verified. Mia and Olivia explain the methodology behind having three independent human software engineers review each issue.5:29–10:45 · The hosts pushing back 4/10 Diagnosing Benchmark Contamination and Overly Narrow Tests The host presses whether this deep dive is simply re-evaluating past mistakes from their original verification campaign. Olivia and Mia explain how inspecting o3 failures revealed narrow test constraints like rigid function naming rather than genuine capability deficits.10:45–12:58 · The hosts pushing back 1/10 Transitioning to SWE-Bench Pro and Contamination Auditing The host highlights Scale's SWE-Bench Pro while Olivia explains their contamination auditor agent that prompted models across labs to reveal memorized solutions and task IDs.12:58–22:31 · The hosts pushing back 2/10 Evaluating Next-Generation Agent Capabilities and Code Quality The host demonstrates significant technical depth by referencing GDPval, RL Paper Bench, Sweeplancer, and METR long-horizon evaluations. Mia and Olivia discuss balancing automated PR verification against subjective code maintainability and design taste.22:31–26:56 · The hosts pushing back 3/10 OpenAI Preparedness Framework and Future Evaluation Wishlist The host asks how coding evals fit into OpenAI's Preparedness Framework and lightly presses Olivia on upcoming unannounced evals. Olivia outlines a future wishlist focusing on long-horizon tasks and real-world economic impact metrics.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 0%27:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 8:45 Defending benchmark lifecycle against host critique

When the host suggests OpenAI should have caught these benchmark flaws originally, Olivia and Mia push back, explaining that auditing failure modes in the abstract is fundamentally harder than comparing against state-of-the-art model outputs.

Hardest push from the hosts ▶ 6:54 Host presses on uniqueness of o3 failure analysis

The host interrupts to challenge whether OpenAI simply repeated their original data cleaning work or conducted a fundamentally different analysis of model failure modes.

Biggest teaching moment ▶ 7:25 Explaining synthetic failure via over-constrained tests

Olivia explains to the host how benchmark tests fail capable models for trivial issues like arbitrary variable naming or checking for unprompted auxiliary features.

The host holds their own ▶ 21:40 Host synthesizes multi-dimensional evaluation metrics

The host demonstrates deep familiarity with the evaluation ecosystem, synthesizing dollar benchmarks, METR's time horizons, and complexity curves into a unified framework.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
SWE-Bench Verified Origins and Open Source Contamination 5411 The host demonstrates good context on the industry impact of SWE-Bench Verified and trends like Qwen's HLE Verified. Mia and Olivia explain the methodology behind having three independent human software engineers review each issue.
Diagnosing Benchmark Contamination and Overly Narrow Tests 6624 The host presses whether this deep dive is simply re-evaluating past mistakes from their original verification campaign. Olivia and Mia explain how inspecting o3 failures revealed narrow test constraints like rigid function naming rather than genuine capability deficits.
Transitioning to SWE-Bench Pro and Contamination Auditing 5511 The host highlights Scale's SWE-Bench Pro while Olivia explains their contamination auditor agent that prompted models across labs to reveal memorized solutions and task IDs.
Evaluating Next-Generation Agent Capabilities and Code Quality 7412 The host demonstrates significant technical depth by referencing GDPval, RL Paper Bench, Sweeplancer, and METR long-horizon evaluations. Mia and Olivia discuss balancing automated PR verification against subjective code maintainability and design taste.
OpenAI Preparedness Framework and Future Evaluation Wishlist 5423 The host asks how coding evals fit into OpenAI's Preparedness Framework and lightly presses Olivia on upcoming unannounced evals. Olivia outlines a future wishlist focusing on long-horizon tasks and real-world economic impact metrics.

Statements from this episode (8)

Opinion
Watkins: SWE-bench Verified is saturated, contaminated, and should be retired
“SweetBenchVerified has been one of the Northstar coding benchmarks that the field has looked at to measure coding progress. But recently we've seen that progress has kind of stalled and this, we realized that this is because the eval is effectively saturated a…”
Olivia Watkins Feb 23, 2026 ▶ 1:08
Assertion Supported
Watkins: OpenAI hired nearly 100 engineers to curate 500 SWE-bench tasks
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Olivia Watkins Feb 23, 2026 ▶ 2:56
Insight
Glaese: Open-source benchmarks cannot use canary strings to avoid contamination
“There's like multiple avenues, but like the problems are sourced from open source repos. So it's not just like when we usually publish evaluations, we publish evaluations, and then we add canary strings to ensure that, you know, they are easily filtered out at…”
Mia Glaese Feb 23, 2026 ▶ 4:52
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia Watkins Feb 23, 2026 ▶ 10:59
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Olivia Watkins Feb 23, 2026 ▶ 11:54
Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Mia Glaese Feb 23, 2026 ▶ 14:34
Disclosure
Watkins: OpenAI will probably not release proprietary AI research coding benchmarks
“Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary. So if we make evals for that, like, we're probably not gonna release them. And it's harder for people in the field to make evals that kind of measure, like, is this a realist…”
Olivia Watkins Feb 23, 2026 ▶ 20:04
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.