Olivia Watkins

Member of Technical Staff, OpenAI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

scientistacademicengineeraliengirlliv.github.io/oliviawatkins ↗

Olivia Watkins is a researcher on the Frontier Evals team at OpenAI, where she evaluates frontier model risks and capabilities. She completed her Ph.D. in Computer Science at UC Berkeley's Center for Human-Compatible Artificial Intelligence under advisor Pieter Abbeel.

6statements → 4claims → 3claims resolved → 100%fully supported → 3.67/5average certainty → 2.33/5average debate potential →

3 supported 0 partly supported 0 contradicted 1 not checkable as stated how the 4 claims stand · each chip opens the sources

4 assertions · 1 opinion · 1 disclosure · every statement was checked. The predictions and assertions are the 4 claims: statements the public record can support or contradict. 3 are resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Olivia argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Olivia Watkins said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Opinion
Watkins: SWE-bench Verified is saturated, contaminated, and should be retired
“SweetBenchVerified has been one of the Northstar coding benchmarks that the field has looked at to measure coding progress. But recently we've seen that progress has kind of stalled and this, we realized that this is because the eval is effectively saturated a…”
Olivia Watkins Feb 23, 2026 ▶ 1:08 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Watkins: Over half of SWE-bench problems investigated by OpenAI had test flaws
“In over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are, like, overly narrow tests where there's some particular implementation detail that the tests were looking for but wa…”
Olivia Watkins Feb 23, 2026 ▶ 7:26 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Not checkable as stated
Watkins: SWE-bench Verified is contaminated across OpenAI, Claude, and Gemini models
“And in SweetBenchVerified, we found many instances of contamination across like, across OpenEye models, across, like, Quad Opus, 4.5, Gemini Flash, and all of these, we saw things like regurgitating the ground truth solutions, things like in some cases giving,…”
Olivia Watkins Feb 23, 2026 ▶ 11:54 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia Watkins Feb 23, 2026 ▶ 10:59 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Disclosure
Watkins: OpenAI will probably not release proprietary AI research coding benchmarks
“Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary. So if we make evals for that, like, we're probably not gonna release them. And it's harder for people in the field to make evals that kind of measure, like, is this a realist…”
Olivia Watkins Feb 23, 2026 ▶ 20:04 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Watkins: OpenAI hired nearly 100 engineers to curate 500 SWE-bench tasks
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Olivia Watkins Feb 23, 2026 ▶ 2:56 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals

Appearances (1)

EpisodeDateSpeaking time
The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals Feb 23, 2026 8m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.