Mia Glaese

VP of Research, OpenAI · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

executivescientistopenai.com ↗

Mia Glaese oversees alignment, human data, capability evaluations, and safety teams at OpenAI, guiding frontier model safeguards and evaluations. Prior to joining OpenAI, she was a research scientist at Google DeepMind, where she co-created the dialogue model Sparrow.

2statements → 1claims → 0claims resolved → 4/5average certainty → 2.5/5average debate potential →

1 not checkable as stated how the 1 claim stands · each chip opens the sources

1 assertion · 1 insight · every statement was checked. The predictions and assertion are the 1 claim: statements the public record can support or contradict. 0 are resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

Everything Mia Glaese said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Glaese: OpenAI no longer trusts further score improvements on SWE-bench Verified
“Issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it, but like it does measure something that is like a real like capability of models.”
Mia Glaese Feb 23, 2026 ▶ 14:34 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Insight
Glaese: Open-source benchmarks cannot use canary strings to avoid contamination
“There's like multiple avenues, but like the problems are sourced from open source repos. So it's not just like when we usually publish evaluations, we publish evaluations, and then we add canary strings to ensure that, you know, they are easily filtered out at…”
Mia Glaese Feb 23, 2026 ▶ 4:52 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals

Appearances (1)

EpisodeDateSpeaking time
The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals Feb 23, 2026 8m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.