SWE-bench Verified, every mention

23 scenes · ← back to SWE-bench Verified

tap a year for its mentions
00133256202420252026episodesmentions
036202420252026episodes it came up in
007.53156202420252026episodesmentions per episode

every year anyone Shawn Wang 8Alessio Fanelli 6Mia Glaese 4Olivia Watkins 3Jesse Hu 3John Yang 2Graham Neubig 1

Verbatim, from the transcripts: the passages where SWE-bench Verified comes up

loading…

When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs Jun 4, 2026 · 1 mention

  • ▶ 7:34 unnamed speaker Like, uh, C bench verified, um, even vending bench one saturated, right?

The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals Feb 23, 2026 · 20 mentions

  • ▶ 0:37 unnamed speaker And, uh, as by my understanding, you were part of the original team that worked on C-Bench Verified as well. 4 times in the scene
  • ▶ 1:52 unnamed speaker I think the, uh, let's, let's sort of reset on, like, what was the original work that you guys did for Sequence Verified, which I think was pretty substantial. 4 times in the scene
  • ▶ 8:57 Mia Glaese I think, I think also like at the time when three bench verified was published, I think it was like a very strong benchmark. 2 times in the scene
  • ▶ 10:46 unnamed speaker We're going to stop reporting CBench Verified, right? 5 times in the scene
  • ▶ 14:13 Mia Glaese They think three bench, three bench verified, obviously measured like some, that measures like some important capability, which is like, given like a description of a GitHub issue, can you produce like a patch that solves that issue, you… 2 times in the scene
  • ▶ 23:35 Olivia Watkins And so, uh, we initially created Sweebench Verified as part of our, like, building out evals for that model autonomy workstream. 3 times in the scene

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang Dec 31, 2025 · 3 mentions

  • ▶ 1:07 unnamed speaker And then SweetBench Verified was, like, maybe last year.
  • ▶ 7:43 John Yang I think the projections are, are quite interesting, and I definitely appreciate them kind of using SweetBench Verified to, to sort of proxy a lot of these things, but
  • ▶ 10:40 John Yang I don't know, but they basically took Sweet Bench Verified, and they changed the issues to make them impossible.

[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor Dec 30, 2025 · 1 mention

  • ▶ 8:42 unnamed speaker So Subay's verified, for sure.

Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits Oct 18, 2025 · 3 mentions

  • ▶ 16:08 unnamed speaker And then for example, you have sweet bench verified, but you don't have a sweet Lancer. 3 times in the scene

The #1 SWE-Bench Verified Agent Apr 2, 2025 · 2 mentions

  • ▶ 1:14 Alessio Fanelli And well, maybe you don't want to say to be humble, but this is going to be the number one sweep bench verified. 2 times in the scene

Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis Jan 28, 2025 · 2 mentions

  • ▶ 26:03 unnamed speaker It's like, you know, there's kind of this, we bench verified benchmark, and then there's like, you know, the more programmatic, how am I going to use this agent? 2 times in the scene

2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents Jan 1, 2025 · 1 mention

Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands) Dec 25, 2024 · 1 mention

  • ▶ 26:33 Graham Neubig Um, right now we have 53% or 55% on sweet bench verified, which is real world GitHub PRS.

The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic Nov 28, 2024 · 7 mentions

  • ▶ 4:14 Shawn Wang So yeah, maybe just give us a context about like why you looked at SweetBench Verified 3 times in the scene
  • ▶ 9:14 Alessio Fanelli How do we get Sweepbench verified to 92%? 4 times in the scene

[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu Oct 19, 2024 · 6 mentions

  • ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 3 times in the scene
  • ▶ 17:06 Jesse Hu Well, so now I'll move on to verified. 3 times in the scene

Building AGI in Real Time (OpenAI Dev Day 2024) Oct 4, 2024 · 2 mentions

  • ▶ 1:07:00 unnamed speaker Special shout-out to listeners like Jesse from Morph Labs when he came on to talk about how he created synthetic datasets to fine-tune the largest lauras that had ever been created for GPT-for-O to post the highest-ever scores on Sweebench… 2 times in the scene

Is finetuning GPT4o worth it? Aug 22, 2024 · 4 mentions

  • ▶ 48:42 Shawn Wang I don't know if you want to comment on, on like that stuff versus, uh, you know, we also have like a, we also want to talk about Sweebench verified.
  • ▶ 51:07 Shawn Wang Uh, Sweebench verified. 3 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.