SWE-bench Verified, every mention
23 scenes · ← back to SWE-bench Verified
tap a year for its mentions
every year anyone Shawn Wang 8Alessio Fanelli 6Mia Glaese 4Olivia Watkins 3Jesse Hu 3John Yang 2Graham Neubig 1
Verbatim, from the transcripts: the passages where SWE-bench Verified comes up
When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
- ▶ 7:34 unnamed speaker Like, uh, C bench verified, um, even vending bench one saturated, right?
The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
- ▶ 0:37 unnamed speaker And, uh, as by my understanding, you were part of the original team that worked on C-Bench Verified as well. 4 times in the scene
- ▶ 1:52 unnamed speaker I think the, uh, let's, let's sort of reset on, like, what was the original work that you guys did for Sequence Verified, which I think was pretty substantial. 4 times in the scene
- ▶ 8:57 Mia Glaese I think, I think also like at the time when three bench verified was published, I think it was like a very strong benchmark. 2 times in the scene
- ▶ 10:46 unnamed speaker We're going to stop reporting CBench Verified, right? 5 times in the scene
- ▶ 14:13 Mia Glaese They think three bench, three bench verified, obviously measured like some, that measures like some important capability, which is like, given like a description of a GitHub issue, can you produce like a patch that solves that issue, you… 2 times in the scene
- ▶ 23:35 Olivia Watkins And so, uh, we initially created Sweebench Verified as part of our, like, building out evals for that model autonomy workstream. 3 times in the scene
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
- ▶ 1:07 unnamed speaker And then SweetBench Verified was, like, maybe last year.
- ▶ 7:43 John Yang I think the projections are, are quite interesting, and I definitely appreciate them kind of using SweetBench Verified to, to sort of proxy a lot of these things, but
- ▶ 10:40 John Yang I don't know, but they basically took Sweet Bench Verified, and they changed the issues to make them impossible.
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
- ▶ 8:42 unnamed speaker So Subay's verified, for sure.
Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
- ▶ 16:08 unnamed speaker And then for example, you have sweet bench verified, but you don't have a sweet Lancer. 3 times in the scene
The #1 SWE-Bench Verified Agent
- ▶ 1:14 Alessio Fanelli And well, maybe you don't want to say to be humble, but this is going to be the number one sweep bench verified. 2 times in the scene
Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
- ▶ 26:03 unnamed speaker It's like, you know, there's kind of this, we bench verified benchmark, and then there's like, you know, the more programmatic, how am I going to use this agent? 2 times in the scene
2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
- ▶ 1:07:40 Shawn Wang We care about Sweebench verified.
Best of 2024 in Agents (from #1 on SWE-Bench Full, Prof. Graham Neubig of OpenHands/AllHands)
- ▶ 26:33 Graham Neubig Um, right now we have 53% or 55% on sweet bench verified, which is real world GitHub PRS.
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 4:14 Shawn Wang So yeah, maybe just give us a context about like why you looked at SweetBench Verified 3 times in the scene
- ▶ 9:14 Alessio Fanelli How do we get Sweepbench verified to 92%? 4 times in the scene
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Building AGI in Real Time (OpenAI Dev Day 2024)
- ▶ 1:07:00 unnamed speaker Special shout-out to listeners like Jesse from Morph Labs when he came on to talk about how he created synthetic datasets to fine-tune the largest lauras that had ever been created for GPT-for-O to post the highest-ever scores on Sweebench… 2 times in the scene
Is finetuning GPT4o worth it?
- ▶ 48:42 Shawn Wang I don't know if you want to comment on, on like that stuff versus, uh, you know, we also have like a, we also want to talk about Sweebench verified.
- ▶ 51:07 Shawn Wang Uh, Sweebench verified. 3 times in the scene