SweetBench
product on 2 shows · 4 statements across 3 episodes
4 statements about SweetBench, every show
Anthropic models hit 72% on SWE-bench, crushing earlier timeline predictions
“He's like, I think we'll be at 90% by the end of 20, 25 or something like that. And sure enough, we're at about 72 now with the new models and. We're at 50% when you made that prediction and it's like continued to scale pretty much like as predicted.”
SWE-bench is the best evaluation for AI programming today
“SweetBench is, you know, the best eval that we have for AI programming today.”
Anthropic releases exact tools and prompt used for SWE-bench agent
“With this blog post we released on SweetBench, we released the exact tools and the prompt that we gave the model to be able to do well.”
Schluntz: Traditional coding evals remain useful alongside SWE-bench
“I think there's definitely a space for these more traditional coding evals that are sort of easy to implement, quick to run and do get you some signal. And maybe hopefully there's just sort of harder versions of human eval that get created.”