Sweebench
product on 3 shows · 7 statements across 6 episodes
Latent Space
No Priors
Big Technology
7 statements about Sweebench, every show
Feinberg: Gemini is obviously worse at coding despite benchmark wins
“So Gemini does pretty well on Sweebench. Sometimes Gemini publishes models that win on some of those software benchmarks. Raise your hand if you're using Gemini to write code right now instead of, you know, the obvious other name competitors. No one. Like, why…”
Lloyd: Upgrading Anthropic models gave Warp only modest SWE-bench gains
“Like, if you take, like, Sonnet four to four five, and we're big partners with Anthropic, they have great models, like, that was, like, a few percentage point increase on Sweebench for us. And we invest, you know, we've invested a decent amount to be one of th…”
Lightcap: GPT-5 beats previous models on SWE-bench and health benchmarks
“It scores better on things like Sweebench. It scores better on all the kind of academic evals that we put it through. This one in particular, we actually made a real emphasis to have it score better on certain health benchmarks. So It's better at medical reaso…”
SWE-bench tasks do not reflect actual enterprise software engineering use cases
“That, and also like just in the enterprise, the use cases are pretty different than those represented in something like Sweebench.”
Running a 100-problem SWE-bench evaluation takes one to two hours
“So a Sweebench eval for me takes about an hour to two hours to run on like a subset of a hundred problems.”