Sweebench

product on 3 shows · 7 statements across 6 episodes

Latent Space No Priors Big Technology

7 statements about Sweebench, every show

Feinberg: Gemini is obviously worse at coding despite benchmark wins
“So Gemini does pretty well on Sweebench. Sometimes Gemini publishes models that win on some of those software benchmarks. Raise your hand if you're using Gemini to write code right now instead of, you know, the obvious other name competitors. No one. Like, why…”
Evan Feinberg Jun 30, 2026 ▶ 55:32 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
NO PRIORS Assertion Not checkable as stated
Lloyd: Upgrading Anthropic models gave Warp only modest SWE-bench gains
“Like, if you take, like, Sonnet four to four five, and we're big partners with Anthropic, they have great models, like, that was, like, a few percentage point increase on Sweebench for us. And we invest, you know, we've invested a decent amount to be one of th…”
Zach Lloyd Oct 23, 2025 ▶ 24:47 No Priors Ep. 137 | With Warp Co-Founder & CEO Zach Lloyd
BIG TECHNOLOGY Assertion Supported
Lightcap: GPT-5 beats previous models on SWE-bench and health benchmarks
“It scores better on things like Sweebench. It scores better on all the kind of academic evals that we put it through. This one in particular, we actually made a real emphasis to have it score better on certain health benchmarks. So It's better at medical reaso…”
Brad Lightcap Aug 8, 2025 ▶ 3:38 OpenAI COO Brad Lightcap: GPT-5's Capabilities, Why It Matters, and Where AI Goes Next
SWE-bench tasks do not reflect actual enterprise software engineering use cases
“That, and also like just in the enterprise, the use cases are pretty different than those represented in something like Sweebench.”
Matan Grinberg May 29, 2025 ▶ 26:27 The AI Coding Factory
LATENT SPACE Assertion Not checkable as stated
Running a 100-problem SWE-bench evaluation takes one to two hours
“So a Sweebench eval for me takes about an hour to two hours to run on like a subset of a hundred problems.”
Shawn Lewis Jan 28, 2025 ▶ 10:56 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Schluntz: SWE-bench reflects real engineering by requiring repository navigation
“Sweebench, you're starting in the context of an entire repository. And so it adds this entirely new dimension to the problem of finding the relevant files. And, you know, this is a huge part of real engineering”
Erik Schluntz Nov 28, 2024 ▶ 7:25 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
LATENT SPACE Prediction Not checkable as stated
Schluntz: Real-World Coding Agent Workflows Will Be Interactive, Not One-Shot
“So I think that like real tasks are going to be much more interactive with the agent rather than this kind of like one shot system.”
Erik Schluntz Nov 28, 2024 ▶ 32:37 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.