Assertion certainty 4/5 debate potential 1/5

Running a 100-problem SWE-bench evaluation takes one to two hours

Shawn Lewis · Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis · Jan 28, 2025 · at 10:56

Shawn Lewis (CTO of Weights & Biases) explains the time and compute requirements for testing his AI coding agent setup on SWE-bench.

0:00 / 0:05exact quote · 5.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So a Sweebench eval for me takes about an hour to two hours to run on like a subset of a hundred problems.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Shawn Lewis

Insight
OpenAI o1 struggles with multi-step agentic tasks compared to GPT-4
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here. I think it's like less trained To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pr…”
Shawn Lewis Jan 28, 2025 ▶ 16:09 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Insight
Automating AI trace inspection is impossible; manual review remains mandatory
“And really you cannot avoid this last part. Like I've tried a lot to automate parts of this by I've tried a lot to automate it to remove the need for me to actually like manually inspect all of these traces. And I'm here to tell you like today, that is still i…”
Shawn Lewis Jan 28, 2025 ▶ 21:58 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Insight
Topping SWE-bench requires multi-trajectory sampling and high compute costs
“You know, if you want to get to the top of the leaderboard at this point, you probably need a strategy like that somewhere and be willing to, like, pay the cost, whatever it is, but it's also, like, it's really important, like, to be able to move up six percen…”
Shawn Lewis Jan 28, 2025 ▶ 31:40 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Prediction Not checkable as stated
Autonomous AI programmers will work effectively within the next two years
“I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.”
Shawn Lewis Jan 28, 2025 ▶ 33:01 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Prediction Not checkable as stated
Model training from scratch will concentrate mostly in major AI labs
“As AI improves, like, fewer people probably need to train models from scratch. It gets concentrated more and more in, in different, like, in the big labs.”
Shawn Lewis Jan 28, 2025 ▶ 3:44 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Assertion Supported
Shawn Lewis: o1 agent achieves 57% single-pass, 64% with parallel rollouts
“It solves, like, something like 57% of problems with a single Rollout and then using parallel rollouts and selecting the best one. With other techniques, we get something like 64%.”
Shawn Lewis Jan 28, 2025 ▶ 15:04 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.