Assertion Partly supported AI assessment confidence: 85% certainty 3/5 debate potential 2/5

Google's SWE-bench submission utilized thousands of trajectories and selection strategies

Shawn Lewis · Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis · Jan 28, 2025 · at 31:33

Weights & Biases CTO Shawn Lewis discusses how competitor submissions on the SWE-bench Verified leaderboard employ multi-trajectory sampling.

0:00 / 0:05exact quote · 5.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The Google submission down below, I think ran thousands of trajectories and then has a strategy for choosing the best.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Shawn Lewis

Insight
OpenAI o1 struggles with multi-step agentic tasks compared to GPT-4
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here. I think it's like less trained To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pr…”
Shawn Lewis Jan 28, 2025 ▶ 16:09 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Insight
Automating AI trace inspection is impossible; manual review remains mandatory
“And really you cannot avoid this last part. Like I've tried a lot to automate parts of this by I've tried a lot to automate it to remove the need for me to actually like manually inspect all of these traces. And I'm here to tell you like today, that is still i…”
Shawn Lewis Jan 28, 2025 ▶ 21:58 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Insight
Topping SWE-bench requires multi-trajectory sampling and high compute costs
“You know, if you want to get to the top of the leaderboard at this point, you probably need a strategy like that somewhere and be willing to, like, pay the cost, whatever it is, but it's also, like, it's really important, like, to be able to move up six percen…”
Shawn Lewis Jan 28, 2025 ▶ 31:40 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Prediction Not checkable as stated
Autonomous AI programmers will work effectively within the next two years
“I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.”
Shawn Lewis Jan 28, 2025 ▶ 33:01 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Prediction Not checkable as stated
Model training from scratch will concentrate mostly in major AI labs
“As AI improves, like, fewer people probably need to train models from scratch. It gets concentrated more and more in, in different, like, in the big labs.”
Shawn Lewis Jan 28, 2025 ▶ 3:44 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Assertion Supported
Shawn Lewis: o1 agent achieves 57% single-pass, 64% with parallel rollouts
“It solves, like, something like 57% of problems with a single Rollout and then using parallel rollouts and selecting the best one. With other techniques, we get something like 64%.”
Shawn Lewis Jan 28, 2025 ▶ 15:04 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.