Jan 28, 2025 · 34m · latent-space

Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis

Shawn Lewis · 23m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Weights & Biases CTO Shawn Lewis joins the Latent Space Podcast to detail how he achieved a record 64.6% solve rate on the SWE-bench Verified leaderboard using OpenAI's o1 model, the Phase Shift framework, and deep observability tooling built on Weave.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.8 Guest teaching 4.4 Guest disagreement 1.1 The hosts pushing back 2.0
05100:0010:0020:0030:001:43–4:57 · The hosts as informed peer 3/10 The Philosophy of Dogfooding and AI Tool Evolution Swix introduces Shawn and sets up a conversational inquiry about why Weights & Biases ventured into SWE-bench agent tooling. Shawn provides an expansive backstory detailing his history with gaze tracking, experiment tracking, and the dogfooding philosophy behind Weave.4:57–7:06 · The hosts as informed peer 5/10 Transitioning from Model Evals to Coding Agent Frameworks Alessio references their past interview with Anthropic regarding SWE-agent evals, prompting Shawn to share his setup. Swix gently intervenes to redirect Shawn from getting too deep into the weeds before establishing the headline results.7:07–13:08 · The hosts as informed peer 3/10 Visualizing Agent Runs and Bottlenecks with Eval Studio Shawn screen-shares Eval Studio and walks through tracking metrics across SWE-bench runs, explaining kinks in reasoning token distributions. The hosts listen as Shawn illustrates the evolution from Google Sheets to integrated Weave evals.13:08–17:18 · The hosts as informed peer 6/10 Leveraging OpenAI o1 for End-to-End Agent Logic Alessio prompts a technical breakdown of Shawn's o1-driven agent architecture. Swix pushes back by citing Paul Gauthier's findings on Aider using o1 as architect and Claude Sonnet as coder, leading Shawn to counter why he found o1 agentic reasoning uniquely challenging and effective.17:19–22:23 · The hosts as informed peer 4/10 Data-Driven Debugging and Linear Agent Traces Shawn explains how aggregating public SWE-bench solve counts yielded per-instance difficulty scores to pinpoint prompt regressions. He demos the linear step-by-step trace viewer and argues that manual inspection of agent data remains essential.22:23–25:04 · The hosts as informed peer 6/10 Configuration Diffing and Framework Tracking in Weave Alessio asks an incisive diagnostic question regarding how to differentiate prompt regressions from tool description failures. Shawn walks through Phase Shift's configuration diffing and code tracking inside Weave to identify prompt phrasing deltas.25:04–27:28 · The hosts as informed peer 5/10 Phase Shift Roadmap and AI-Assisted Tool Development Alessio asks whether Phase Shift can be polished by its own agent and used for real-world tasks beyond benchmarks. Shawn explains that the UI was actually built with Cursor and notes the framework is currently tuned specifically for SWE-bench verified.27:29–32:30 · The hosts as informed peer 6/10 SWE-Bench Trajectories and the Cross-Check Selection Mechanism Swix brings up critique from competitor agent builders who dismiss Shawn's score as expensive rejection sampling. Shawn defends the validity and necessity of parallel rollouts and cross-check selection, noting every top leaderboard participant relies on trajectory filtering at this frontier.1:43–4:57 · Guest teaching 2/10 The Philosophy of Dogfooding and AI Tool Evolution Swix introduces Shawn and sets up a conversational inquiry about why Weights & Biases ventured into SWE-bench agent tooling. Shawn provides an expansive backstory detailing his history with gaze tracking, experiment tracking, and the dogfooding philosophy behind Weave.4:57–7:06 · Guest teaching 2/10 Transitioning from Model Evals to Coding Agent Frameworks Alessio references their past interview with Anthropic regarding SWE-agent evals, prompting Shawn to share his setup. Swix gently intervenes to redirect Shawn from getting too deep into the weeds before establishing the headline results.7:07–13:08 · Guest teaching 5/10 Visualizing Agent Runs and Bottlenecks with Eval Studio Shawn screen-shares Eval Studio and walks through tracking metrics across SWE-bench runs, explaining kinks in reasoning token distributions. The hosts listen as Shawn illustrates the evolution from Google Sheets to integrated Weave evals.13:08–17:18 · Guest teaching 5/10 Leveraging OpenAI o1 for End-to-End Agent Logic Alessio prompts a technical breakdown of Shawn's o1-driven agent architecture. Swix pushes back by citing Paul Gauthier's findings on Aider using o1 as architect and Claude Sonnet as coder, leading Shawn to counter why he found o1 agentic reasoning uniquely challenging and effective.17:19–22:23 · Guest teaching 6/10 Data-Driven Debugging and Linear Agent Traces Shawn explains how aggregating public SWE-bench solve counts yielded per-instance difficulty scores to pinpoint prompt regressions. He demos the linear step-by-step trace viewer and argues that manual inspection of agent data remains essential.22:23–25:04 · Guest teaching 5/10 Configuration Diffing and Framework Tracking in Weave Alessio asks an incisive diagnostic question regarding how to differentiate prompt regressions from tool description failures. Shawn walks through Phase Shift's configuration diffing and code tracking inside Weave to identify prompt phrasing deltas.25:04–27:28 · Guest teaching 4/10 Phase Shift Roadmap and AI-Assisted Tool Development Alessio asks whether Phase Shift can be polished by its own agent and used for real-world tasks beyond benchmarks. Shawn explains that the UI was actually built with Cursor and notes the framework is currently tuned specifically for SWE-bench verified.27:29–32:30 · Guest teaching 6/10 SWE-Bench Trajectories and the Cross-Check Selection Mechanism Swix brings up critique from competitor agent builders who dismiss Shawn's score as expensive rejection sampling. Shawn defends the validity and necessity of parallel rollouts and cross-check selection, noting every top leaderboard participant relies on trajectory filtering at this frontier.1:43–4:57 · Guest disagreement 0/10 The Philosophy of Dogfooding and AI Tool Evolution Swix introduces Shawn and sets up a conversational inquiry about why Weights & Biases ventured into SWE-bench agent tooling. Shawn provides an expansive backstory detailing his history with gaze tracking, experiment tracking, and the dogfooding philosophy behind Weave.4:57–7:06 · Guest disagreement 1/10 Transitioning from Model Evals to Coding Agent Frameworks Alessio references their past interview with Anthropic regarding SWE-agent evals, prompting Shawn to share his setup. Swix gently intervenes to redirect Shawn from getting too deep into the weeds before establishing the headline results.7:07–13:08 · Guest disagreement 0/10 Visualizing Agent Runs and Bottlenecks with Eval Studio Shawn screen-shares Eval Studio and walks through tracking metrics across SWE-bench runs, explaining kinks in reasoning token distributions. The hosts listen as Shawn illustrates the evolution from Google Sheets to integrated Weave evals.13:08–17:18 · Guest disagreement 4/10 Leveraging OpenAI o1 for End-to-End Agent Logic Alessio prompts a technical breakdown of Shawn's o1-driven agent architecture. Swix pushes back by citing Paul Gauthier's findings on Aider using o1 as architect and Claude Sonnet as coder, leading Shawn to counter why he found o1 agentic reasoning uniquely challenging and effective.17:19–22:23 · Guest disagreement 0/10 Data-Driven Debugging and Linear Agent Traces Shawn explains how aggregating public SWE-bench solve counts yielded per-instance difficulty scores to pinpoint prompt regressions. He demos the linear step-by-step trace viewer and argues that manual inspection of agent data remains essential.22:23–25:04 · Guest disagreement 0/10 Configuration Diffing and Framework Tracking in Weave Alessio asks an incisive diagnostic question regarding how to differentiate prompt regressions from tool description failures. Shawn walks through Phase Shift's configuration diffing and code tracking inside Weave to identify prompt phrasing deltas.25:04–27:28 · Guest disagreement 0/10 Phase Shift Roadmap and AI-Assisted Tool Development Alessio asks whether Phase Shift can be polished by its own agent and used for real-world tasks beyond benchmarks. Shawn explains that the UI was actually built with Cursor and notes the framework is currently tuned specifically for SWE-bench verified.27:29–32:30 · Guest disagreement 4/10 SWE-Bench Trajectories and the Cross-Check Selection Mechanism Swix brings up critique from competitor agent builders who dismiss Shawn's score as expensive rejection sampling. Shawn defends the validity and necessity of parallel rollouts and cross-check selection, noting every top leaderboard participant relies on trajectory filtering at this frontier.1:43–4:57 · The hosts pushing back 0/10 The Philosophy of Dogfooding and AI Tool Evolution Swix introduces Shawn and sets up a conversational inquiry about why Weights & Biases ventured into SWE-bench agent tooling. Shawn provides an expansive backstory detailing his history with gaze tracking, experiment tracking, and the dogfooding philosophy behind Weave.4:57–7:06 · The hosts pushing back 2/10 Transitioning from Model Evals to Coding Agent Frameworks Alessio references their past interview with Anthropic regarding SWE-agent evals, prompting Shawn to share his setup. Swix gently intervenes to redirect Shawn from getting too deep into the weeds before establishing the headline results.7:07–13:08 · The hosts pushing back 1/10 Visualizing Agent Runs and Bottlenecks with Eval Studio Shawn screen-shares Eval Studio and walks through tracking metrics across SWE-bench runs, explaining kinks in reasoning token distributions. The hosts listen as Shawn illustrates the evolution from Google Sheets to integrated Weave evals.13:08–17:18 · The hosts pushing back 5/10 Leveraging OpenAI o1 for End-to-End Agent Logic Alessio prompts a technical breakdown of Shawn's o1-driven agent architecture. Swix pushes back by citing Paul Gauthier's findings on Aider using o1 as architect and Claude Sonnet as coder, leading Shawn to counter why he found o1 agentic reasoning uniquely challenging and effective.17:19–22:23 · The hosts pushing back 1/10 Data-Driven Debugging and Linear Agent Traces Shawn explains how aggregating public SWE-bench solve counts yielded per-instance difficulty scores to pinpoint prompt regressions. He demos the linear step-by-step trace viewer and argues that manual inspection of agent data remains essential.22:23–25:04 · The hosts pushing back 2/10 Configuration Diffing and Framework Tracking in Weave Alessio asks an incisive diagnostic question regarding how to differentiate prompt regressions from tool description failures. Shawn walks through Phase Shift's configuration diffing and code tracking inside Weave to identify prompt phrasing deltas.25:04–27:28 · The hosts pushing back 1/10 Phase Shift Roadmap and AI-Assisted Tool Development Alessio asks whether Phase Shift can be polished by its own agent and used for real-world tasks beyond benchmarks. Shawn explains that the UI was actually built with Cursor and notes the framework is currently tuned specifically for SWE-bench verified.27:29–32:30 · The hosts pushing back 4/10 SWE-Bench Trajectories and the Cross-Check Selection Mechanism Swix brings up critique from competitor agent builders who dismiss Shawn's score as expensive rejection sampling. Shawn defends the validity and necessity of parallel rollouts and cross-check selection, noting every top leaderboard participant relies on trajectory filtering at this frontier.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 31:11 Rebutting criticism of multi-trajectory rollouts

Shawn firmly rejects the dismissive claim that his result is trivial brute forcing, explaining that every top lab uses parallel rollouts and that gaining 6% at the top of the benchmark is exponentially difficult.

Hardest push from the hosts ▶ 15:52 Swix challenges pure o1 architecture with Aider benchmark

Swix challenges Shawn's reliance on pure o1 by citing Paul Gauthier's proven Aider setup combining o1 as an architect with Sonnet as a code generator.

Biggest teaching moment ▶ 18:22 Quantifying instance difficulty from public leaderboard solves

Shawn teaches the hosts a clever evaluation technique: calculating an empirical difficulty score for benchmark instances by aggregating total public solves across leaderboard submissions.

The host holds their own ▶ 15:52 Host introduces competitor metagame dynamics

Swix demonstrates deep industry awareness of agent architectures by bringing up specific multi-model configurations on competing coding benchmarks to test Shawn's architectural choices.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Philosophy of Dogfooding and AI Tool Evolution 3200 Swix introduces Shawn and sets up a conversational inquiry about why Weights & Biases ventured into SWE-bench agent tooling. Shawn provides an expansive backstory detailing his history with gaze tracking, experiment tracking, and the dogfooding philosophy behind Weave.
Transitioning from Model Evals to Coding Agent Frameworks 5212 Alessio references their past interview with Anthropic regarding SWE-agent evals, prompting Shawn to share his setup. Swix gently intervenes to redirect Shawn from getting too deep into the weeds before establishing the headline results.
Visualizing Agent Runs and Bottlenecks with Eval Studio 3501 Shawn screen-shares Eval Studio and walks through tracking metrics across SWE-bench runs, explaining kinks in reasoning token distributions. The hosts listen as Shawn illustrates the evolution from Google Sheets to integrated Weave evals.
Leveraging OpenAI o1 for End-to-End Agent Logic 6545 Alessio prompts a technical breakdown of Shawn's o1-driven agent architecture. Swix pushes back by citing Paul Gauthier's findings on Aider using o1 as architect and Claude Sonnet as coder, leading Shawn to counter why he found o1 agentic reasoning uniquely challenging and effective.
Data-Driven Debugging and Linear Agent Traces 4601 Shawn explains how aggregating public SWE-bench solve counts yielded per-instance difficulty scores to pinpoint prompt regressions. He demos the linear step-by-step trace viewer and argues that manual inspection of agent data remains essential.
Configuration Diffing and Framework Tracking in Weave 6502 Alessio asks an incisive diagnostic question regarding how to differentiate prompt regressions from tool description failures. Shawn walks through Phase Shift's configuration diffing and code tracking inside Weave to identify prompt phrasing deltas.
Phase Shift Roadmap and AI-Assisted Tool Development 5401 Alessio asks whether Phase Shift can be polished by its own agent and used for real-world tasks beyond benchmarks. Shawn explains that the UI was actually built with Cursor and notes the framework is currently tuned specifically for SWE-bench verified.
SWE-Bench Trajectories and the Cross-Check Selection Mechanism 6644 Swix brings up critique from competitor agent builders who dismiss Shawn's score as expensive rejection sampling. Shawn defends the validity and necessity of parallel rollouts and cross-check selection, noting every top leaderboard participant relies on trajectory filtering at this frontier.

Statements from this episode (13)

Prediction Not checkable as stated
Model training from scratch will concentrate mostly in major AI labs
“As AI improves, like, fewer people probably need to train models from scratch. It gets concentrated more and more in, in different, like, in the big labs.”
Shawn Lewis Jan 28, 2025 ▶ 3:44
Opinion
SWE-bench is the best evaluation for AI programming today
“SweetBench is, you know, the best eval that we have for AI programming today.”
Shawn Lewis Jan 28, 2025 ▶ 7:46
Assertion Not checkable as stated
Running a 100-problem SWE-bench evaluation takes one to two hours
“So a Sweebench eval for me takes about an hour to two hours to run on like a subset of a hundred problems.”
Shawn Lewis Jan 28, 2025 ▶ 10:56
Assertion Supported
Shawn Lewis: o1 agent achieves 57% single-pass, 64% with parallel rollouts
“It solves, like, something like 57% of problems with a single Rollout and then using parallel rollouts and selecting the best one. With other techniques, we get something like 64%.”
Shawn Lewis Jan 28, 2025 ▶ 15:04
Insight
OpenAI o1 struggles with multi-step agentic tasks compared to GPT-4
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here. I think it's like less trained To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pr…”
Shawn Lewis Jan 28, 2025 ▶ 16:09
Disclosure
Lewis: Ran approximately 1,000 evaluations while developing SWE-bench agent
“You can see in the course of this, I did something like a thousand evals.”
Shawn Lewis Jan 28, 2025 ▶ 20:02
Insight
Automating AI trace inspection is impossible; manual review remains mandatory
“And really you cannot avoid this last part. Like I've tried a lot to automate parts of this by I've tried a lot to automate it to remove the need for me to actually like manually inspect all of these traces. And I'm here to tell you like today, that is still i…”
Shawn Lewis Jan 28, 2025 ▶ 21:58
Insight
Debug agent regressions by qualitatively clustering failures across execution traces
“So it's really, like, I'll flip through these traces and kind of, like for each one, I'll write down notes about, like, what I thought went wrong there, and I'll do that for, like, say, 20 or so, and then I kind of go, okay, what's the biggest problem that we …”
Shawn Lewis Jan 28, 2025 ▶ 22:50
Assertion Not checkable as stated
Phase Shift's Eval Studio user interface was entirely written using Cursor
“Everything in this in the phase shift UI, this, or this eval studio UI that I showed you was written by AI. So this entire UI was written by AI, but it was not written by. Oh my God. They shipped it was written by cursor.”
Shawn Lewis Jan 28, 2025 ▶ 26:22
Assertion Partly supported
Google's SWE-bench submission utilized thousands of trajectories and selection strategies
“The Google submission down below, I think ran thousands of trajectories and then has a strategy for choosing the best.”
Shawn Lewis Jan 28, 2025 ▶ 31:33
Insight
Topping SWE-bench requires multi-trajectory sampling and high compute costs
“You know, if you want to get to the top of the leaderboard at this point, you probably need a strategy like that somewhere and be willing to, like, pay the cost, whatever it is, but it's also, like, it's really important, like, to be able to move up six percen…”
Shawn Lewis Jan 28, 2025 ▶ 31:40
Prediction Not checkable as stated
Autonomous AI programmers will work effectively within the next two years
“I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.”
Shawn Lewis Jan 28, 2025 ▶ 33:01
Opinion
AI agent moats lie in business interfaces, where Devin leads significantly
“I think that may be where most of the, like, if there's any mode here, it's gonna be around, like, the interfaces into humans and their businesses. And Devon, like, has a major, major lead on, on making that work really well.”
Shawn Lewis Jan 28, 2025 ▶ 33:22
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.