Insight certainty 3/5 debate potential 3/5

OpenAI o1 struggles with multi-step agentic tasks compared to GPT-4

Shawn Lewis · Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis · Jan 28, 2025 · at 16:09

Weights & Biases CTO Shawn Lewis compares using OpenAI's o1 model for agent orchestration versus GPT-4 when building coding agents.

0:00 / 0:26exact quote · 26.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here. I think it's like less trained To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pretty decent at like taking a long sequence of steps to solve a problem and being able to refer back to like prior steps in the right order and stuff like that. Whereas it feels like it kind of gets confused.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Shawn Lewis

Insight
Automating AI trace inspection is impossible; manual review remains mandatory
“And really you cannot avoid this last part. Like I've tried a lot to automate parts of this by I've tried a lot to automate it to remove the need for me to actually like manually inspect all of these traces. And I'm here to tell you like today, that is still i…”
Shawn Lewis Jan 28, 2025 ▶ 21:58 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Insight
Topping SWE-bench requires multi-trajectory sampling and high compute costs
“You know, if you want to get to the top of the leaderboard at this point, you probably need a strategy like that somewhere and be willing to, like, pay the cost, whatever it is, but it's also, like, it's really important, like, to be able to move up six percen…”
Shawn Lewis Jan 28, 2025 ▶ 31:40 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Prediction Not checkable as stated
Autonomous AI programmers will work effectively within the next two years
“I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.”
Shawn Lewis Jan 28, 2025 ▶ 33:01 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Prediction Not checkable as stated
Model training from scratch will concentrate mostly in major AI labs
“As AI improves, like, fewer people probably need to train models from scratch. It gets concentrated more and more in, in different, like, in the big labs.”
Shawn Lewis Jan 28, 2025 ▶ 3:44 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Assertion Supported
Shawn Lewis: o1 agent achieves 57% single-pass, 64% with parallel rollouts
“It solves, like, something like 57% of problems with a single Rollout and then using parallel rollouts and selecting the best one. With other techniques, we get something like 64%.”
Shawn Lewis Jan 28, 2025 ▶ 15:04 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Opinion
AI agent moats lie in business interfaces, where Devin leads significantly
“I think that may be where most of the, like, if there's any mode here, it's gonna be around, like, the interfaces into humans and their businesses. And Devon, like, has a major, major lead on, on making that work really well.”
Shawn Lewis Jan 28, 2025 ▶ 33:22 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.