OpenAI o1 struggles with multi-step agentic tasks compared to GPT-4
Shawn Lewis · Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis · Jan 28, 2025 · at 16:09
Weights & Biases CTO Shawn Lewis compares using OpenAI's o1 model for agent orchestration versus GPT-4 when building coding agents.
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here. I think it's like less trained To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pretty decent at like taking a long sequence of steps to solve a problem and being able to refer back to like prior steps in the right order and stuff like that. Whereas it feels like it kind of gets confused.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →