Weights & Biases CTO Shawn Lewis defends using multi-trajectory sampling and selection algorithms after taking the #1 ranking on SWE-bench Verified.
Insight
OpenAI o1 struggles with multi-step agentic tasks compared to GPT-4
“O-One is very good at programming, but it's kind of, the agent part was the harder part to get it to do here.
I think it's like less trained
To take the next step in like an agentic task, whereas GPT-IV for like the last two years has been really, you know, pr…”
Insight
Automating AI trace inspection is impossible; manual review remains mandatory
“And really you cannot avoid this last part. Like I've tried a lot to automate parts of this by I've tried a lot to automate it to remove the need for me to actually like manually inspect all of these traces. And I'm here to tell you like today, that is still i…”
Prediction Not checkable as stated
Autonomous AI programmers will work effectively within the next two years
“I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.”
Prediction Not checkable as stated
Model training from scratch will concentrate mostly in major AI labs
“As AI improves, like, fewer people probably need to train models from scratch. It gets concentrated more and more in, in different, like, in the big labs.”
Assertion Supported
Shawn Lewis: o1 agent achieves 57% single-pass, 64% with parallel rollouts
“It solves, like, something like 57% of problems with a single
Rollout and then using parallel rollouts and selecting the best one.
With other techniques, we get something like 64%.”
Opinion
AI agent moats lie in business interfaces, where Devin leads significantly
“I think that may be where most of the, like, if there's any mode here, it's gonna be around, like, the interfaces into humans and their businesses. And Devon, like, has a major, major lead on, on making that work really well.”