Mike Krieger, CPO at Anthropic, explains the gap between synthetic coding benchmark scores (SWE-bench) and user-perceived model reliability during model development.
“Even when it was already outperforming Opus, for example, on sweet bench, people still didn't feel it was better, but then it continued to train and it was like now better than Opus and people don't want to switch back.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Mike Krieger
Opinion
Krieger: Knowledge work AI progress follows code's exponential curve on a delay
“And I think of knowledge work as being on a similar sort of exponential as code has been on, but just time shifted, right?”
Mike KriegerSep 30, 2025▶ 19:50⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
AssertionNot checkable as stated
Sonnet 4.5 ran autonomously for 30 hours versus 7 for Opus 4
“So this, you know, we had a customer and internally, we also got like a 30 hour plus kind of execution versus I think Opus four was seven hours.”
Mike KriegerSep 30, 2025▶ 20:51⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
Opinion
Krieger: Claude Sonnet 4.5 Outperforms Opus at Generating 3D Games
“This is like officially good. It's like better than Opus at this. It's like, It generated this, like, great split-screen stereoscopic thing, three-dimensional, like, thing.”
Mike KriegerSep 30, 2025▶ 4:24⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
Opinion
Krieger: Current vision models lack the precision of skilled visual designers
“The models don't see as well as they could. They see, okay, you know, you ask them analyze a complex photo and they're able to do it, but I want them to be as persnickety as a like really good visual designer. Like, no, that looks, the baseline looks a little …”
Mike KriegerSep 30, 2025▶ 10:07⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
Insight
Krieger: AI Agents Must Support Both MCP and Visual Computer Use
“And that thing's never gonna have an MCP around it. Like, it's just like, who knows if the company created is even around much less like ready to sort of expose their kind of underlying constructs as API. So I think you will need to be able to do both.”
Mike KriegerSep 30, 2025▶ 13:00⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
AssertionNot checkable as stated
Krieger: Claude Sonnet 4.5 day-one traffic eclipsed Sonnet 4
“We have more traffic on Sonnet 4.5 than we had on Sonnet four. So basically it's already eclipsed Sonnet four.”
Mike KriegerSep 30, 2025▶ 1:15⚡️Claude Sonnet 4.5 and Anthropic's roadmap for Agents and Developers — Mike Krieger, Anthropic
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.