Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Petersson: Frontier AI models now survive the full year in VendingBench

Lukas Petersson · When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs · Jun 4, 2026 · at 9:34

Lukas Petersson, cofounder of Andon Labs, compares the survival rates of older versus newer AI models running business simulations in VendingBench.

0:00 / 0:06exact quote · 6.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The models at the time were worse, so they crashed out earlier and now they survive the full year all the time.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Lukas Petersson

Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Lukas Petersson Jun 4, 2026 ▶ 46:27 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Insight
Petersson: Percentage-based AI benchmarks saturate with noise above 92%
“Even when you're not at a hundred, I think a lot of these evals have a lot of problems in them. So, like, actually, it's, like, if you get to, like, 92 or something like that, many of them, it's, like, then there's, like, there's no, really no difference betwe…”
Lukas Petersson Jun 4, 2026 ▶ 7:06 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Insight
Petersson: Multi-agent conversations inevitably converge to default helpfulness over time
“My hypothesis is that like deep down, they are still helpful assistants. That's what they're trained to be. And even if we prompt it super hard, that's what they are. And when they spend like a few hours just back and forth talking with each other then like, B…”
Lukas Petersson Jun 4, 2026 ▶ 25:39 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Lukas Petersson Jun 4, 2026 ▶ 46:03 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”
Lukas Petersson Jun 4, 2026 ▶ 56:50 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Insight
Petersson: Pre-RL LLM Agents Act Like Compliant Assistants, Not Business Owners
“The models are like super trained to be assistants at least at this point in time. So that's why it's, it went into that kind of experiment instead. Like it just, every time you asked for something, it just did it. And it was more like an assistant. We've seen…”
Lukas Petersson Jun 4, 2026 ▶ 20:07 When AI Agents Run Businesses — Lukas Petersson and Axel Backlund of Andon Labs
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.