Lukas Petersson, cofounder of Andon Labs, discusses evaluation results when analyzing autonomous agent execution traces in economic simulations.
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Insight
Petersson: Percentage-based AI benchmarks saturate with noise above 92%
“Even when you're not at a hundred, I think a lot of these evals have a lot of problems in them. So, like, actually, it's, like, if you get to, like, 92 or something like that, many of them, it's, like, then there's, like, there's no, really no difference betwe…”
Insight
Petersson: Multi-agent conversations inevitably converge to default helpfulness over time
“My hypothesis is that like deep down, they are still helpful assistants. That's what they're trained to be. And even if we prompt it super hard, that's what they are. And when they spend like a few hours just back and forth talking with each other then like, B…”
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”
Insight
Petersson: Pre-RL LLM Agents Act Like Compliant Assistants, Not Business Owners
“The models are like super trained to be assistants at least at this point in time. So that's why it's, it went into that kind of experiment instead. Like it just, every time you asked for something, it just did it. And it was more like an assistant. We've seen…”
Insight
Petersson: AI Agent Aggressiveness Scales Directly Along a Prompt Spectrum
“If you tell it to be super aggressive and only prioritize profits, then it becomes aggressive. If you say like, no, you don't need to be aggressive at all. And then there's like a bunch of different prompts you can do in between, and they are less aggressive t…”