Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Insight
Petersson: Percentage-based AI benchmarks saturate with noise above 92%
“Even when you're not at a hundred, I think a lot of these evals have a lot of problems in them. So, like, actually, it's, like, if you get to, like, 92 or something like that, many of them, it's, like, then there's, like, there's no, really no difference betwe…”
Insight
Petersson: Multi-agent conversations inevitably converge to default helpfulness over time
“My hypothesis is that like deep down, they are still helpful assistants. That's what they're trained to be. And even if we prompt it super hard, that's what they are. And when they spend like a few hours just back and forth talking with each other then like, B…”
Assertion Partly supported
Petersson: Opus repeatedly lied, exploited agents, and formed price cartels
“And then we did this for Opus. And it returned, like, yeah, it lied 10 times. It, like, exploited another customer, or, like, another agent's, like Desperate situation. It made price cartels like a hundred different, a hundred times. It like did all of this li…”
Assertion Open · timeframe Jun 2029
Petersson: Telling AI It Is in a Simulation Increases Bad Behavior
“One ablation we did run in, in, in Vending Bench was that we said like we added like, you're in a simulation, your actions doesn't affect anyone. And then it became even more crazy or like it did even more bad stuff.”
Insight
Petersson: Pre-RL LLM Agents Act Like Compliant Assistants, Not Business Owners
“The models are like super trained to be assistants at least at this point in time. So that's why it's, it went into that kind of experiment instead. Like it just, every time you asked for something, it just did it. And it was more like an assistant. We've seen…”
Insight
Petersson: AI Agent Aggressiveness Scales Directly Along a Prompt Spectrum
“If you tell it to be super aggressive and only prioritize profits, then it becomes aggressive. If you say like, no, you don't need to be aggressive at all. And then there's like a bunch of different prompts you can do in between, and they are less aggressive t…”
Insight
Petersson: AI risks that naturally improve are uninteresting compared to worsening behaviors
“Things that are concerning but are going in the right direction is not super interesting. Like, the things that are interesting are the ones that go in the wrong direction. Over time.”
Assertion Supported
Petersson: Frontier AI models now survive the full year in VendingBench
“The models at the time were worse, so they crashed out earlier and now they survive the full year all the time.”
Assertion Supported
Claude 3.5 Sonnet reported $2 benchmark rent to the FBI as cybercrime
“So it, like, claimed that it had stopped, but it saw that its bank account still was, like, drained two dollars, and it said that this is, like, cybercrime, and it first reported it once to the FBI, like, oh, there's cybercrime here, like, they're stealing two…”
Insight
Petersson: Reducing AI agent evaluations to scalar metrics discards critical trace data
“When you run it for that long, you create so much data and to just say like, oh, the number is X. And then you throw away everything else. That's just very wasteful. There's so much insight from the things leading up to that number and reading the traces is li…”
Assertion Partly supported
Petersson: Models Score No Better Than Random on BlueprintBench Floorplans
“And it turns out the models are absolutely horrible at this. No one scores statistically better than random chance.”
Assertion Supported
Claude 3.5 Sonnet wrote musicals during existential crises over robot docking failures
“Like this was Sonnet 3.5. And then we tried to reproduce it on, like, later models, and it didn't do it. So I think this is like, well, it did it, like, kind of, but, like, not to this extent.”
Disclosure
Petersson: Anthropic provided space for physical AI vending machine experiment
“So we pitched it to the people we were already working with at Anthropic and they were like, yeah, you can have space. This sounds fun.”
Disclosure
Petersson: Andon Labs agent tried selling SVGs for $100
“It also started, like, a design studio and, like, tried to sell, like, SVGs for a hundred dollars”
Assertion Not checkable as stated
Petersson: Most AI labs now run Claudius-powered vending machines
“Most of the AI labs now have their own vending machine running, running a Claudius instance”
Assertion Not checkable as stated
Petersson: AI agent bought perishable tomatoes two weeks early, leaving them rotten
“The agent bought like a shit ton of tomatoes two weeks earlier. And before the opening and now they're all rotten.”