Everything Alex Duffy said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Duffy: Playable AI game benchmarks teach people how LLMs operate
“If we make this playable, you know, then it kind of can teach people how to use AI, like language models just by playing. Cause you'll like understand how they work. You have to negotiate against them. You see their responses.”
Duffy: AI benchmarks follow a lifecycle from initial idea to saturation
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”
Duffy: Effective LLM writing requires editing earlier context to avoid pollution
“If you're not reflecting on a message coming out and editing above, then like that's something you need to do because you don't want to pollute the context with anything that's not exactly what you're trying to say.”
Gemini 2.5 Flash runs AI Diplomacy for $1-$5, undercutting competitors
“Running games at two five flash was instant and one to five bucks of, you know, way like, I don't know, 20 to a hundred with the other models.”
Duffy: AI Diplomacy trace logs and dataset are publicly released
“I posted a video on X that shows you how to actually like, we released all the data, all the trace logs.”
Duffy: LLM context should contain enough information for a human to play
“If you're looking at the context that's being sent to the language model, could you play the game?”
Duffy: Every wants to host a human vs. AI Diplomacy tournament
“I'd love to have a human versus AI diplomacy tournament.”
Duffy: Every is releasing Monologue, a local Whisper-style voice tool
“Naveen's actually releasing something called monologue. You know, it's kind of like, I think a better whisper flow for the things that we do can run locally.”
Duffy: AI Diplomacy used try-catch blocks instead of Pydantic for JSON
“I know one of our GitHub issues already called us out for not using Pydantic parsing and instead of having like eight different try catches for the different JSON versions that come out of these.”
Duffy: AI Diplomacy games take 4 to 18 hours to play out
“No, human time, watching it, I think it's, like, anywhere from four to 12, four to 18 hours, something, to play out.”