Jun 11, 2025 · 34m · latent-space

⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy

Alex Duffy · 22m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Alex Duffy, Head of AI at Every, breaks down the creation and strategic insights of the AI Diplomacy game benchmark while discussing Every's multidisciplinary culture, practical AI tooling, and human-in-the-loop workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.3 Guest teaching 3.2 Guest disagreement 0.2 The hosts pushing back 0.7
05100:0010:0020:0030:001:48–5:00 · The hosts as informed peer 3/10 Every's Product Ecosystem: Cora, Sparkle, and Spiral The host asks about Every's suite of AI products and Alex's consulting practice. Alex provides an overview of internal tools like Cora, Sparkle, and Spiral in a very collaborative introductory exchange.5:00–9:01 · The hosts as informed peer 4/10 Origins and Design of the AI Diplomacy Benchmark The host brings up the Diplomacy benchmark stream and references Noam Brown and DiploBench. Alex shares the project's background and community contributions across the AI ecosystem.9:02–14:20 · The hosts as informed peer 5/10 Games, Self-Play, and the Purpose of Benchmarking The host provides historical context on game benchmarks like AlphaGo and Dota before asking about reasoning models. Alex explains how o3 deceives opponents in its secret diary while Claude refuses to betray allies.14:21–17:11 · The hosts as informed peer 4/10 Open Trace Logs and Public Game Data Analysis The host calls out duplicate messages on the UI stream, prompting Alex to laughingly admit it was a bug rushed out for the conference. Alex then walks through the public trace logs and dataset structure.17:12–20:39 · The hosts as informed peer 5/10 Context Window Optimization, Prompting, and Developer Tooling The host inquires about long-range context management and developer tooling. Alex shares practical strategies like folder-level llms.txt files and structured experiment logging, which the host validates with his own dev practices.20:39–25:36 · The hosts as informed peer 6/10 Benchmarks as Memes and Democratizing Practical AI Literacy Alex discusses his conference talk on benchmarks as cultural memes and demystifying AI for non-technical users. The host adds his framework separating decision benchmarks for base models from product-specific agent evaluations.25:37–28:30 · The hosts as informed peer 4/10 Human-in-the-Loop Creative Writing Workflows with LLMs The host asks about evaluating creative writing with LLMs. Alex describes Every's iterative human-in-the-loop workflow combining voice dictation, style guides, and repeated prompt refining.28:31–30:46 · The hosts as informed peer 4/10 Reflections and Takeaways from the AI Engineer Conference The host reflects on organizing the AI Engineer conference and the challenges of the AI in Action track. Alex shares positive takeaways on keynotes and praises the event's high density of ideas.30:46–34:18 · The hosts as informed peer 4/10 Future Roadmap: Playable Tournaments and Human Rematches The host and guest discuss future plans for a human versus AI Diplomacy tournament. The host is startled by the 4 to 18 hour game runtime and brainstorms time limits and format rules.1:48–5:00 · Guest teaching 2/10 Every's Product Ecosystem: Cora, Sparkle, and Spiral The host asks about Every's suite of AI products and Alex's consulting practice. Alex provides an overview of internal tools like Cora, Sparkle, and Spiral in a very collaborative introductory exchange.5:00–9:01 · Guest teaching 3/10 Origins and Design of the AI Diplomacy Benchmark The host brings up the Diplomacy benchmark stream and references Noam Brown and DiploBench. Alex shares the project's background and community contributions across the AI ecosystem.9:02–14:20 · Guest teaching 5/10 Games, Self-Play, and the Purpose of Benchmarking The host provides historical context on game benchmarks like AlphaGo and Dota before asking about reasoning models. Alex explains how o3 deceives opponents in its secret diary while Claude refuses to betray allies.14:21–17:11 · Guest teaching 3/10 Open Trace Logs and Public Game Data Analysis The host calls out duplicate messages on the UI stream, prompting Alex to laughingly admit it was a bug rushed out for the conference. Alex then walks through the public trace logs and dataset structure.17:12–20:39 · Guest teaching 4/10 Context Window Optimization, Prompting, and Developer Tooling The host inquires about long-range context management and developer tooling. Alex shares practical strategies like folder-level llms.txt files and structured experiment logging, which the host validates with his own dev practices.20:39–25:36 · Guest teaching 3/10 Benchmarks as Memes and Democratizing Practical AI Literacy Alex discusses his conference talk on benchmarks as cultural memes and demystifying AI for non-technical users. The host adds his framework separating decision benchmarks for base models from product-specific agent evaluations.25:37–28:30 · Guest teaching 4/10 Human-in-the-Loop Creative Writing Workflows with LLMs The host asks about evaluating creative writing with LLMs. Alex describes Every's iterative human-in-the-loop workflow combining voice dictation, style guides, and repeated prompt refining.28:31–30:46 · Guest teaching 2/10 Reflections and Takeaways from the AI Engineer Conference The host reflects on organizing the AI Engineer conference and the challenges of the AI in Action track. Alex shares positive takeaways on keynotes and praises the event's high density of ideas.30:46–34:18 · Guest teaching 3/10 Future Roadmap: Playable Tournaments and Human Rematches The host and guest discuss future plans for a human versus AI Diplomacy tournament. The host is startled by the 4 to 18 hour game runtime and brainstorms time limits and format rules.1:48–5:00 · Guest disagreement 0/10 Every's Product Ecosystem: Cora, Sparkle, and Spiral The host asks about Every's suite of AI products and Alex's consulting practice. Alex provides an overview of internal tools like Cora, Sparkle, and Spiral in a very collaborative introductory exchange.5:00–9:01 · Guest disagreement 0/10 Origins and Design of the AI Diplomacy Benchmark The host brings up the Diplomacy benchmark stream and references Noam Brown and DiploBench. Alex shares the project's background and community contributions across the AI ecosystem.9:02–14:20 · Guest disagreement 0/10 Games, Self-Play, and the Purpose of Benchmarking The host provides historical context on game benchmarks like AlphaGo and Dota before asking about reasoning models. Alex explains how o3 deceives opponents in its secret diary while Claude refuses to betray allies.14:21–17:11 · Guest disagreement 1/10 Open Trace Logs and Public Game Data Analysis The host calls out duplicate messages on the UI stream, prompting Alex to laughingly admit it was a bug rushed out for the conference. Alex then walks through the public trace logs and dataset structure.17:12–20:39 · Guest disagreement 0/10 Context Window Optimization, Prompting, and Developer Tooling The host inquires about long-range context management and developer tooling. Alex shares practical strategies like folder-level llms.txt files and structured experiment logging, which the host validates with his own dev practices.20:39–25:36 · Guest disagreement 1/10 Benchmarks as Memes and Democratizing Practical AI Literacy Alex discusses his conference talk on benchmarks as cultural memes and demystifying AI for non-technical users. The host adds his framework separating decision benchmarks for base models from product-specific agent evaluations.25:37–28:30 · Guest disagreement 0/10 Human-in-the-Loop Creative Writing Workflows with LLMs The host asks about evaluating creative writing with LLMs. Alex describes Every's iterative human-in-the-loop workflow combining voice dictation, style guides, and repeated prompt refining.28:31–30:46 · Guest disagreement 0/10 Reflections and Takeaways from the AI Engineer Conference The host reflects on organizing the AI Engineer conference and the challenges of the AI in Action track. Alex shares positive takeaways on keynotes and praises the event's high density of ideas.30:46–34:18 · Guest disagreement 0/10 Future Roadmap: Playable Tournaments and Human Rematches The host and guest discuss future plans for a human versus AI Diplomacy tournament. The host is startled by the 4 to 18 hour game runtime and brainstorms time limits and format rules.1:48–5:00 · The hosts pushing back 0/10 Every's Product Ecosystem: Cora, Sparkle, and Spiral The host asks about Every's suite of AI products and Alex's consulting practice. Alex provides an overview of internal tools like Cora, Sparkle, and Spiral in a very collaborative introductory exchange.5:00–9:01 · The hosts pushing back 0/10 Origins and Design of the AI Diplomacy Benchmark The host brings up the Diplomacy benchmark stream and references Noam Brown and DiploBench. Alex shares the project's background and community contributions across the AI ecosystem.9:02–14:20 · The hosts pushing back 0/10 Games, Self-Play, and the Purpose of Benchmarking The host provides historical context on game benchmarks like AlphaGo and Dota before asking about reasoning models. Alex explains how o3 deceives opponents in its secret diary while Claude refuses to betray allies.14:21–17:11 · The hosts pushing back 3/10 Open Trace Logs and Public Game Data Analysis The host calls out duplicate messages on the UI stream, prompting Alex to laughingly admit it was a bug rushed out for the conference. Alex then walks through the public trace logs and dataset structure.17:12–20:39 · The hosts pushing back 0/10 Context Window Optimization, Prompting, and Developer Tooling The host inquires about long-range context management and developer tooling. Alex shares practical strategies like folder-level llms.txt files and structured experiment logging, which the host validates with his own dev practices.20:39–25:36 · The hosts pushing back 2/10 Benchmarks as Memes and Democratizing Practical AI Literacy Alex discusses his conference talk on benchmarks as cultural memes and demystifying AI for non-technical users. The host adds his framework separating decision benchmarks for base models from product-specific agent evaluations.25:37–28:30 · The hosts pushing back 0/10 Human-in-the-Loop Creative Writing Workflows with LLMs The host asks about evaluating creative writing with LLMs. Alex describes Every's iterative human-in-the-loop workflow combining voice dictation, style guides, and repeated prompt refining.28:31–30:46 · The hosts pushing back 0/10 Reflections and Takeaways from the AI Engineer Conference The host reflects on organizing the AI Engineer conference and the challenges of the AI in Action track. Alex shares positive takeaways on keynotes and praises the event's high density of ideas.30:46–34:18 · The hosts pushing back 1/10 Future Roadmap: Playable Tournaments and Human Rematches The host and guest discuss future plans for a human versus AI Diplomacy tournament. The host is startled by the 4 to 18 hour game runtime and brainstorms time limits and format rules.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 14:28 Playful reaction to bug callout

Alex humorously pushes back when the host points out duplicate messages streaming on the frontend, asking why the host had to call him out on a launch bug.

Hardest push from the hosts ▶ 14:21 Host challenges frontend streaming behavior

The host directly questions why the live demo is streaming duplicated messages, pressing on whether it is an intentional strategy or an error.

Biggest teaching moment ▶ 12:05 Detailed breakdown of model personalities and deceit

Alex reveals nuanced empirical findings about how different frontier models play Diplomacy, showing how o3 plots betrayals in its hidden diary while Claude consistently declines deceptive plays.

The host holds their own ▶ 23:29 Host categorizes benchmarks versus agent evals

The host articulates an architectural framework distinguishing minimal-scaffold decision benchmarks for assessing raw model traits from heavily scaffolded evals needed for production agents.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Every's Product Ecosystem: Cora, Sparkle, and Spiral 3200 The host asks about Every's suite of AI products and Alex's consulting practice. Alex provides an overview of internal tools like Cora, Sparkle, and Spiral in a very collaborative introductory exchange.
Origins and Design of the AI Diplomacy Benchmark 4300 The host brings up the Diplomacy benchmark stream and references Noam Brown and DiploBench. Alex shares the project's background and community contributions across the AI ecosystem.
Games, Self-Play, and the Purpose of Benchmarking 5500 The host provides historical context on game benchmarks like AlphaGo and Dota before asking about reasoning models. Alex explains how o3 deceives opponents in its secret diary while Claude refuses to betray allies.
Open Trace Logs and Public Game Data Analysis 4313 The host calls out duplicate messages on the UI stream, prompting Alex to laughingly admit it was a bug rushed out for the conference. Alex then walks through the public trace logs and dataset structure.
Context Window Optimization, Prompting, and Developer Tooling 5400 The host inquires about long-range context management and developer tooling. Alex shares practical strategies like folder-level llms.txt files and structured experiment logging, which the host validates with his own dev practices.
Benchmarks as Memes and Democratizing Practical AI Literacy 6312 Alex discusses his conference talk on benchmarks as cultural memes and demystifying AI for non-technical users. The host adds his framework separating decision benchmarks for base models from product-specific agent evaluations.
Human-in-the-Loop Creative Writing Workflows with LLMs 4400 The host asks about evaluating creative writing with LLMs. Alex describes Every's iterative human-in-the-loop workflow combining voice dictation, style guides, and repeated prompt refining.
Reflections and Takeaways from the AI Engineer Conference 4200 The host reflects on organizing the AI Engineer conference and the challenges of the AI in Action track. Alex shares positive takeaways on keynotes and praises the event's high density of ideas.
Future Roadmap: Playable Tournaments and Human Rematches 4301 The host and guest discuss future plans for a human versus AI Diplomacy tournament. The host is startled by the 4 to 18 hour game runtime and brainstorms time limits and format rules.

Statements from this episode (12)

Disclosure
Duffy: Every is releasing Monologue, a local Whisper-style voice tool
“Naveen's actually releasing something called monologue. You know, it's kind of like, I think a better whisper flow for the things that we do can run locally.”
Alex Duffy Jun 11, 2025 ▶ 2:54
Insight
Duffy: Playable AI game benchmarks teach people how LLMs operate
“If we make this playable, you know, then it kind of can teach people how to use AI, like language models just by playing. Cause you'll like understand how they work. You have to negotiate against them. You see their responses.”
Alex Duffy Jun 11, 2025 ▶ 6:36
Assertion Not checkable as stated
Gemini 2.5 Flash runs AI Diplomacy for $1-$5, undercutting competitors
“Running games at two five flash was instant and one to five bucks of, you know, way like, I don't know, 20 to a hundred with the other models.”
Alex Duffy Jun 11, 2025 ▶ 11:45
Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36
Disclosure
Duffy: AI Diplomacy trace logs and dataset are publicly released
“I posted a video on X that shows you how to actually like, we released all the data, all the trace logs.”
Alex Duffy Jun 11, 2025 ▶ 14:51
Insight
Duffy: LLM context should contain enough information for a human to play
“If you're looking at the context that's being sent to the language model, could you play the game?”
Alex Duffy Jun 11, 2025 ▶ 17:24
Disclosure
Duffy: AI Diplomacy used try-catch blocks instead of Pydantic for JSON
“I know one of our GitHub issues already called us out for not using Pydantic parsing and instead of having like eight different try catches for the different JSON versions that come out of these.”
Alex Duffy Jun 11, 2025 ▶ 20:22
Insight
Duffy: AI benchmarks follow a lifecycle from initial idea to saturation
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”
Alex Duffy Jun 11, 2025 ▶ 21:11
Insight
Duffy: Effective LLM writing requires editing earlier context to avoid pollution
“If you're not reflecting on a message coming out and editing above, then like that's something you need to do because you don't want to pollute the context with anything that's not exactly what you're trying to say.”
Alex Duffy Jun 11, 2025 ▶ 27:28
Disclosure
Duffy: Every wants to host a human vs. AI Diplomacy tournament
“I'd love to have a human versus AI diplomacy tournament.”
Alex Duffy Jun 11, 2025 ▶ 31:16
Assertion Partly supported
Duffy: AI Diplomacy games take 4 to 18 hours to play out
“No, human time, watching it, I think it's, like, anywhere from four to 12, four to 18 hours, something, to play out.”
Alex Duffy Jun 11, 2025 ▶ 32:09
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.