Jun 11, 2025 · 34m · latent-space
⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Alex Duffy, Head of AI at Every, breaks down the creation and strategic insights of the AI Diplomacy game benchmark while discussing Every's multidisciplinary culture, practical AI tooling, and human-in-the-loop workflows.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Alex humorously pushes back when the host points out duplicate messages streaming on the frontend, asking why the host had to call him out on a launch bug.
Hardest push from the hosts ▶ 14:21 Host challenges frontend streaming behaviorThe host directly questions why the live demo is streaming duplicated messages, pressing on whether it is an intentional strategy or an error.
Biggest teaching moment ▶ 12:05 Detailed breakdown of model personalities and deceitAlex reveals nuanced empirical findings about how different frontier models play Diplomacy, showing how o3 plots betrayals in its hidden diary while Claude consistently declines deceptive plays.
The host holds their own ▶ 23:29 Host categorizes benchmarks versus agent evalsThe host articulates an architectural framework distinguishing minimal-scaffold decision benchmarks for assessing raw model traits from heavily scaffolded evals needed for production agents.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Every's Product Ecosystem: Cora, Sparkle, and Spiral | 3 | 2 | 0 | 0 | The host asks about Every's suite of AI products and Alex's consulting practice. Alex provides an overview of internal tools like Cora, Sparkle, and Spiral in a very collaborative introductory exchange. | |
| Origins and Design of the AI Diplomacy Benchmark | 4 | 3 | 0 | 0 | The host brings up the Diplomacy benchmark stream and references Noam Brown and DiploBench. Alex shares the project's background and community contributions across the AI ecosystem. | |
| Games, Self-Play, and the Purpose of Benchmarking | 5 | 5 | 0 | 0 | The host provides historical context on game benchmarks like AlphaGo and Dota before asking about reasoning models. Alex explains how o3 deceives opponents in its secret diary while Claude refuses to betray allies. | |
| Open Trace Logs and Public Game Data Analysis | 4 | 3 | 1 | 3 | The host calls out duplicate messages on the UI stream, prompting Alex to laughingly admit it was a bug rushed out for the conference. Alex then walks through the public trace logs and dataset structure. | |
| Context Window Optimization, Prompting, and Developer Tooling | 5 | 4 | 0 | 0 | The host inquires about long-range context management and developer tooling. Alex shares practical strategies like folder-level llms.txt files and structured experiment logging, which the host validates with his own dev practices. | |
| Benchmarks as Memes and Democratizing Practical AI Literacy | 6 | 3 | 1 | 2 | Alex discusses his conference talk on benchmarks as cultural memes and demystifying AI for non-technical users. The host adds his framework separating decision benchmarks for base models from product-specific agent evaluations. | |
| Human-in-the-Loop Creative Writing Workflows with LLMs | 4 | 4 | 0 | 0 | The host asks about evaluating creative writing with LLMs. Alex describes Every's iterative human-in-the-loop workflow combining voice dictation, style guides, and repeated prompt refining. | |
| Reflections and Takeaways from the AI Engineer Conference | 4 | 2 | 0 | 0 | The host reflects on organizing the AI Engineer conference and the challenges of the AI in Action track. Alex shares positive takeaways on keynotes and praises the event's high density of ideas. | |
| Future Roadmap: Playable Tournaments and Human Rematches | 4 | 3 | 0 | 1 | The host and guest discuss future plans for a human versus AI Diplomacy tournament. The host is startled by the 4 to 18 hour game runtime and brainstorms time limits and format rules. |