Alex Duffy

Co-Founder & CEO, Good Start Labs · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderengineerexecutiveauthor@alxai_ ↗LinkedIn ↗alxai.com ↗

Duffy founded Good Start Labs, a startup spun out of Every that uses gaming environments and gameplay to train and evaluate AI models. He is widely known for creating the AI Diplomacy benchmark—an evaluation harness that pits large language models against one another in full-press strategic negotiation.

12statements → 4claims → 3claims resolved → 67%fully supported → 3.67/5average certainty → 1.58/5average debate potential →

2 supported 1 partly supported 0 contradicted 1 not checkable as stated how the 4 claims stand · each chip opens the sources

4 assertions · 4 insights · 4 disclosures · every statement was checked. The predictions and assertions are the 4 claims: statements the public record can support or contradict. 3 are resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Alex argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
75% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything Alex Duffy said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Insight
Duffy: Playable AI game benchmarks teach people how LLMs operate
“If we make this playable, you know, then it kind of can teach people how to use AI, like language models just by playing. Cause you'll like understand how they work. You have to negotiate against them. You see their responses.”
Alex Duffy Jun 11, 2025 ▶ 6:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Insight
Duffy: AI benchmarks follow a lifecycle from initial idea to saturation
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”
Alex Duffy Jun 11, 2025 ▶ 21:11 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Insight
Duffy: Effective LLM writing requires editing earlier context to avoid pollution
“If you're not reflecting on a message coming out and editing above, then like that's something you need to do because you don't want to pollute the context with anything that's not exactly what you're trying to say.”
Alex Duffy Jun 11, 2025 ▶ 27:28 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Not checkable as stated
Gemini 2.5 Flash runs AI Diplomacy for $1-$5, undercutting competitors
“Running games at two five flash was instant and one to five bucks of, you know, way like, I don't know, 20 to a hundred with the other models.”
Alex Duffy Jun 11, 2025 ▶ 11:45 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Disclosure
Duffy: AI Diplomacy trace logs and dataset are publicly released
“I posted a video on X that shows you how to actually like, we released all the data, all the trace logs.”
Alex Duffy Jun 11, 2025 ▶ 14:51 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Insight
Duffy: LLM context should contain enough information for a human to play
“If you're looking at the context that's being sent to the language model, could you play the game?”
Alex Duffy Jun 11, 2025 ▶ 17:24 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Disclosure
Duffy: Every wants to host a human vs. AI Diplomacy tournament
“I'd love to have a human versus AI diplomacy tournament.”
Alex Duffy Jun 11, 2025 ▶ 31:16 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Disclosure
Duffy: Every is releasing Monologue, a local Whisper-style voice tool
“Naveen's actually releasing something called monologue. You know, it's kind of like, I think a better whisper flow for the things that we do can run locally.”
Alex Duffy Jun 11, 2025 ▶ 2:54 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Disclosure
Duffy: AI Diplomacy used try-catch blocks instead of Pydantic for JSON
“I know one of our GitHub issues already called us out for not using Pydantic parsing and instead of having like eight different try catches for the different JSON versions that come out of these.”
Alex Duffy Jun 11, 2025 ▶ 20:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Partly supported
Duffy: AI Diplomacy games take 4 to 18 hours to play out
“No, human time, watching it, I think it's, like, anywhere from four to 12, four to 18 hours, something, to play out.”
Alex Duffy Jun 11, 2025 ▶ 32:09 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy

Appearances (1)

EpisodeDateSpeaking time
⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy Jun 11, 2025 22m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.