People, every show

Alex Duffy

Co-Founder & CEO, Good Start Labs. On 1 show, 1 appearance. The Shows tab opens the full record on each.

founderengineerexecutiveauthor@alxai_ ↗LinkedIn ↗alxai.com ↗

Duffy founded Good Start Labs, a startup spun out of Every that uses gaming environments and gameplay to train and evaluate AI models. He is widely known for creating the AI Diplomacy benchmark—an evaluation harness that pits large language models against one another in full-press strategic negotiation.

1shows
1appearances
12statements
3resolved
2supported
0contradicted
67%fully supported

Everything Alex Duffy said on any show that made the record, most notable first. Each card names its show and opens the statement there.

LATENT SPACE Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
LATENT SPACE Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Duffy: Playable AI game benchmarks teach people how LLMs operate
“If we make this playable, you know, then it kind of can teach people how to use AI, like language models just by playing. Cause you'll like understand how they work. You have to negotiate against them. You see their responses.”
Alex Duffy Jun 11, 2025 ▶ 6:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Duffy: AI benchmarks follow a lifecycle from initial idea to saturation
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”
Alex Duffy Jun 11, 2025 ▶ 21:11 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Duffy: Effective LLM writing requires editing earlier context to avoid pollution
“If you're not reflecting on a message coming out and editing above, then like that's something you need to do because you don't want to pollute the context with anything that's not exactly what you're trying to say.”
Alex Duffy Jun 11, 2025 ▶ 27:28 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
LATENT SPACE Assertion Not checkable as stated
Gemini 2.5 Flash runs AI Diplomacy for $1-$5, undercutting competitors
“Running games at two five flash was instant and one to five bucks of, you know, way like, I don't know, 20 to a hundred with the other models.”
Alex Duffy Jun 11, 2025 ▶ 11:45 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
LATENT SPACE Disclosure
Duffy: AI Diplomacy trace logs and dataset are publicly released
“I posted a video on X that shows you how to actually like, we released all the data, all the trace logs.”
Alex Duffy Jun 11, 2025 ▶ 14:51 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Duffy: LLM context should contain enough information for a human to play
“If you're looking at the context that's being sent to the language model, could you play the game?”
Alex Duffy Jun 11, 2025 ▶ 17:24 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
LATENT SPACE Disclosure
Duffy: Every wants to host a human vs. AI Diplomacy tournament
“I'd love to have a human versus AI diplomacy tournament.”
Alex Duffy Jun 11, 2025 ▶ 31:16 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
LATENT SPACE Disclosure
Duffy: Every is releasing Monologue, a local Whisper-style voice tool
“Naveen's actually releasing something called monologue. You know, it's kind of like, I think a better whisper flow for the things that we do can run locally.”
Alex Duffy Jun 11, 2025 ▶ 2:54 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
LATENT SPACE Disclosure
Duffy: AI Diplomacy used try-catch blocks instead of Pydantic for JSON
“I know one of our GitHub issues already called us out for not using Pydantic parsing and instead of having like eight different try catches for the different JSON versions that come out of these.”
Alex Duffy Jun 11, 2025 ▶ 20:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
LATENT SPACE Assertion Partly supported
Duffy: AI Diplomacy games take 4 to 18 hours to play out
“No, human time, watching it, I think it's, like, anywhere from four to 12, four to 18 hours, something, to play out.”
Alex Duffy Jun 11, 2025 ▶ 32:09 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy

One line per show, most statements first. The link opens Alex's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER Co-Founder & CEO, Good Start Labs 1 12 67% 2/3 full record on Latent Space →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.