Insight certainty 3/5 debate potential 2/5

Duffy: AI benchmarks follow a lifecycle from initial idea to saturation

Alex Duffy · ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy · Jun 11, 2025 · at 21:11

Alex Duffy, Head of AI at Every, discusses the nature of AI evals and benchmarks as cultural memes spreading across the machine learning community.

0:00 / 0:05exact quote · 5.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Essentially there's, I think, a life cycle of a benchmark, right? It starts with an idea, then it gets adopted, and then it gets saturated.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Alex Duffy

Assertion Supported
Duffy: OpenAI's o3 actively deceives opponents and plots betrayals in AI Diplomacy
“Oh, three was one of the few that will actually send a message to another power saying that they're planning to do something. And then like in their diary diary, right? Oh, they fell for it. Hook, line and sinker. Totally gonna betray him and take it over.”
Alex Duffy Jun 11, 2025 ▶ 12:22 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Supported
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Alex Duffy Jun 11, 2025 ▶ 12:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Insight
Duffy: Playable AI game benchmarks teach people how LLMs operate
“If we make this playable, you know, then it kind of can teach people how to use AI, like language models just by playing. Cause you'll like understand how they work. You have to negotiate against them. You see their responses.”
Alex Duffy Jun 11, 2025 ▶ 6:36 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Insight
Duffy: Effective LLM writing requires editing earlier context to avoid pollution
“If you're not reflecting on a message coming out and editing above, then like that's something you need to do because you don't want to pollute the context with anything that's not exactly what you're trying to say.”
Alex Duffy Jun 11, 2025 ▶ 27:28 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Assertion Not checkable as stated
Gemini 2.5 Flash runs AI Diplomacy for $1-$5, undercutting competitors
“Running games at two five flash was instant and one to five bucks of, you know, way like, I don't know, 20 to a hundred with the other models.”
Alex Duffy Jun 11, 2025 ▶ 11:45 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Disclosure
Duffy: AI Diplomacy trace logs and dataset are publicly released
“I posted a video on X that shows you how to actually like, we released all the data, all the trace logs.”
Alex Duffy Jun 11, 2025 ▶ 14:51 ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.