TAU-bench, every mention

8 scenes · ← back to TAU-bench

tap a year for its mentions
0082153202420252026episodesmentions
023202420252026episodes it came up in
0021.543202420252026episodesmentions per episode

every year anyone Pratik Bhavsar 8Shawn Wang 4John Yang 2Bill Chen 2Alessio Fanelli 2George Cameron 1

Verbatim, from the transcripts: the passages where TAU-bench comes up

loading…

The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI Jul 22, 2026 · 2 mentions

  • ▶ 58:25 unnamed speaker like on certain benchmarks, like the tilebench one, like you're actually state of the art. 2 times in the scene

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith Jan 9, 2026 · 5 mentions

  • ▶ 1:09:14 George Cameron Interestingly, in Tau, um, Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the, um, some of the GBD five, for…
  • ▶ 1:16:14 Shawn Wang Uh, there, there's a little bit of debate over the, the accuracy of tile bench. 4 times in the scene

[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang Dec 31, 2025 · 2 mentions

  • ▶ 8:24 John Yang I'm personally think it's quite interesting to think about the user simulator stuff, so like, Tao Bench. 2 times in the scene

⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI Dec 26, 2025 · 2 mentions

  • ▶ 20:44 Bill Chen And then there are, like, also academic benchmarks already does this in some ways, like CowBench, and now we have, like, Tau SquareBench 2 times in the scene

⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo Jul 14, 2025 · 8 mentions

  • ▶ 12:58 Pratik Bhavsar I do not remember exactly which one has it and which one doesn't, but there were these discrepancies and we really enjoyed, like I really loved the Tau benchmark also cause it had 5 times in the scene
  • ▶ 26:33 Pratik Bhavsar So we kind of mixed up the dataset during that time. 3 times in the scene

Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph Sep 27, 2024 · 2 mentions

  • ▶ 1:16:06 Alessio Fanelli Just one thing I want to highlight from your work, we don't have to go into it, is, uh, Tao Bench. 2 times in the scene
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.