TAU-bench, every mention
8 scenes · ← back to TAU-bench
tap a year for its mentions
every year anyone Pratik Bhavsar 8Shawn Wang 4John Yang 2Bill Chen 2Alessio Fanelli 2George Cameron 1
Verbatim, from the transcripts: the passages where TAU-bench comes up
The AI Frontier: from open weights to open research — Eiso Kant, Poolside AI
- ▶ 58:25 unnamed speaker like on certain benchmarks, like the tilebench one, like you're actually state of the art. 2 times in the scene
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 1:09:14 George Cameron Interestingly, in Tau, um, Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the, um, some of the GBD five, for…
- ▶ 1:16:14 Shawn Wang Uh, there, there's a little bit of debate over the, the accuracy of tile bench. 4 times in the scene
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust — Brian Fioca + Bill Chen, OpenAI
⚡️Ranking Agentic LLMs — Pratik Bhavsar, Galileo
- ▶ 12:58 Pratik Bhavsar I do not remember exactly which one has it and which one doesn't, but there were these discrepancies and we really enjoyed, like I really loved the Tau benchmark also cause it had 5 times in the scene
- ▶ 26:33 Pratik Bhavsar So we kind of mixed up the dataset during that time. 3 times in the scene
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 1:16:06 Alessio Fanelli Just one thing I want to highlight from your work, we don't have to go into it, is, uh, Tao Bench. 2 times in the scene