TAU-bench

2 statements across 2 episodes · 1 bullish · 0 bearish · 2 people on the record · first statement Sep 27, 2024 by Shunyu Yao · said 21 times in 6 episodes since 2024 · across every show →

Mentions by year

brought up most by Pratik Bhavsar (8), Shawn Wang (4), John Yang (2), Bill Chen (2), Alessio Fanelli (2), George Cameron (1)

tap a year for its mentions
0082153202420252026episodesmentions
023202420252026episodes it came up in
0021.543202420252026episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about TAU-bench, oldest first

Sep 27, 2024
Insight
Yao: Enterprise customer support AI requires 99% reliability over simple tasks, not search
“It's very different from coding or web agent or whatever people are doing, because it's more about how can you do simple things reliably It's not about, you know, can you sample a hundred times and you find one good mass proof or kill solution. It's more about…”
Shunyu Yao Sep 27, 2024 ▶ 1:17:44 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Jan 9, 2026 positive
Assertion Contradicted
Frontier models are cheaper for agentic tasks because they require fewer turns
“Interestingly, in Tau Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the some of the GBD five, for instance got to the answer faster. And …”
George Cameron Jan 9, 2026 ▶ 1:09:14 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.