OpenAI researcher Shunyu Yao explains the motivation behind the Tau-Bench benchmark and how enterprise support agents differ from coding or reasoning benchmarks.
“It's very different from coding or web agent or whatever people are doing, because it's more about how can you do simple things reliably
It's not about, you know, can you sample a hundred times and you find one good mass proof or kill solution.
It's more about you chat with a hundred different users on very simple things.
Can you be robust to solve like 99% of the time, right?”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Shunyu Yao
PredictionNot checkable as stated
Shunyu Yao predicts training models on human computer trajectories achieves AGI
“The simplest way to achieve AGI is literally just record the re-actuatory of every human being and just put them together, you know, like what do you have thought about? What do you have done? Let's say on the computer, right? Imagine like solid experiment. Li…”
Shunyu YaoSep 27, 2024▶ 1:02:07Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
AssertionNot checkable as stated
Shunyu Yao says Ilya Sutskever claimed GPT-1 had solved language
“Back in OpenAI, they did this GPT-ONE together, and Ilya just said, Karthik, you should stay, because we just solved the language.”
Shunyu YaoSep 27, 2024▶ 2:12Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao advises developers to default to minimalist prompting for AI agents
“And I think in terms of the actual prompting method to use for a particular problem, I'm I think we should all be in the minimum list kind of camp, right? You should try the minimum thing and see if it works and if it doesn't work and there's absolute reason t…”
Shunyu YaoSep 27, 2024▶ 27:28Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao argues modern LLMs make prompt engineering tricks obsolete
“I feel like in some sense, I feel like prompt engineering, even it's like a slightly negative word at the time, because it refers to all those kind of weird tricks that you have to apply. But I think we don't have to do that anymore. Like given today's progres…”
Shunyu YaoSep 27, 2024▶ 29:31Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Yao: Lack of realistic benchmarks is AI's primary bottleneck
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”
Shunyu YaoSep 27, 2024▶ 31:08Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao believes coding is the best application for AI agents
“Obviously coding is the best application for agents because it's all the gradable. It's super important. You can make everything like API or code action, right?”
Shunyu YaoSep 27, 2024▶ 37:04Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.