Insight certainty 3/5 debate potential 3/5

Yao: Lack of realistic benchmarks is AI's primary bottleneck

Shunyu Yao · Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph · Sep 27, 2024 · at 31:08

Shunyu Yao, AI researcher and creator of ReAct and SWE-bench, explains why benchmarking is lagging behind inference and search methodologies.

0:00 / 0:07exact quote · 7.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Shunyu Yao

Prediction Not checkable as stated
Shunyu Yao predicts training models on human computer trajectories achieves AGI
“The simplest way to achieve AGI is literally just record the re-actuatory of every human being and just put them together, you know, like what do you have thought about? What do you have done? Let's say on the computer, right? Imagine like solid experiment. Li…”
Shunyu Yao Sep 27, 2024 ▶ 1:02:07 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Assertion Not checkable as stated
Shunyu Yao says Ilya Sutskever claimed GPT-1 had solved language
“Back in OpenAI, they did this GPT-ONE together, and Ilya just said, Karthik, you should stay, because we just solved the language.”
Shunyu Yao Sep 27, 2024 ▶ 2:12 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao advises developers to default to minimalist prompting for AI agents
“And I think in terms of the actual prompting method to use for a particular problem, I'm I think we should all be in the minimum list kind of camp, right? You should try the minimum thing and see if it works and if it doesn't work and there's absolute reason t…”
Shunyu Yao Sep 27, 2024 ▶ 27:28 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao argues modern LLMs make prompt engineering tricks obsolete
“I feel like in some sense, I feel like prompt engineering, even it's like a slightly negative word at the time, because it refers to all those kind of weird tricks that you have to apply. But I think we don't have to do that anymore. Like given today's progres…”
Shunyu Yao Sep 27, 2024 ▶ 29:31 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao believes coding is the best application for AI agents
“Obviously coding is the best application for agents because it's all the gradable. It's super important. You can make everything like API or code action, right?”
Shunyu Yao Sep 27, 2024 ▶ 37:04 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Yao: Reliable tool design accounts for 90% of agent performance
“I think making the tool good and reliable is probably like 90% of the whole agent. Once the tool is actually good, then the agent design can be much, much simpler. On the other hand, if the tool is bad, then no matter how much you put into the agent design pla…”
Shunyu Yao Sep 27, 2024 ▶ 45:35 Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.