“I think one way to think of reflection is that the traditional idea of reinforcement learning is you have a scalar reward, and then you somehow back propagate the signal of the scalar reward. To the rest of your neural network through whatever algorithm, like policy gradient or A to C or whatever. And if you think about the real life, you know, most of the reward signal is not scalar. It's like your boss told you know, you should have done a better job in this, but a good job on that or whatever, right? It's not like a scalar reward, like 29 or something. I think in general, human do more deal more with, you know, long scalar reward, or you can say language feedback, right? And the way That they deal with language feedback also have this kind of back propagation kind of process, right? Because you start from this, you did a good job on job B and then you reflect, you know, what could have done different to change, to make it better. And you kind of change your prompt, right? Basically you change your prompt on how to do job A and how to do job B. And then you do the whole thing again. So it's really like a pipeline of language where it's self-created descent. You have something like tax reasoning to replace Those gradient descent algorithms.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Shunyu Yao
PredictionNot checkable as stated
Shunyu Yao predicts training models on human computer trajectories achieves AGI
“The simplest way to achieve AGI is literally just record the re-actuatory of every human being and just put them together, you know, like what do you have thought about? What do you have done? Let's say on the computer, right? Imagine like solid experiment. Li…”
Shunyu YaoSep 27, 2024▶ 1:02:07Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
AssertionNot checkable as stated
Shunyu Yao says Ilya Sutskever claimed GPT-1 had solved language
“Back in OpenAI, they did this GPT-ONE together, and Ilya just said, Karthik, you should stay, because we just solved the language.”
Shunyu YaoSep 27, 2024▶ 2:12Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao advises developers to default to minimalist prompting for AI agents
“And I think in terms of the actual prompting method to use for a particular problem, I'm I think we should all be in the minimum list kind of camp, right? You should try the minimum thing and see if it works and if it doesn't work and there's absolute reason t…”
Shunyu YaoSep 27, 2024▶ 27:28Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao argues modern LLMs make prompt engineering tricks obsolete
“I feel like in some sense, I feel like prompt engineering, even it's like a slightly negative word at the time, because it refers to all those kind of weird tricks that you have to apply. But I think we don't have to do that anymore. Like given today's progres…”
Shunyu YaoSep 27, 2024▶ 29:31Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Yao: Lack of realistic benchmarks is AI's primary bottleneck
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”
Shunyu YaoSep 27, 2024▶ 31:08Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Insight
Shunyu Yao believes coding is the best application for AI agents
“Obviously coding is the best application for agents because it's all the gradable. It's super important. You can make everything like API or code action, right?”
Shunyu YaoSep 27, 2024▶ 37:04Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.