Sep 27, 2024 · 1h 26m · latent-space

Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph

Shunyu Yao · 45m spoken Harrison Chase · 18m spoken Shawn Wang · 10m spoken Alessio Fanelli · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

LangChain creator Harrison Chase and researcher Shunyu Yao join the Latent Space podcast to discuss the evolution, cognitive architectures, evaluation benchmarks, and tooling design powering modern language agents. They explore foundational concepts from ReAct and Reflexion to Agent-Computer Interfaces (ACI) and stateful production orchestration with LangGraph.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 20.7% of the talking time here. How this is scored →

The hosts as informed peer 6.5 Guest teaching 5.2 Guest disagreement 1.7 The hosts pushing back 2.0
05100:0020:0040:001:00:001:20:000:03–9:08 · The hosts as informed peer 6/10 Welcome and Introductions with Harrison Chase and Shunyu Yao The hosts open with deep familiarity regarding Shunyu Yao's PhD defense and Harrison Chase's early inspiration from the ReAct paper. Shunyu provides historical context on text-adventure games like Zork I and the shift from RL to LLM-driven reasoning. The exchange is warm, collaborative, and appreciative.9:09–13:56 · The hosts as informed peer 6/10 ReAct's Legacy, Tool Calling, and the Modern Agent Loop Harrison and Shunyu discuss whether modern function-calling loops still reflect ReAct principles. Shunyu clarifies that inner monologues remain crucial when tool APIs diverge from pre-training data distributions, while Alessio asks how thinking steps are being internalized into base model weights.13:56–23:36 · The hosts as informed peer 7/10 Reflexion, Language-Based Feedback, and Agent Memory Architectures Shunyu explains Reflexion as substituting scalar RL rewards with rich linguistic feedback acting as linguistic gradient descent. The hosts and Harrison connect this to LangMem, Voyager, and cognitive science categorizations of semantic versus procedural memory.23:38–30:43 · The hosts as informed peer 6/10 Tree of Thoughts, Search Algorithms, and Prompting Philosophy The discussion covers Tree of Thoughts, search vs. reactive tasks, and the minimalist prompting philosophy. Shunyu rejects complex emotional prompt hacks in favor of plain human-to-human communication, prompting Harrison and Swyx to analyze user communication hurdles.30:44–35:58 · The hosts as informed peer 6/10 MCTS, The Benchmark Bottleneck, and Evaluation Realism Shunyu critiques the academic habit of applying over-complex methods to trivial benchmarks rather than building scalable, realistic tasks. He unpacks the trilemma of benchmark design between auto-gradability, realism, and scalability.35:59–45:59 · The hosts as informed peer 7/10 Interactive Coding, SWE-Agent, and Agent-Computer Interfaces (ACI) Shunyu details the progression from InterCode to SWE-bench and SWE-agent, highlighting that optimizing the Agent-Computer Interface (ACI) yields far higher returns than raw planning algorithms. The hosts draw parallels to Cognition's Devin and user interaction paradigms.46:00–54:51 · The hosts as informed peer 8/10 Optimizing Interfaces for LLMs: ACI vs. HCI and Machine Limits Swyx pushes back against the premise that HCI and ACI design completely overlap, citing his empirical experience with structured output field naming and verbose error traces. Shunyu contrasts human working memory constraints with machine context limits, while Harrison argues for allowing agent interfaces to diverge early.54:53–1:04:45 · The hosts as informed peer 7/10 Separating Intelligence from Knowledge and Training Trajectory Data Swyx poses whether intelligence can be cleanly separated from knowledge, pointing to on-device LoRA adapters. Shunyu pushes back, arguing from historical AI and Hinton's perspective that intelligence emerges with knowledge, though framing knowledge as a cache for intelligence.1:04:46–1:16:04 · The hosts as informed peer 7/10 CoALA Framework, Action Spaces, and Stateful Orchestration Shunyu walks through the CoALA framework across memory, action space, and decision-making. Harrison and Swyx integrate LangGraph's stateful cross-thread persistence model into the architecture, debating whether developers or agents should select memory tooling.1:16:05–1:23:22 · The hosts as informed peer 6/10 Tau-Bench, Customer Simulation, and Practical Enterprise Agents Alessio and Shunyu discuss Tau-Bench and customer service agent simulation where LLMs model realistic user behavior under information asymmetry. Harrison answers Shunyu's question about enterprise applications, pointing to customer support, SDR data enrichment, and spreadsheet agents.1:23:22–1:26:22 · The hosts as informed peer 6/10 LangGraph Studio, Agent IDEs, and Future Tooling Shunyu questions whether low-code tooling is truly ready for non-programmers, and Harrison clarifies that LangGraph Studio acts as an IDE to bridge developer architecture with PM prompt refinement. The group wraps up with reflections on future agent UX.0:03–9:08 · Guest teaching 5/10 Welcome and Introductions with Harrison Chase and Shunyu Yao The hosts open with deep familiarity regarding Shunyu Yao's PhD defense and Harrison Chase's early inspiration from the ReAct paper. Shunyu provides historical context on text-adventure games like Zork I and the shift from RL to LLM-driven reasoning. The exchange is warm, collaborative, and appreciative.9:09–13:56 · Guest teaching 5/10 ReAct's Legacy, Tool Calling, and the Modern Agent Loop Harrison and Shunyu discuss whether modern function-calling loops still reflect ReAct principles. Shunyu clarifies that inner monologues remain crucial when tool APIs diverge from pre-training data distributions, while Alessio asks how thinking steps are being internalized into base model weights.13:56–23:36 · Guest teaching 6/10 Reflexion, Language-Based Feedback, and Agent Memory Architectures Shunyu explains Reflexion as substituting scalar RL rewards with rich linguistic feedback acting as linguistic gradient descent. The hosts and Harrison connect this to LangMem, Voyager, and cognitive science categorizations of semantic versus procedural memory.23:38–30:43 · Guest teaching 4/10 Tree of Thoughts, Search Algorithms, and Prompting Philosophy The discussion covers Tree of Thoughts, search vs. reactive tasks, and the minimalist prompting philosophy. Shunyu rejects complex emotional prompt hacks in favor of plain human-to-human communication, prompting Harrison and Swyx to analyze user communication hurdles.30:44–35:58 · Guest teaching 6/10 MCTS, The Benchmark Bottleneck, and Evaluation Realism Shunyu critiques the academic habit of applying over-complex methods to trivial benchmarks rather than building scalable, realistic tasks. He unpacks the trilemma of benchmark design between auto-gradability, realism, and scalability.35:59–45:59 · Guest teaching 6/10 Interactive Coding, SWE-Agent, and Agent-Computer Interfaces (ACI) Shunyu details the progression from InterCode to SWE-bench and SWE-agent, highlighting that optimizing the Agent-Computer Interface (ACI) yields far higher returns than raw planning algorithms. The hosts draw parallels to Cognition's Devin and user interaction paradigms.46:00–54:51 · Guest teaching 5/10 Optimizing Interfaces for LLMs: ACI vs. HCI and Machine Limits Swyx pushes back against the premise that HCI and ACI design completely overlap, citing his empirical experience with structured output field naming and verbose error traces. Shunyu contrasts human working memory constraints with machine context limits, while Harrison argues for allowing agent interfaces to diverge early.54:53–1:04:45 · Guest teaching 6/10 Separating Intelligence from Knowledge and Training Trajectory Data Swyx poses whether intelligence can be cleanly separated from knowledge, pointing to on-device LoRA adapters. Shunyu pushes back, arguing from historical AI and Hinton's perspective that intelligence emerges with knowledge, though framing knowledge as a cache for intelligence.1:04:46–1:16:04 · Guest teaching 5/10 CoALA Framework, Action Spaces, and Stateful Orchestration Shunyu walks through the CoALA framework across memory, action space, and decision-making. Harrison and Swyx integrate LangGraph's stateful cross-thread persistence model into the architecture, debating whether developers or agents should select memory tooling.1:16:05–1:23:22 · Guest teaching 5/10 Tau-Bench, Customer Simulation, and Practical Enterprise Agents Alessio and Shunyu discuss Tau-Bench and customer service agent simulation where LLMs model realistic user behavior under information asymmetry. Harrison answers Shunyu's question about enterprise applications, pointing to customer support, SDR data enrichment, and spreadsheet agents.1:23:22–1:26:22 · Guest teaching 4/10 LangGraph Studio, Agent IDEs, and Future Tooling Shunyu questions whether low-code tooling is truly ready for non-programmers, and Harrison clarifies that LangGraph Studio acts as an IDE to bridge developer architecture with PM prompt refinement. The group wraps up with reflections on future agent UX.0:03–9:08 · Guest disagreement 1/10 Welcome and Introductions with Harrison Chase and Shunyu Yao The hosts open with deep familiarity regarding Shunyu Yao's PhD defense and Harrison Chase's early inspiration from the ReAct paper. Shunyu provides historical context on text-adventure games like Zork I and the shift from RL to LLM-driven reasoning. The exchange is warm, collaborative, and appreciative.9:09–13:56 · Guest disagreement 2/10 ReAct's Legacy, Tool Calling, and the Modern Agent Loop Harrison and Shunyu discuss whether modern function-calling loops still reflect ReAct principles. Shunyu clarifies that inner monologues remain crucial when tool APIs diverge from pre-training data distributions, while Alessio asks how thinking steps are being internalized into base model weights.13:56–23:36 · Guest disagreement 1/10 Reflexion, Language-Based Feedback, and Agent Memory Architectures Shunyu explains Reflexion as substituting scalar RL rewards with rich linguistic feedback acting as linguistic gradient descent. The hosts and Harrison connect this to LangMem, Voyager, and cognitive science categorizations of semantic versus procedural memory.23:38–30:43 · Guest disagreement 2/10 Tree of Thoughts, Search Algorithms, and Prompting Philosophy The discussion covers Tree of Thoughts, search vs. reactive tasks, and the minimalist prompting philosophy. Shunyu rejects complex emotional prompt hacks in favor of plain human-to-human communication, prompting Harrison and Swyx to analyze user communication hurdles.30:44–35:58 · Guest disagreement 2/10 MCTS, The Benchmark Bottleneck, and Evaluation Realism Shunyu critiques the academic habit of applying over-complex methods to trivial benchmarks rather than building scalable, realistic tasks. He unpacks the trilemma of benchmark design between auto-gradability, realism, and scalability.35:59–45:59 · Guest disagreement 1/10 Interactive Coding, SWE-Agent, and Agent-Computer Interfaces (ACI) Shunyu details the progression from InterCode to SWE-bench and SWE-agent, highlighting that optimizing the Agent-Computer Interface (ACI) yields far higher returns than raw planning algorithms. The hosts draw parallels to Cognition's Devin and user interaction paradigms.46:00–54:51 · Guest disagreement 3/10 Optimizing Interfaces for LLMs: ACI vs. HCI and Machine Limits Swyx pushes back against the premise that HCI and ACI design completely overlap, citing his empirical experience with structured output field naming and verbose error traces. Shunyu contrasts human working memory constraints with machine context limits, while Harrison argues for allowing agent interfaces to diverge early.54:53–1:04:45 · Guest disagreement 3/10 Separating Intelligence from Knowledge and Training Trajectory Data Swyx poses whether intelligence can be cleanly separated from knowledge, pointing to on-device LoRA adapters. Shunyu pushes back, arguing from historical AI and Hinton's perspective that intelligence emerges with knowledge, though framing knowledge as a cache for intelligence.1:04:46–1:16:04 · Guest disagreement 1/10 CoALA Framework, Action Spaces, and Stateful Orchestration Shunyu walks through the CoALA framework across memory, action space, and decision-making. Harrison and Swyx integrate LangGraph's stateful cross-thread persistence model into the architecture, debating whether developers or agents should select memory tooling.1:16:05–1:23:22 · Guest disagreement 1/10 Tau-Bench, Customer Simulation, and Practical Enterprise Agents Alessio and Shunyu discuss Tau-Bench and customer service agent simulation where LLMs model realistic user behavior under information asymmetry. Harrison answers Shunyu's question about enterprise applications, pointing to customer support, SDR data enrichment, and spreadsheet agents.1:23:22–1:26:22 · Guest disagreement 2/10 LangGraph Studio, Agent IDEs, and Future Tooling Shunyu questions whether low-code tooling is truly ready for non-programmers, and Harrison clarifies that LangGraph Studio acts as an IDE to bridge developer architecture with PM prompt refinement. The group wraps up with reflections on future agent UX.0:03–9:08 · The hosts pushing back 1/10 Welcome and Introductions with Harrison Chase and Shunyu Yao The hosts open with deep familiarity regarding Shunyu Yao's PhD defense and Harrison Chase's early inspiration from the ReAct paper. Shunyu provides historical context on text-adventure games like Zork I and the shift from RL to LLM-driven reasoning. The exchange is warm, collaborative, and appreciative.9:09–13:56 · The hosts pushing back 2/10 ReAct's Legacy, Tool Calling, and the Modern Agent Loop Harrison and Shunyu discuss whether modern function-calling loops still reflect ReAct principles. Shunyu clarifies that inner monologues remain crucial when tool APIs diverge from pre-training data distributions, while Alessio asks how thinking steps are being internalized into base model weights.13:56–23:36 · The hosts pushing back 2/10 Reflexion, Language-Based Feedback, and Agent Memory Architectures Shunyu explains Reflexion as substituting scalar RL rewards with rich linguistic feedback acting as linguistic gradient descent. The hosts and Harrison connect this to LangMem, Voyager, and cognitive science categorizations of semantic versus procedural memory.23:38–30:43 · The hosts pushing back 2/10 Tree of Thoughts, Search Algorithms, and Prompting Philosophy The discussion covers Tree of Thoughts, search vs. reactive tasks, and the minimalist prompting philosophy. Shunyu rejects complex emotional prompt hacks in favor of plain human-to-human communication, prompting Harrison and Swyx to analyze user communication hurdles.30:44–35:58 · The hosts pushing back 1/10 MCTS, The Benchmark Bottleneck, and Evaluation Realism Shunyu critiques the academic habit of applying over-complex methods to trivial benchmarks rather than building scalable, realistic tasks. He unpacks the trilemma of benchmark design between auto-gradability, realism, and scalability.35:59–45:59 · The hosts pushing back 2/10 Interactive Coding, SWE-Agent, and Agent-Computer Interfaces (ACI) Shunyu details the progression from InterCode to SWE-bench and SWE-agent, highlighting that optimizing the Agent-Computer Interface (ACI) yields far higher returns than raw planning algorithms. The hosts draw parallels to Cognition's Devin and user interaction paradigms.46:00–54:51 · The hosts pushing back 4/10 Optimizing Interfaces for LLMs: ACI vs. HCI and Machine Limits Swyx pushes back against the premise that HCI and ACI design completely overlap, citing his empirical experience with structured output field naming and verbose error traces. Shunyu contrasts human working memory constraints with machine context limits, while Harrison argues for allowing agent interfaces to diverge early.54:53–1:04:45 · The hosts pushing back 3/10 Separating Intelligence from Knowledge and Training Trajectory Data Swyx poses whether intelligence can be cleanly separated from knowledge, pointing to on-device LoRA adapters. Shunyu pushes back, arguing from historical AI and Hinton's perspective that intelligence emerges with knowledge, though framing knowledge as a cache for intelligence.1:04:46–1:16:04 · The hosts pushing back 2/10 CoALA Framework, Action Spaces, and Stateful Orchestration Shunyu walks through the CoALA framework across memory, action space, and decision-making. Harrison and Swyx integrate LangGraph's stateful cross-thread persistence model into the architecture, debating whether developers or agents should select memory tooling.1:16:05–1:23:22 · The hosts pushing back 1/10 Tau-Bench, Customer Simulation, and Practical Enterprise Agents Alessio and Shunyu discuss Tau-Bench and customer service agent simulation where LLMs model realistic user behavior under information asymmetry. Harrison answers Shunyu's question about enterprise applications, pointing to customer support, SDR data enrichment, and spreadsheet agents.1:23:22–1:26:22 · The hosts pushing back 2/10 LangGraph Studio, Agent IDEs, and Future Tooling Shunyu questions whether low-code tooling is truly ready for non-programmers, and Harrison clarifies that LangGraph Studio acts as an IDE to bridge developer architecture with PM prompt refinement. The group wraps up with reflections on future agent UX.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 50.7% · guest 49.3%0:00 · the hosts 50.7% · guest 49.3%3:00 · the hosts 34% · guest 66%3:00 · the hosts 34% · guest 66%6:00 · the hosts 7.2% · guest 92.8%6:00 · the hosts 7.2% · guest 92.8%9:00 · the hosts 7.7% · guest 92.3%9:00 · the hosts 7.7% · guest 92.3%12:00 · the hosts 20.2% · guest 79.8%12:00 · the hosts 20.2% · guest 79.8%15:00 · the hosts 3.8% · guest 96.2%15:00 · the hosts 3.8% · guest 96.2%18:00 · the hosts 17.1% · guest 82.9%18:00 · the hosts 17.1% · guest 82.9%21:00 · the hosts 21.3% · guest 78.7%21:00 · the hosts 21.3% · guest 78.7%24:00 · the hosts 18.1% · guest 81.9%24:00 · the hosts 18.1% · guest 81.9%27:00 · the hosts 15% · guest 85%27:00 · the hosts 15% · guest 85%30:00 · the hosts 9.1% · guest 90.9%30:00 · the hosts 9.1% · guest 90.9%33:00 · the hosts 4% · guest 96%33:00 · the hosts 4% · guest 96%36:00 · the hosts 10.8% · guest 89.2%36:00 · the hosts 10.8% · guest 89.2%39:00 · the hosts 15.2% · guest 84.8%39:00 · the hosts 15.2% · guest 84.8%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 40.8% · guest 59.2%45:00 · the hosts 40.8% · guest 59.2%48:00 · the hosts 56.7% · guest 43.3%48:00 · the hosts 56.7% · guest 43.3%51:00 · the hosts 26.1% · guest 73.9%51:00 · the hosts 26.1% · guest 73.9%54:00 · the hosts 34.3% · guest 65.7%54:00 · the hosts 34.3% · guest 65.7%57:00 · the hosts 32.4% · guest 67.6%57:00 · the hosts 32.4% · guest 67.6%1:00:00 · the hosts 37.4% · guest 62.6%1:00:00 · the hosts 37.4% · guest 62.6%1:03:00 · the hosts 28.7% · guest 71.3%1:03:00 · the hosts 28.7% · guest 71.3%1:06:00 · the hosts 5.8% · guest 94.2%1:06:00 · the hosts 5.8% · guest 94.2%1:09:00 · the hosts 7% · guest 93%1:09:00 · the hosts 7% · guest 93%1:12:00 · the hosts 22.2% · guest 77.8%1:12:00 · the hosts 22.2% · guest 77.8%1:15:00 · the hosts 37.3% · guest 62.7%1:15:00 · the hosts 37.3% · guest 62.7%1:18:00 · the hosts 7.1% · guest 92.9%1:18:00 · the hosts 7.1% · guest 92.9%1:21:00 · the hosts 9.7% · guest 90.3%1:21:00 · the hosts 9.7% · guest 90.3%1:24:00 · the hosts 26.5% · guest 73.5%1:24:00 · the hosts 26.5% · guest 73.5%
Sharpest disagreement ▶ 57:19 Shunyu rejecting complete intelligence-knowledge separation

Shunyu directly challenges Swyx's thesis on isolating intelligence from knowledge, arguing that intelligence inherently emerges through acquired knowledge rather than existing in isolation.

Hardest push from the hosts ▶ 47:21 Swyx pushing back on HCI and ACI equivalence

Swyx challenges the idea that interfaces designed for humans cleanly translate to AI agents, providing detailed counterexamples from structured output schemas and verbose compiler error loops.

Biggest teaching moment ▶ 51:42 Shunyu breaking down fundamental memory limits of ACI vs HCI

Shunyu educates the hosts on the foundational cognitive contrast between narrow human working memory (requiring sequential steps) and expansive LLM context windows (benefiting from parallel batch results).

The host holds their own ▶ 47:31 Swyx detailing specific schema field optimizations for LLM reasoning

Swyx demonstrates hands-on engineering expertise by showing how altering JSON key names to candidate topics triggers better chain-of-thought behavior in frontier models.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Welcome and Introductions with Harrison Chase and Shunyu Yao 6511 The hosts open with deep familiarity regarding Shunyu Yao's PhD defense and Harrison Chase's early inspiration from the ReAct paper. Shunyu provides historical context on text-adventure games like Zork I and the shift from RL to LLM-driven reasoning. The exchange is warm, collaborative, and appreciative.
ReAct's Legacy, Tool Calling, and the Modern Agent Loop 6522 Harrison and Shunyu discuss whether modern function-calling loops still reflect ReAct principles. Shunyu clarifies that inner monologues remain crucial when tool APIs diverge from pre-training data distributions, while Alessio asks how thinking steps are being internalized into base model weights.
Reflexion, Language-Based Feedback, and Agent Memory Architectures 7612 Shunyu explains Reflexion as substituting scalar RL rewards with rich linguistic feedback acting as linguistic gradient descent. The hosts and Harrison connect this to LangMem, Voyager, and cognitive science categorizations of semantic versus procedural memory.
Tree of Thoughts, Search Algorithms, and Prompting Philosophy 6422 The discussion covers Tree of Thoughts, search vs. reactive tasks, and the minimalist prompting philosophy. Shunyu rejects complex emotional prompt hacks in favor of plain human-to-human communication, prompting Harrison and Swyx to analyze user communication hurdles.
MCTS, The Benchmark Bottleneck, and Evaluation Realism 6621 Shunyu critiques the academic habit of applying over-complex methods to trivial benchmarks rather than building scalable, realistic tasks. He unpacks the trilemma of benchmark design between auto-gradability, realism, and scalability.
Interactive Coding, SWE-Agent, and Agent-Computer Interfaces (ACI) 7612 Shunyu details the progression from InterCode to SWE-bench and SWE-agent, highlighting that optimizing the Agent-Computer Interface (ACI) yields far higher returns than raw planning algorithms. The hosts draw parallels to Cognition's Devin and user interaction paradigms.
Optimizing Interfaces for LLMs: ACI vs. HCI and Machine Limits 8534 Swyx pushes back against the premise that HCI and ACI design completely overlap, citing his empirical experience with structured output field naming and verbose error traces. Shunyu contrasts human working memory constraints with machine context limits, while Harrison argues for allowing agent interfaces to diverge early.
Separating Intelligence from Knowledge and Training Trajectory Data 7633 Swyx poses whether intelligence can be cleanly separated from knowledge, pointing to on-device LoRA adapters. Shunyu pushes back, arguing from historical AI and Hinton's perspective that intelligence emerges with knowledge, though framing knowledge as a cache for intelligence.
CoALA Framework, Action Spaces, and Stateful Orchestration 7512 Shunyu walks through the CoALA framework across memory, action space, and decision-making. Harrison and Swyx integrate LangGraph's stateful cross-thread persistence model into the architecture, debating whether developers or agents should select memory tooling.
Tau-Bench, Customer Simulation, and Practical Enterprise Agents 6511 Alessio and Shunyu discuss Tau-Bench and customer service agent simulation where LLMs model realistic user behavior under information asymmetry. Harrison answers Shunyu's question about enterprise applications, pointing to customer support, SDR data enrichment, and spreadsheet agents.
LangGraph Studio, Agent IDEs, and Future Tooling 6422 Shunyu questions whether low-code tooling is truly ready for non-programmers, and Harrison clarifies that LangGraph Studio acts as an IDE to bridge developer architecture with PM prompt refinement. The group wraps up with reflections on future agent UX.

Statements from this episode (30)

Assertion Not checkable as stated
Shunyu Yao says Ilya Sutskever claimed GPT-1 had solved language
“Back in OpenAI, they did this GPT-ONE together, and Ilya just said, Karthik, you should stay, because we just solved the language.”
Shunyu Yao Sep 27, 2024 ▶ 2:12
Assertion Not checkable as stated
Shunyu Yao built the ReAct prototype before Chain-of-Thought existed
“The prototype I think was around November of 2021. So that's even before like chain of thought or whatever came up.”
Shunyu Yao Sep 27, 2024 ▶ 8:13
Opinion
Yao: Text adventure games remain very hard even for GPT-4
“Like those texting are just too hard. I think today it's still very hard. Like if you used to be before to solve it, it's still very hard.”
Shunyu Yao Sep 27, 2024 ▶ 8:29
Insight
Harrison Chase says production AI agents rely on three main defaults
“And there's such a long tail of other ones, but in practice, like, when people go to production, they generally have their own tools, or maybe one of those three, maybe some other ones, but, like, very, very few other ones.”
Harrison Chase Sep 27, 2024 ▶ 9:49
Insight
Yao: Pairing reasoning with tool use is essential for unfamiliar tools
“And I think the second contribution is this idea of what people call like inner monologue or thinking or reasoning or whatever to be paired with tool use. I think that's still not trivial because if you look at the default function calling or whatever, like th…”
Shunyu Yao Sep 27, 2024 ▶ 11:38
Assertion Supported
Harrison Chase says OpenAI recommends adding a thought field to tool schemas
“I think open AI even recommended, like when you're doing tool calling, it's sometimes helpful to put like a thought field in the tool along with all the actual acquired arguments and then have that one first. So it fills out that first and then, and that's, th…”
Harrison Chase Sep 27, 2024 ▶ 12:10
Insight
Yao: Reflexion replaces scalar RL rewards with verbal gradient descent
“I think one way to think of reflection is that the traditional idea of reinforcement learning is you have a scalar reward, and then you somehow back propagate the signal of the scalar reward. To the rest of your neural network through whatever algorithm, like …”
Shunyu Yao Sep 27, 2024 ▶ 15:35
Assertion Supported
Chase: LangChain and other frameworks lack off-the-shelf Reflexion implementations
“I don't think we have like an off the shelf kind of like implementation of reflection and kind of like the general sense. I think the concepts like absolutely we see used in different kind of like specific cognitive architectures, but I don't think we have one…”
Harrison Chase Sep 27, 2024 ▶ 17:05
Insight
Yao: Evaluator quality is the key bottleneck for agent self-reflection
“I think a key bottleneck is the evaluator, right? Basically you need to have a good sense of the signal. So for example, like if you are trying to do a very hard reasoning task, say mathematics, For example, and you don't have any tools, right? It's operating …”
Shunyu Yao Sep 27, 2024 ▶ 17:56
Opinion
Harrison Chase says ReAct is the most popular agent prompting framework
“I would say like reacts probably like the most popular. I think there's aspects of reflection that Get used. Tree of thought, probably like the least so.”
Harrison Chase Sep 27, 2024 ▶ 25:23
Insight
Shunyu Yao advises developers to default to minimalist prompting for AI agents
“And I think in terms of the actual prompting method to use for a particular problem, I'm I think we should all be in the minimum list kind of camp, right? You should try the minimum thing and see if it works and if it doesn't work and there's absolute reason t…”
Shunyu Yao Sep 27, 2024 ▶ 27:28
Insight
Shunyu Yao argues modern LLMs make prompt engineering tricks obsolete
“I feel like in some sense, I feel like prompt engineering, even it's like a slightly negative word at the time, because it refers to all those kind of weird tricks that you have to apply. But I think we don't have to do that anymore. Like given today's progres…”
Shunyu Yao Sep 27, 2024 ▶ 29:31
Insight
Yao: Lack of realistic benchmarks is AI's primary bottleneck
“So I think right now the problem is not even that we don't have good methodologies, it's more about we don't have good tasks.”
Shunyu Yao Sep 27, 2024 ▶ 31:08
Opinion
Shunyu Yao says academic AI research overcomplicates methods on simplistic tasks
“And I think in general, what people do in academia that I think is not good is they choose a very simple task, like Alford, and then they apply overly complex methods and to show the improved two percent I think like you should probably match, you know, the le…”
Shunyu Yao Sep 27, 2024 ▶ 31:32
Insight
Yao: SWE-bench succeeded by balancing auto-grading, practicality, and scalability
“And I think part of the reason that Sweetbench is so popular now is it kind of hits the balance between these three dimensions, right? Easy to evaluate and being actually practical and being scalable.”
Shunyu Yao Sep 27, 2024 ▶ 34:40
Insight
Shunyu Yao believes coding is the best application for AI agents
“Obviously coding is the best application for agents because it's all the gradable. It's super important. You can make everything like API or code action, right?”
Shunyu Yao Sep 27, 2024 ▶ 37:04
Opinion
Shawn Wang says Devin's breakthrough was Agent-Computer Interfaces, not advanced planning
“The planner is like actually pretty simple, but ACI. That they book through on.”
Shawn Wang Sep 27, 2024 ▶ 45:28
Insight
Yao: Reliable tool design accounts for 90% of agent performance
“I think making the tool good and reliable is probably like 90% of the whole agent. Once the tool is actually good, then the agent design can be much, much simpler. On the other hand, if the tool is bad, then no matter how much you put into the agent design pla…”
Shunyu Yao Sep 27, 2024 ▶ 45:35
Assertion Supported
Harrison Chase says TypeScript yields better LLM tool-calling performance than JSON
“I saw some paper that used TypeScript notation instead of JSON notation for tool calling and it got a lot better performance.”
Harrison Chase Sep 27, 2024 ▶ 46:42
Insight
Shunyu Yao says agent interfaces should leverage large context over temporal steps
“If you look at find or whatever terminal command, you know, you can only look at one thing at a time, or that's because we have a very small working memory. You can only deal with one thing at a time. You can only look at one paragraph of text at the same time…”
Shunyu Yao Sep 27, 2024 ▶ 51:50
Insight
Yao: Human cognition should serve as an AI reference point, not a blueprint
“I don't think we should copy exactly what's going on with human all the way, but I think it's good to have a reference point because this is a working example of how intelligence works.”
Shunyu Yao Sep 27, 2024 ▶ 54:11
Opinion
Yao: Improving AI agents requires better data, not architectural changes
“I think it's data. I think it's data because like changing architecture now is too hard and we don't have a good, better alternative solution now. I think it's mostly about data and agent data is obviously hard because People just write down the final result o…”
Shunyu Yao Sep 27, 2024 ▶ 58:54
Prediction Not checkable as stated
Shunyu Yao predicts training models on human computer trajectories achieves AGI
“The simplest way to achieve AGI is literally just record the re-actuatory of every human being and just put them together, you know, like what do you have thought about? What do you have done? Let's say on the computer, right? Imagine like solid experiment. Li…”
Shunyu Yao Sep 27, 2024 ▶ 1:02:07
Opinion
Chase: Few-shot prompting works better than detailed instructions for agent trajectories
“I'm pretty bullish on it, to be honest, for a few reasons. Like, one, I think it can maybe help for more complex things, but then also, two, like, it's a form of prompting, and prompting is just Communicating with the model what you want it to do. And sometime…”
Harrison Chase Sep 27, 2024 ▶ 1:03:16
Insight
Harrison Chase says developers should guide agent planning explicitly in code
“Sometimes I say that like the LLMs aren't [4137] Great at planning yet. [4138] So we can help them plan by telling them how to plan and code. [4140] Cause that's very explicit and that's a good way of communicating how they should plan and stuff like that.”
Harrison Chase Sep 27, 2024 ▶ 1:08:52
Disclosure
Harrison Chase admits LangChain's memory service lacked product-market fit
“The memory service we launched, I don't think really found product market fit.”
Harrison Chase Sep 27, 2024 ▶ 1:12:12
Disclosure
Chase: LangGraph is adding cross-thread state persistence for agent memory
“Right now they're all persistent for a single thread. [4392] We're going to add the ability to persist that between threads. [4395] So then if you basically want to scope a memory to a user ID or to an assistant or to an organization, then you can do that.”
Harrison Chase Sep 27, 2024 ▶ 1:13:10
Insight
Yao: Enterprise customer support AI requires 99% reliability over simple tasks, not search
“It's very different from coding or web agent or whatever people are doing, because it's more about how can you do simple things reliably It's not about, you know, can you sample a hundred times and you find one good mass proof or kill solution. It's more about…”
Shunyu Yao Sep 27, 2024 ▶ 1:17:44
Opinion
Chase: Customer support is a clear production success area for AI agents
“One big area where there's clearly been success is in customer support. both companies doing that as a service, but also larger enterprises doing that and building that functionality inside.”
Harrison Chase Sep 27, 2024 ▶ 1:20:38
Opinion
Harrison Chase says coding AI agents are not yet a proven success
“There's a bunch of people doing coding stuff. We've already talked about that. I think that's a little bit, I wouldn't say that's a success yet, but there's a lot of excitement and stuff there.”
Harrison Chase Sep 27, 2024 ▶ 1:20:50
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.