Mar 28, 2024 · 27m · no-priors

No Priors Ep. 57 | With LangChain CEO and Co-Founder Harrison Chase

Harrison Chase · 19m spoken Sarah Guo · 3m spoken Elad Gil · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, LangChain CEO Harrison Chase discusses the evolution of LLM application development, exploring the transition from linear chains to cyclical agent graphs, practical memory architectures, and the ongoing necessity of RAG and prompt optimization in production AI systems.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.6% of the talking time here. How this is scored →

The hosts as informed peer 5.0 Guest teaching 3.7 Guest disagreement 1.4 The hosts pushing back 1.1
05100:0010:0020:001:42–4:25 · The hosts as informed peer 4/10 Framework Evolution: From Simple Chains to LangGraph Sarah asks how LangChain manages architectural stability versus rapid ecosystem shifts. Harrison explains how the framework evolved from simple chains to cyclical state graphs like LangGraph to support production agent requirements.4:26–7:19 · The hosts as informed peer 4/10 Overcoming Planning and UX Bottlenecks in AI Agents Elad asks what foundational components are still missing to make autonomous agents performant. Harrison categorizes the current blockers into UX ambiguity, underlying LLM planning deficiencies, and builder evaluation workflows.7:20–10:00 · The hosts as informed peer 4/10 Categorizing Memory: Procedural Flywheels and Personalization Elad prompts a discussion on how memory should be structured in agentic systems. Harrison provides a clear taxonomy separating system-level procedural memory (using tool-use flywheels) from personalized episodic memory.10:00–14:06 · The hosts as informed peer 6/10 Production AI Architectures and High-Impact Agent Domains Sarah highlights recent breakthroughs in multi-step RAG, tree search sampling, and coding agents like Cognition's Devin. Harrison agrees and elaborates on how controlled state machines and advanced query analysis are replacing naive loops.14:07–19:53 · The hosts as informed peer 7/10 Model Interoperability, Switching Costs, and Context Windows vs. RAG Sarah and Elad probe model switching costs, prompt portability, and whether expanding million-token context windows render RAG obsolete. Harrison details why RAG remains vital for multi-needle synthesis, iterative environment feedback, and cost management at scale.19:53–22:56 · The hosts as informed peer 5/10 Evaluating Fine-Tuning Adoption and Open-Source Reasoning Models Elad inquires why fine-tuning is rarely deployed in production compared to prompt engineering. Harrison validates this observation and offers a contrarian take that open-source reasoning models still fail to live up to online hype compared to leading frontier models.22:56–27:05 · The hosts as informed peer 5/10 Continual Learning, Optimization Loops, and Future Frontiers Sarah asks about unexplored application frontiers, prompting Harrison to outline continual learning through automated few-shot optimization loops, drawing comparisons to Stanford's DSPy framework alongside lighthearted banter on its pronunciation.1:42–4:25 · Guest teaching 3/10 Framework Evolution: From Simple Chains to LangGraph Sarah asks how LangChain manages architectural stability versus rapid ecosystem shifts. Harrison explains how the framework evolved from simple chains to cyclical state graphs like LangGraph to support production agent requirements.4:26–7:19 · Guest teaching 4/10 Overcoming Planning and UX Bottlenecks in AI Agents Elad asks what foundational components are still missing to make autonomous agents performant. Harrison categorizes the current blockers into UX ambiguity, underlying LLM planning deficiencies, and builder evaluation workflows.7:20–10:00 · Guest teaching 5/10 Categorizing Memory: Procedural Flywheels and Personalization Elad prompts a discussion on how memory should be structured in agentic systems. Harrison provides a clear taxonomy separating system-level procedural memory (using tool-use flywheels) from personalized episodic memory.10:00–14:06 · Guest teaching 3/10 Production AI Architectures and High-Impact Agent Domains Sarah highlights recent breakthroughs in multi-step RAG, tree search sampling, and coding agents like Cognition's Devin. Harrison agrees and elaborates on how controlled state machines and advanced query analysis are replacing naive loops.14:07–19:53 · Guest teaching 4/10 Model Interoperability, Switching Costs, and Context Windows vs. RAG Sarah and Elad probe model switching costs, prompt portability, and whether expanding million-token context windows render RAG obsolete. Harrison details why RAG remains vital for multi-needle synthesis, iterative environment feedback, and cost management at scale.19:53–22:56 · Guest teaching 4/10 Evaluating Fine-Tuning Adoption and Open-Source Reasoning Models Elad inquires why fine-tuning is rarely deployed in production compared to prompt engineering. Harrison validates this observation and offers a contrarian take that open-source reasoning models still fail to live up to online hype compared to leading frontier models.22:56–27:05 · Guest teaching 3/10 Continual Learning, Optimization Loops, and Future Frontiers Sarah asks about unexplored application frontiers, prompting Harrison to outline continual learning through automated few-shot optimization loops, drawing comparisons to Stanford's DSPy framework alongside lighthearted banter on its pronunciation.1:42–4:25 · Guest disagreement 1/10 Framework Evolution: From Simple Chains to LangGraph Sarah asks how LangChain manages architectural stability versus rapid ecosystem shifts. Harrison explains how the framework evolved from simple chains to cyclical state graphs like LangGraph to support production agent requirements.4:26–7:19 · Guest disagreement 1/10 Overcoming Planning and UX Bottlenecks in AI Agents Elad asks what foundational components are still missing to make autonomous agents performant. Harrison categorizes the current blockers into UX ambiguity, underlying LLM planning deficiencies, and builder evaluation workflows.7:20–10:00 · Guest disagreement 1/10 Categorizing Memory: Procedural Flywheels and Personalization Elad prompts a discussion on how memory should be structured in agentic systems. Harrison provides a clear taxonomy separating system-level procedural memory (using tool-use flywheels) from personalized episodic memory.10:00–14:06 · Guest disagreement 1/10 Production AI Architectures and High-Impact Agent Domains Sarah highlights recent breakthroughs in multi-step RAG, tree search sampling, and coding agents like Cognition's Devin. Harrison agrees and elaborates on how controlled state machines and advanced query analysis are replacing naive loops.14:07–19:53 · Guest disagreement 2/10 Model Interoperability, Switching Costs, and Context Windows vs. RAG Sarah and Elad probe model switching costs, prompt portability, and whether expanding million-token context windows render RAG obsolete. Harrison details why RAG remains vital for multi-needle synthesis, iterative environment feedback, and cost management at scale.19:53–22:56 · Guest disagreement 3/10 Evaluating Fine-Tuning Adoption and Open-Source Reasoning Models Elad inquires why fine-tuning is rarely deployed in production compared to prompt engineering. Harrison validates this observation and offers a contrarian take that open-source reasoning models still fail to live up to online hype compared to leading frontier models.22:56–27:05 · Guest disagreement 1/10 Continual Learning, Optimization Loops, and Future Frontiers Sarah asks about unexplored application frontiers, prompting Harrison to outline continual learning through automated few-shot optimization loops, drawing comparisons to Stanford's DSPy framework alongside lighthearted banter on its pronunciation.1:42–4:25 · The hosts pushing back 1/10 Framework Evolution: From Simple Chains to LangGraph Sarah asks how LangChain manages architectural stability versus rapid ecosystem shifts. Harrison explains how the framework evolved from simple chains to cyclical state graphs like LangGraph to support production agent requirements.4:26–7:19 · The hosts pushing back 1/10 Overcoming Planning and UX Bottlenecks in AI Agents Elad asks what foundational components are still missing to make autonomous agents performant. Harrison categorizes the current blockers into UX ambiguity, underlying LLM planning deficiencies, and builder evaluation workflows.7:20–10:00 · The hosts pushing back 1/10 Categorizing Memory: Procedural Flywheels and Personalization Elad prompts a discussion on how memory should be structured in agentic systems. Harrison provides a clear taxonomy separating system-level procedural memory (using tool-use flywheels) from personalized episodic memory.10:00–14:06 · The hosts pushing back 1/10 Production AI Architectures and High-Impact Agent Domains Sarah highlights recent breakthroughs in multi-step RAG, tree search sampling, and coding agents like Cognition's Devin. Harrison agrees and elaborates on how controlled state machines and advanced query analysis are replacing naive loops.14:07–19:53 · The hosts pushing back 2/10 Model Interoperability, Switching Costs, and Context Windows vs. RAG Sarah and Elad probe model switching costs, prompt portability, and whether expanding million-token context windows render RAG obsolete. Harrison details why RAG remains vital for multi-needle synthesis, iterative environment feedback, and cost management at scale.19:53–22:56 · The hosts pushing back 1/10 Evaluating Fine-Tuning Adoption and Open-Source Reasoning Models Elad inquires why fine-tuning is rarely deployed in production compared to prompt engineering. Harrison validates this observation and offers a contrarian take that open-source reasoning models still fail to live up to online hype compared to leading frontier models.22:56–27:05 · The hosts pushing back 1/10 Continual Learning, Optimization Loops, and Future Frontiers Sarah asks about unexplored application frontiers, prompting Harrison to outline continual learning through automated few-shot optimization loops, drawing comparisons to Stanford's DSPy framework alongside lighthearted banter on its pronunciation.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 30.7% · guest 69.3%0:00 · the hosts 30.7% · guest 69.3%3:00 · the hosts 9.9% · guest 90.1%3:00 · the hosts 9.9% · guest 90.1%6:00 · the hosts 8.4% · guest 91.6%6:00 · the hosts 8.4% · guest 91.6%9:00 · the hosts 16.4% · guest 83.6%9:00 · the hosts 16.4% · guest 83.6%12:00 · the hosts 41.8% · guest 58.2%12:00 · the hosts 41.8% · guest 58.2%15:00 · the hosts 23.4% · guest 76.6%15:00 · the hosts 23.4% · guest 76.6%18:00 · the hosts 16.9% · guest 83.1%18:00 · the hosts 16.9% · guest 83.1%21:00 · the hosts 28.1% · guest 71.9%21:00 · the hosts 28.1% · guest 71.9%24:00 · the hosts 10.2% · guest 89.8%24:00 · the hosts 10.2% · guest 89.8%27:00 · the hosts 76.8% · guest 23.2%27:00 · the hosts 76.8% · guest 23.2%
Sharpest disagreement ▶ 22:25 Pushing back against open-source reasoning hype

Harrison delivers a contrarian 'hot take', bluntly stating that open-source models lag significantly behind GPT-4 and Claude 3 in reasoning and do not live up to Twitter excitement.

Hardest push from the hosts ▶ 14:06 Challenging model portability assumptions

Sarah questions the popular industry narrative that developers can seamlessly switch backends between model providers given prompt sensitivities and behavioral differences.

Biggest teaching moment ▶ 18:05 Reframing long-context benchmarks vs. real-world RAG

Harrison educates the audience on why standard needle-in-a-haystack evaluations are misleading, explaining that real RAG requires multi-document synthesis and iterative environmental reasoning.

The host holds their own ▶ 12:22 Sarah detailing tree search and efficient sampling

Sarah demonstrates strong technical authority by analyzing recent agent design patterns combining tree search, efficient sampling, and practical execution loops.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Framework Evolution: From Simple Chains to LangGraph 4311 Sarah asks how LangChain manages architectural stability versus rapid ecosystem shifts. Harrison explains how the framework evolved from simple chains to cyclical state graphs like LangGraph to support production agent requirements.
Overcoming Planning and UX Bottlenecks in AI Agents 4411 Elad asks what foundational components are still missing to make autonomous agents performant. Harrison categorizes the current blockers into UX ambiguity, underlying LLM planning deficiencies, and builder evaluation workflows.
Categorizing Memory: Procedural Flywheels and Personalization 4511 Elad prompts a discussion on how memory should be structured in agentic systems. Harrison provides a clear taxonomy separating system-level procedural memory (using tool-use flywheels) from personalized episodic memory.
Production AI Architectures and High-Impact Agent Domains 6311 Sarah highlights recent breakthroughs in multi-step RAG, tree search sampling, and coding agents like Cognition's Devin. Harrison agrees and elaborates on how controlled state machines and advanced query analysis are replacing naive loops.
Model Interoperability, Switching Costs, and Context Windows vs. RAG 7422 Sarah and Elad probe model switching costs, prompt portability, and whether expanding million-token context windows render RAG obsolete. Harrison details why RAG remains vital for multi-needle synthesis, iterative environment feedback, and cost management at scale.
Evaluating Fine-Tuning Adoption and Open-Source Reasoning Models 5431 Elad inquires why fine-tuning is rarely deployed in production compared to prompt engineering. Harrison validates this observation and offers a contrarian take that open-source reasoning models still fail to live up to online hype compared to leading frontier models.
Continual Learning, Optimization Loops, and Future Frontiers 5311 Sarah asks about unexplored application frontiers, prompting Harrison to outline continual learning through automated few-shot optimization loops, drawing comparisons to Stanford's DSPy framework alongside lighthearted banter on its pronunciation.

Statements from this episode (14)

Insight
Chase: AI agents fundamentally require cyclical graphs, not linear architectures
“So, you know, all these agents are basically running an LLM in a loop. You need cycles and so lane graph helps with that.”
Harrison Chase Mar 28, 2024 ▶ 3:21
Insight
Chase: Working AI agents require hardcoded domain structure, not LLM autonomy
“I think when we see people building agents that work right now, it's often breaking it down into a bunch of smaller components and kind of like imparting their domain knowledge about how information should Flow through these components. Because I think the ele…”
Harrison Chase Mar 28, 2024 ▶ 5:28
Insight
Chase: Few-shot prompting effectively builds procedural memory for AI agents
“So on the procedural side, I think the main thing that we see people doing and that we think is pretty effective is few-shot prompting and maybe fine-tuning for how to use for how to use tools, because that's basically what it comes down to. What's the right w…”
Harrison Chase Mar 28, 2024 ▶ 8:22
Opinion
Chase: Passive background insight extraction will drive AI personalization memory
“I also think one thing that I'm bullish on is a more kind of, like passive background process that kind of looks at conversations and almost, like, extracts insights. And then you can use those insights in kind of, like, future conversations.”
Harrison Chase Mar 28, 2024 ▶ 9:26
Opinion
Chase: AI agent memory is extremely nascent and lacks interesting developments
“I feel it's like a field that's just, like, super, super nascent. Like, I don't, I actually am underwhelmed at the amount of, like, really interesting stuff that's going on there.”
Harrison Chase Mar 28, 2024 ▶ 9:44
Insight
Chase: Successful production AI agents operate as controlled state machines
“And I think the things that we see making it into production and informed a lot of the development of laying graph is, or is something in the middle where it's like this controlled state machine type thing.”
Harrison Chase Mar 28, 2024 ▶ 12:01
Insight
Chase: Code execution gives AI coding agents a direct feedback loop
“But those type like coding, coding problems in general, we see a lot of people working on. I think there's a really nice feedback loop that you can get by just like executing the code and seeing if it works.”
Harrison Chase Mar 28, 2024 ▶ 13:13
Prediction Not checkable as stated
Chase: LLM prompts will likely converge as models become more intelligent
“I do think the prompts will probably start to converge in the sense that if you think the models are getting more and more intelligent than like, hopefully these small idiosyncratic sees don't matter as much.”
Harrison Chase Mar 28, 2024 ▶ 15:05
Prediction Not checkable as stated
Chase: Long context windows will not replace chaining and AI agents
“There are also things where it requires iterations. You need to like decide what to do, interact with the environment, get that back. So this whole idea of chaining and agents, I don't like, That's less around context windows and more around interacting with t…”
Harrison Chase Mar 28, 2024 ▶ 18:07
Insight
Chase: Needle-in-a-haystack benchmarks fail to represent real RAG reasoning
“That, that actually really doesn't reflect a lot of RAG use cases in, in my opinion, because like that's the needle in the haystack is like, okay, given this long context, can I find a single information point? But oftentimes RAG is about seeing multiple infor…”
Harrison Chase Mar 28, 2024 ▶ 18:44
Assertion Not checkable as stated
Chase: Developers only implement model fine-tuning after reaching critical scale
“We see people experimenting with it. I think the only real place where they're doing it is when they've reached like really critical scale which I still don't think is that many applications to date.”
Harrison Chase Mar 28, 2024 ▶ 20:24
Opinion
Chase: Open-source models still lag behind Claude 3 and GPT-4
“Like there's, I think we see increasingly interest in open source, but the reasoning abilities are still just like lagging behind Cloud three or GPT four. And I think like for a lot of the applications that it kind of, it probably depends on the types of appli…”
Harrison Chase Mar 28, 2024 ▶ 22:06
Opinion
Chase: New AI startups should build applications leveraging long-term memory
“If I wasn't doing LinkedIn, if I was starting a company right now, I'd probably start something at the application layer, and it would probably be something that really takes advantage of, like, long-term memory.”
Harrison Chase Mar 28, 2024 ▶ 23:34
Insight
Chase: Few-shot example datasets are faster and cheaper than model fine-tuning
“Building up few shot example data sets and really using those. I think it's much faster and cheaper than fine tuning models. It's easier to do than trying to like. Programmatically change the prompt in some way.”
Harrison Chase Mar 28, 2024 ▶ 24:23
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.