Oct 21, 2023 · 1h 12m · latent-space

Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue

Kanjun Qiu · 48m spoken Shawn Wang · 11m spoken Alessio Fanelli · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Imbue CEO Kanjun Qiu examines why current autonomous AI agents encounter severe reliability bottlenecks, detailing how foundation models optimized for explicit natural language reasoning, synthetic code data, and robust software abstractions will transform personal computing.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 25.3% of the talking time here. How this is scored →

The hosts as informed peer 5.4 Guest teaching 3.8 Guest disagreement 1.7 The hosts pushing back 2.8
05100:0015:0030:0045:001:00:002:57–7:13 · The hosts as informed peer 5/10 Early Startups and Lessons from Sorceress AI Swyx draws a sharp analogy between Sorceress's recruiting business and the dating app business model of endogenous churn. Kanjun validates the comparison and explains how trust barriers informed Imbue's interface approach.7:13–11:28 · The hosts as informed peer 4/10 Founding Generally Intelligent and Free Intellectual Energy Alessio asks about starting Generally Intelligent in 2021 before agents were mainstream. Kanjun explains her foundational thesis on free intellectual energy analogous to the Industrial Revolution and early self-supervised learning breakthroughs.11:29–15:24 · The hosts as informed peer 5/10 Reasoning Models and Personal Computing Vision Alessio and Swyx explore Imbue's focus on reasoning models and ask how they navigate between a pure research lab and a product company. Kanjun draws parallels to Apple's long multi-touch development cycles.15:24–17:52 · The hosts as informed peer 4/10 Imbue's Internal Culture and Engineering Structure Alessio queries how Imbue manages its unique dual culture after raising 200 million dollars. Kanjun outlines her social process philosophy of treating employees as creative agents rather than disposable corporate assets.17:52–22:31 · The hosts as informed peer 7/10 Optimizing Pre-Training and Synthetic Data Generation Swyx brings strong domain depth regarding compute versus data spending in foundation models, explicitly questioning whether synthetic model data causes distribution collapse. Kanjun explains that Imbue intentionally wants a spiky reasoning distribution rather than matching web text.22:32–27:36 · The hosts as informed peer 5/10 Limits of Reinforcement Learning and Avalon Simulation Alessio asks about the Avalon simulation benchmark, playfully prompting Kanjun to correct a quote. Kanjun delivers a thorough technical breakdown of why pure reinforcement learning without curricula fails at planning and why reasoning must happen in language.27:37–32:13 · The hosts as informed peer 5/10 Current Agent Limitations and Architectural Abstractions Swyx presses on whether simple single-task agents face an asymptote ceiling. Kanjun readily agrees, acknowledging that current agent abstractions are severely leaky and that pile-of-hacks solutions fail production reliability bars.32:14–37:53 · The hosts as informed peer 6/10 Defining Reasoning and Natural Language Programming Alessio references George Hotz's claim that sufficient compute will automatically emerge reasoning from human data. Kanjun explicitly rejects that premise, arguing that human reasoning strategies must be generated intentionally and that code serves as a structured curriculum.37:54–46:54 · The hosts as informed peer 7/10 Interface Paradigms, Agent Protocols, and Evaluation The hosts and Kanjun discuss skeuomorphic chat interfaces, historical computing transitions, agent communication protocols, and recent research like MetaGPT and Voyager. Both sides trade specific literature references comfortably.46:55–54:14 · The hosts as informed peer 7/10 Context Memory Limitations and Imbue's Positioning Swyx pushes hard with direct competitive questions, asking why RAG is insufficient for complex tasks and confronting Kanjun on why candidates should pick Imbue over Adept. Kanjun provides a structured defense centered on non-leaky abstractions and dogfooding.54:16–1:07:07 · The hosts as informed peer 6/10 Curiosity, Online Learning, and Cultivating Scenius Alessio and Swyx explore Kanjun's philosophical website questions and deep involvement in building hacker house ecosystems like The Archive and South Park Commons. Kanjun and Swyx discuss the concept of scenius and idea flux.1:07:07–1:12:28 · The hosts as informed peer 4/10 Lightning Round, AI Governance, and Safety Takeaways In the lightning round, Kanjun deflects Swyx's hypothetical alternative company question by asserting Imbue is the only thing worth building. She concludes with Imbue's multi-pillar approach to AI safety and empirical policy work with the Department of Commerce.2:57–7:13 · Guest teaching 3/10 Early Startups and Lessons from Sorceress AI Swyx draws a sharp analogy between Sorceress's recruiting business and the dating app business model of endogenous churn. Kanjun validates the comparison and explains how trust barriers informed Imbue's interface approach.7:13–11:28 · Guest teaching 4/10 Founding Generally Intelligent and Free Intellectual Energy Alessio asks about starting Generally Intelligent in 2021 before agents were mainstream. Kanjun explains her foundational thesis on free intellectual energy analogous to the Industrial Revolution and early self-supervised learning breakthroughs.11:29–15:24 · Guest teaching 3/10 Reasoning Models and Personal Computing Vision Alessio and Swyx explore Imbue's focus on reasoning models and ask how they navigate between a pure research lab and a product company. Kanjun draws parallels to Apple's long multi-touch development cycles.15:24–17:52 · Guest teaching 3/10 Imbue's Internal Culture and Engineering Structure Alessio queries how Imbue manages its unique dual culture after raising 200 million dollars. Kanjun outlines her social process philosophy of treating employees as creative agents rather than disposable corporate assets.17:52–22:31 · Guest teaching 4/10 Optimizing Pre-Training and Synthetic Data Generation Swyx brings strong domain depth regarding compute versus data spending in foundation models, explicitly questioning whether synthetic model data causes distribution collapse. Kanjun explains that Imbue intentionally wants a spiky reasoning distribution rather than matching web text.22:32–27:36 · Guest teaching 6/10 Limits of Reinforcement Learning and Avalon Simulation Alessio asks about the Avalon simulation benchmark, playfully prompting Kanjun to correct a quote. Kanjun delivers a thorough technical breakdown of why pure reinforcement learning without curricula fails at planning and why reasoning must happen in language.27:37–32:13 · Guest teaching 4/10 Current Agent Limitations and Architectural Abstractions Swyx presses on whether simple single-task agents face an asymptote ceiling. Kanjun readily agrees, acknowledging that current agent abstractions are severely leaky and that pile-of-hacks solutions fail production reliability bars.32:14–37:53 · Guest teaching 5/10 Defining Reasoning and Natural Language Programming Alessio references George Hotz's claim that sufficient compute will automatically emerge reasoning from human data. Kanjun explicitly rejects that premise, arguing that human reasoning strategies must be generated intentionally and that code serves as a structured curriculum.37:54–46:54 · Guest teaching 3/10 Interface Paradigms, Agent Protocols, and Evaluation The hosts and Kanjun discuss skeuomorphic chat interfaces, historical computing transitions, agent communication protocols, and recent research like MetaGPT and Voyager. Both sides trade specific literature references comfortably.46:55–54:14 · Guest teaching 4/10 Context Memory Limitations and Imbue's Positioning Swyx pushes hard with direct competitive questions, asking why RAG is insufficient for complex tasks and confronting Kanjun on why candidates should pick Imbue over Adept. Kanjun provides a structured defense centered on non-leaky abstractions and dogfooding.54:16–1:07:07 · Guest teaching 3/10 Curiosity, Online Learning, and Cultivating Scenius Alessio and Swyx explore Kanjun's philosophical website questions and deep involvement in building hacker house ecosystems like The Archive and South Park Commons. Kanjun and Swyx discuss the concept of scenius and idea flux.1:07:07–1:12:28 · Guest teaching 3/10 Lightning Round, AI Governance, and Safety Takeaways In the lightning round, Kanjun deflects Swyx's hypothetical alternative company question by asserting Imbue is the only thing worth building. She concludes with Imbue's multi-pillar approach to AI safety and empirical policy work with the Department of Commerce.2:57–7:13 · Guest disagreement 1/10 Early Startups and Lessons from Sorceress AI Swyx draws a sharp analogy between Sorceress's recruiting business and the dating app business model of endogenous churn. Kanjun validates the comparison and explains how trust barriers informed Imbue's interface approach.7:13–11:28 · Guest disagreement 1/10 Founding Generally Intelligent and Free Intellectual Energy Alessio asks about starting Generally Intelligent in 2021 before agents were mainstream. Kanjun explains her foundational thesis on free intellectual energy analogous to the Industrial Revolution and early self-supervised learning breakthroughs.11:29–15:24 · Guest disagreement 1/10 Reasoning Models and Personal Computing Vision Alessio and Swyx explore Imbue's focus on reasoning models and ask how they navigate between a pure research lab and a product company. Kanjun draws parallels to Apple's long multi-touch development cycles.15:24–17:52 · Guest disagreement 1/10 Imbue's Internal Culture and Engineering Structure Alessio queries how Imbue manages its unique dual culture after raising 200 million dollars. Kanjun outlines her social process philosophy of treating employees as creative agents rather than disposable corporate assets.17:52–22:31 · Guest disagreement 2/10 Optimizing Pre-Training and Synthetic Data Generation Swyx brings strong domain depth regarding compute versus data spending in foundation models, explicitly questioning whether synthetic model data causes distribution collapse. Kanjun explains that Imbue intentionally wants a spiky reasoning distribution rather than matching web text.22:32–27:36 · Guest disagreement 2/10 Limits of Reinforcement Learning and Avalon Simulation Alessio asks about the Avalon simulation benchmark, playfully prompting Kanjun to correct a quote. Kanjun delivers a thorough technical breakdown of why pure reinforcement learning without curricula fails at planning and why reasoning must happen in language.27:37–32:13 · Guest disagreement 1/10 Current Agent Limitations and Architectural Abstractions Swyx presses on whether simple single-task agents face an asymptote ceiling. Kanjun readily agrees, acknowledging that current agent abstractions are severely leaky and that pile-of-hacks solutions fail production reliability bars.32:14–37:53 · Guest disagreement 4/10 Defining Reasoning and Natural Language Programming Alessio references George Hotz's claim that sufficient compute will automatically emerge reasoning from human data. Kanjun explicitly rejects that premise, arguing that human reasoning strategies must be generated intentionally and that code serves as a structured curriculum.37:54–46:54 · Guest disagreement 2/10 Interface Paradigms, Agent Protocols, and Evaluation The hosts and Kanjun discuss skeuomorphic chat interfaces, historical computing transitions, agent communication protocols, and recent research like MetaGPT and Voyager. Both sides trade specific literature references comfortably.46:55–54:14 · Guest disagreement 2/10 Context Memory Limitations and Imbue's Positioning Swyx pushes hard with direct competitive questions, asking why RAG is insufficient for complex tasks and confronting Kanjun on why candidates should pick Imbue over Adept. Kanjun provides a structured defense centered on non-leaky abstractions and dogfooding.54:16–1:07:07 · Guest disagreement 1/10 Curiosity, Online Learning, and Cultivating Scenius Alessio and Swyx explore Kanjun's philosophical website questions and deep involvement in building hacker house ecosystems like The Archive and South Park Commons. Kanjun and Swyx discuss the concept of scenius and idea flux.1:07:07–1:12:28 · Guest disagreement 2/10 Lightning Round, AI Governance, and Safety Takeaways In the lightning round, Kanjun deflects Swyx's hypothetical alternative company question by asserting Imbue is the only thing worth building. She concludes with Imbue's multi-pillar approach to AI safety and empirical policy work with the Department of Commerce.2:57–7:13 · The hosts pushing back 2/10 Early Startups and Lessons from Sorceress AI Swyx draws a sharp analogy between Sorceress's recruiting business and the dating app business model of endogenous churn. Kanjun validates the comparison and explains how trust barriers informed Imbue's interface approach.7:13–11:28 · The hosts pushing back 1/10 Founding Generally Intelligent and Free Intellectual Energy Alessio asks about starting Generally Intelligent in 2021 before agents were mainstream. Kanjun explains her foundational thesis on free intellectual energy analogous to the Industrial Revolution and early self-supervised learning breakthroughs.11:29–15:24 · The hosts pushing back 2/10 Reasoning Models and Personal Computing Vision Alessio and Swyx explore Imbue's focus on reasoning models and ask how they navigate between a pure research lab and a product company. Kanjun draws parallels to Apple's long multi-touch development cycles.15:24–17:52 · The hosts pushing back 1/10 Imbue's Internal Culture and Engineering Structure Alessio queries how Imbue manages its unique dual culture after raising 200 million dollars. Kanjun outlines her social process philosophy of treating employees as creative agents rather than disposable corporate assets.17:52–22:31 · The hosts pushing back 5/10 Optimizing Pre-Training and Synthetic Data Generation Swyx brings strong domain depth regarding compute versus data spending in foundation models, explicitly questioning whether synthetic model data causes distribution collapse. Kanjun explains that Imbue intentionally wants a spiky reasoning distribution rather than matching web text.22:32–27:36 · The hosts pushing back 2/10 Limits of Reinforcement Learning and Avalon Simulation Alessio asks about the Avalon simulation benchmark, playfully prompting Kanjun to correct a quote. Kanjun delivers a thorough technical breakdown of why pure reinforcement learning without curricula fails at planning and why reasoning must happen in language.27:37–32:13 · The hosts pushing back 3/10 Current Agent Limitations and Architectural Abstractions Swyx presses on whether simple single-task agents face an asymptote ceiling. Kanjun readily agrees, acknowledging that current agent abstractions are severely leaky and that pile-of-hacks solutions fail production reliability bars.32:14–37:53 · The hosts pushing back 4/10 Defining Reasoning and Natural Language Programming Alessio references George Hotz's claim that sufficient compute will automatically emerge reasoning from human data. Kanjun explicitly rejects that premise, arguing that human reasoning strategies must be generated intentionally and that code serves as a structured curriculum.37:54–46:54 · The hosts pushing back 3/10 Interface Paradigms, Agent Protocols, and Evaluation The hosts and Kanjun discuss skeuomorphic chat interfaces, historical computing transitions, agent communication protocols, and recent research like MetaGPT and Voyager. Both sides trade specific literature references comfortably.46:55–54:14 · The hosts pushing back 6/10 Context Memory Limitations and Imbue's Positioning Swyx pushes hard with direct competitive questions, asking why RAG is insufficient for complex tasks and confronting Kanjun on why candidates should pick Imbue over Adept. Kanjun provides a structured defense centered on non-leaky abstractions and dogfooding.54:16–1:07:07 · The hosts pushing back 2/10 Curiosity, Online Learning, and Cultivating Scenius Alessio and Swyx explore Kanjun's philosophical website questions and deep involvement in building hacker house ecosystems like The Archive and South Park Commons. Kanjun and Swyx discuss the concept of scenius and idea flux.1:07:07–1:12:28 · The hosts pushing back 2/10 Lightning Round, AI Governance, and Safety Takeaways In the lightning round, Kanjun deflects Swyx's hypothetical alternative company question by asserting Imbue is the only thing worth building. She concludes with Imbue's multi-pillar approach to AI safety and empirical policy work with the Department of Commerce.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 50.5% · guest 49.5%0:00 · the hosts 50.5% · guest 49.5%3:00 · the hosts 18.1% · guest 81.9%3:00 · the hosts 18.1% · guest 81.9%6:00 · the hosts 8.9% · guest 91.1%6:00 · the hosts 8.9% · guest 91.1%9:00 · the hosts 9.2% · guest 90.8%9:00 · the hosts 9.2% · guest 90.8%12:00 · the hosts 3.9% · guest 96.1%12:00 · the hosts 3.9% · guest 96.1%15:00 · the hosts 12.6% · guest 87.4%15:00 · the hosts 12.6% · guest 87.4%18:00 · the hosts 37.1% · guest 62.9%18:00 · the hosts 37.1% · guest 62.9%21:00 · the hosts 45.3% · guest 54.7%21:00 · the hosts 45.3% · guest 54.7%24:00 · the hosts 8.5% · guest 91.5%24:00 · the hosts 8.5% · guest 91.5%27:00 · the hosts 16.5% · guest 83.5%27:00 · the hosts 16.5% · guest 83.5%30:00 · the hosts 29.1% · guest 70.9%30:00 · the hosts 29.1% · guest 70.9%33:00 · the hosts 40.5% · guest 59.5%33:00 · the hosts 40.5% · guest 59.5%36:00 · the hosts 16.4% · guest 83.6%36:00 · the hosts 16.4% · guest 83.6%39:00 · the hosts 17.4% · guest 82.6%39:00 · the hosts 17.4% · guest 82.6%42:00 · the hosts 42.8% · guest 57.2%42:00 · the hosts 42.8% · guest 57.2%45:00 · the hosts 46.8% · guest 53.2%45:00 · the hosts 46.8% · guest 53.2%48:00 · the hosts 9.6% · guest 90.4%48:00 · the hosts 9.6% · guest 90.4%51:00 · the hosts 34% · guest 66%51:00 · the hosts 34% · guest 66%54:00 · the hosts 44.6% · guest 55.4%54:00 · the hosts 44.6% · guest 55.4%57:00 · the hosts 28.8% · guest 71.2%57:00 · the hosts 28.8% · guest 71.2%1:00:00 · the hosts 37.4% · guest 62.6%1:00:00 · the hosts 37.4% · guest 62.6%1:03:00 · the hosts 28.8% · guest 71.2%1:03:00 · the hosts 28.8% · guest 71.2%1:06:00 · the hosts 21.9% · guest 78.1%1:06:00 · the hosts 21.9% · guest 78.1%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 35.5% · guest 64.5%1:12:00 · the hosts 35.5% · guest 64.5%
Sharpest disagreement ▶ 32:45 Kanjun firmly rejects George Hotz's scale-solves-reasoning thesis

Kanjun directly counters the premise introduced from George Hotz, firmly stating that reasoning cannot be acquired by merely throwing massive data and compute at black boxes.

Hardest push from the hosts ▶ 49:49 Swyx directly presses Kanjun on Imbue versus Adept

Swyx asks an unvarnished competitive question, pushing Kanjun to justify why top talent should join Imbue instead of its most direct venture-backed peer Adept.

Biggest teaching moment ▶ 24:35 Kanjun details the structural failure modes of pure reinforcement learning

Kanjun provides a clear tutorial on why reinforcement learning algorithms fail completely without explicit curricula and lack the capacity for high-level planning.

The host holds their own ▶ 21:09 Swyx challenges synthetic data on distribution resampling bias

Swyx brings rigorous machine learning skepticism, contrasting foundation model spending profiles and questioning the statistical validity of training models on synthetic model outputs.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Early Startups and Lessons from Sorceress AI 5312 Swyx draws a sharp analogy between Sorceress's recruiting business and the dating app business model of endogenous churn. Kanjun validates the comparison and explains how trust barriers informed Imbue's interface approach.
Founding Generally Intelligent and Free Intellectual Energy 4411 Alessio asks about starting Generally Intelligent in 2021 before agents were mainstream. Kanjun explains her foundational thesis on free intellectual energy analogous to the Industrial Revolution and early self-supervised learning breakthroughs.
Reasoning Models and Personal Computing Vision 5312 Alessio and Swyx explore Imbue's focus on reasoning models and ask how they navigate between a pure research lab and a product company. Kanjun draws parallels to Apple's long multi-touch development cycles.
Imbue's Internal Culture and Engineering Structure 4311 Alessio queries how Imbue manages its unique dual culture after raising 200 million dollars. Kanjun outlines her social process philosophy of treating employees as creative agents rather than disposable corporate assets.
Optimizing Pre-Training and Synthetic Data Generation 7425 Swyx brings strong domain depth regarding compute versus data spending in foundation models, explicitly questioning whether synthetic model data causes distribution collapse. Kanjun explains that Imbue intentionally wants a spiky reasoning distribution rather than matching web text.
Limits of Reinforcement Learning and Avalon Simulation 5622 Alessio asks about the Avalon simulation benchmark, playfully prompting Kanjun to correct a quote. Kanjun delivers a thorough technical breakdown of why pure reinforcement learning without curricula fails at planning and why reasoning must happen in language.
Current Agent Limitations and Architectural Abstractions 5413 Swyx presses on whether simple single-task agents face an asymptote ceiling. Kanjun readily agrees, acknowledging that current agent abstractions are severely leaky and that pile-of-hacks solutions fail production reliability bars.
Defining Reasoning and Natural Language Programming 6544 Alessio references George Hotz's claim that sufficient compute will automatically emerge reasoning from human data. Kanjun explicitly rejects that premise, arguing that human reasoning strategies must be generated intentionally and that code serves as a structured curriculum.
Interface Paradigms, Agent Protocols, and Evaluation 7323 The hosts and Kanjun discuss skeuomorphic chat interfaces, historical computing transitions, agent communication protocols, and recent research like MetaGPT and Voyager. Both sides trade specific literature references comfortably.
Context Memory Limitations and Imbue's Positioning 7426 Swyx pushes hard with direct competitive questions, asking why RAG is insufficient for complex tasks and confronting Kanjun on why candidates should pick Imbue over Adept. Kanjun provides a structured defense centered on non-leaky abstractions and dogfooding.
Curiosity, Online Learning, and Cultivating Scenius 6312 Alessio and Swyx explore Kanjun's philosophical website questions and deep involvement in building hacker house ecosystems like The Archive and South Park Commons. Kanjun and Swyx discuss the concept of scenius and idea flux.
Lightning Round, AI Governance, and Safety Takeaways 4322 In the lightning round, Kanjun deflects Swyx's hypothetical alternative company question by asserting Imbue is the only thing worth building. She concludes with Imbue's multi-pillar approach to AI safety and empirical policy work with the Department of Commerce.

Statements from this episode (20)

Insight
Qiu: Recruiting software suffers high churn because success halts hiring
“Recruiting is a market that is endogenously high churn. Which means, because people start hiring, and then we hire the role for them, and they stop hiring.”
Kanjun Qiu Oct 21, 2023 ▶ 5:42
Insight
Qiu: AI will provide free intellectual energy like the Industrial Revolution
“And what's happened over the last 150 years since the Industrial Revolution is we've kind of gotten free energy. Like, energy is way more free than it was a 150 years ago. And so, as a result, we've built all these technologies, like the stove, and the dishwas…”
Kanjun Qiu Oct 21, 2023 ▶ 8:48
Opinion
Qiu: Reasoning is the single biggest blocker for AI agents
“Reasoning is actually, we believe the biggest blocker to agents or systems that can do these larger goals.”
Kanjun Qiu Oct 21, 2023 ▶ 11:53
Insight
Qiu: Models lack reasoning because the internet lacks explicit reasoning data
“And models today, they're not optimized for reasoning. It turns out that there's not actually that much explicit reasoning data on the internet.”
Kanjun Qiu Oct 21, 2023 ▶ 13:03
Assertion Not publicly verifiable
Imbue's CARBS hyperparameter optimizer originated from plasma physics concepts
“CARBS, our hyperparameter optimizer, came from Abe trying to automate his own research process doing hyperparameter optimization, and he actually pulled some ideas from plasma physics, he's a plasma physicist, to make the local search work.”
Kanjun Qiu Oct 21, 2023 ▶ 16:39
Insight
Kanjun Qiu: Agent reliability's final 20% is as hard as self-driving cars
“Agents haven't been productized yet for, partly for this reason, is that, like, the abstractions are very leaky. You know, we can get, like, 80% of the way there, but, like, self-driving cars, like, the remaining 20% is actually really difficult.”
Kanjun Qiu Oct 21, 2023 ▶ 18:44
Insight
Qiu: Synthetic code data can improve reasoning better than human code
“Code also is a big piece of improving reasoning. So yeah generated code is not That much worse than, like, regular human written code. You might even say it could be better in a lot of ways.”
Kanjun Qiu Oct 21, 2023 ▶ 22:17
Insight
Qiu: RL algorithms do not work at all without a curriculum
“One is with no curriculum, RL algorithms don't work at all.”
Kanjun Qiu Oct 21, 2023 ▶ 24:30
Insight
Qiu: Pure reinforcement learning cannot deliver planning and reasoning
“The second thing we learned is that reinforcement learning is not a good vehicle. Like, pure reinforcement learning is not a good vehicle for planning and reasoning.”
Kanjun Qiu Oct 21, 2023 ▶ 25:16
Opinion
Qiu: Language reasoning in agents is primarily for human inspectability
“I personally think of it as it's much more important for us, the human user. So I think you probably could get end to end agents that work and are fairly general at some point in the future. But I think you don't want that. Like, we actually want agents that w…”
Kanjun Qiu Oct 21, 2023 ▶ 26:46
Insight
Kanjun Qiu: Most autonomous AI agents do not work well today
“So to your question of, like, what agents work well and what doesn't work well, like, most of the agents don't work well, and we're slowly making them work better by improving the underlying model and improving these.”
Kanjun Qiu Oct 21, 2023 ▶ 30:05
Prediction Not checkable as stated
Kanjun Qiu: Agent development will evolve beyond bare metal to higher abstractions
“And I think it's basically a similar route here where we're like in the like bare metal phase of agent building, and we will eventually get to something with much nicer abstractions.”
Kanjun Qiu Oct 21, 2023 ▶ 32:03
Disclosure
Qiu: Imbue generates specific reasoning data rather than relying on web data
“So I think internally, yeah, we have a lot of thoughts on what reasoning is, and we generate a lot more specific data. We're not just like, oh, it'll figure out reasoning from this black box or like, it'll figure out reasoning from the data that, that exists.”
Kanjun Qiu Oct 21, 2023 ▶ 33:45
Insight
Qiu: Code is the most explicit reasoning curriculum for AI models
“Code is the most explicit example of reasoning data on the internet. Yeah. And it's not only structured, it's actually very explicit, which is nice. You know, it says this variable means this and then it uses this variable and then the function does this. Like…”
Kanjun Qiu Oct 21, 2023 ▶ 36:16
Insight
Kanjun Qiu: Chat is a skeuomorphic, primitive interface for AI agents
“Chat as an interface is skeuomorphic. So in the early days, when we made word processors on our computers, they had notepad lines because that's what we understood you know, these like objects to be chat. Like texting someone is something we understand. So tex…”
Kanjun Qiu Oct 21, 2023 ▶ 38:39
Opinion
Kanjun Qiu: Standardizing agent protocols is premature because agents don't work yet
“Part of why I think it's early is because the issue with agents is it's not quite like the internet where you could like make a website and the website would appear. The issue with agents is that they don't work. And so it may be a bit early to figure out what…”
Kanjun Qiu Oct 21, 2023 ▶ 42:32
Prediction Not checkable as stated
Qiu: Memory limitations block AI agents from handling complex, long-running tasks
“I think what we'll see is we'll get like relatively simplistic agents pretty soon, and they will get more and more complex. And there's like a future wave in which they are able to do these like really difficult, really long running tasks. And the blocker to t…”
Kanjun Qiu Oct 21, 2023 ▶ 47:34
Insight
Qiu: RAG is inadequate for scientific AI reasoning and cumulative synthesis
“I don't think RAG is enough for that kind of thing. But RAG is certainly enough for, like, user preferences and things like that.”
Kanjun Qiu Oct 21, 2023 ▶ 49:33
Assertion Supported
Kanjun Qiu: Larger models fine-tune faster and with higher sample efficiency
“As models get bigger, they fine tune faster. So they're more sample efficient as they get bigger.”
Kanjun Qiu Oct 21, 2023 ▶ 59:05
Assertion Partly supported
Imbue built an AI agent to analyze 20,000 Commerce Department policy proposals
“We built an agent that helped us analyze the, like, 20,000 pages of policy proposals submitted to the Department of Commerce request for AI policy proposals, and we, like, looked at what were the problems people brought up, and what were the solutions they pre…”
Kanjun Qiu Oct 21, 2023 ▶ 1:11:00
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.