Oct 21, 2023 · 1h 12m · latent-space
Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Imbue CEO Kanjun Qiu examines why current autonomous AI agents encounter severe reliability bottlenecks, detailing how foundation models optimized for explicit natural language reasoning, synthetic code data, and robust software abstractions will transform personal computing.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 25.3% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Kanjun directly counters the premise introduced from George Hotz, firmly stating that reasoning cannot be acquired by merely throwing massive data and compute at black boxes.
Hardest push from the hosts ▶ 49:49 Swyx directly presses Kanjun on Imbue versus AdeptSwyx asks an unvarnished competitive question, pushing Kanjun to justify why top talent should join Imbue instead of its most direct venture-backed peer Adept.
Biggest teaching moment ▶ 24:35 Kanjun details the structural failure modes of pure reinforcement learningKanjun provides a clear tutorial on why reinforcement learning algorithms fail completely without explicit curricula and lack the capacity for high-level planning.
The host holds their own ▶ 21:09 Swyx challenges synthetic data on distribution resampling biasSwyx brings rigorous machine learning skepticism, contrasting foundation model spending profiles and questioning the statistical validity of training models on synthetic model outputs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Early Startups and Lessons from Sorceress AI | 5 | 3 | 1 | 2 | Swyx draws a sharp analogy between Sorceress's recruiting business and the dating app business model of endogenous churn. Kanjun validates the comparison and explains how trust barriers informed Imbue's interface approach. | |
| Founding Generally Intelligent and Free Intellectual Energy | 4 | 4 | 1 | 1 | Alessio asks about starting Generally Intelligent in 2021 before agents were mainstream. Kanjun explains her foundational thesis on free intellectual energy analogous to the Industrial Revolution and early self-supervised learning breakthroughs. | |
| Reasoning Models and Personal Computing Vision | 5 | 3 | 1 | 2 | Alessio and Swyx explore Imbue's focus on reasoning models and ask how they navigate between a pure research lab and a product company. Kanjun draws parallels to Apple's long multi-touch development cycles. | |
| Imbue's Internal Culture and Engineering Structure | 4 | 3 | 1 | 1 | Alessio queries how Imbue manages its unique dual culture after raising 200 million dollars. Kanjun outlines her social process philosophy of treating employees as creative agents rather than disposable corporate assets. | |
| Optimizing Pre-Training and Synthetic Data Generation | 7 | 4 | 2 | 5 | Swyx brings strong domain depth regarding compute versus data spending in foundation models, explicitly questioning whether synthetic model data causes distribution collapse. Kanjun explains that Imbue intentionally wants a spiky reasoning distribution rather than matching web text. | |
| Limits of Reinforcement Learning and Avalon Simulation | 5 | 6 | 2 | 2 | Alessio asks about the Avalon simulation benchmark, playfully prompting Kanjun to correct a quote. Kanjun delivers a thorough technical breakdown of why pure reinforcement learning without curricula fails at planning and why reasoning must happen in language. | |
| Current Agent Limitations and Architectural Abstractions | 5 | 4 | 1 | 3 | Swyx presses on whether simple single-task agents face an asymptote ceiling. Kanjun readily agrees, acknowledging that current agent abstractions are severely leaky and that pile-of-hacks solutions fail production reliability bars. | |
| Defining Reasoning and Natural Language Programming | 6 | 5 | 4 | 4 | Alessio references George Hotz's claim that sufficient compute will automatically emerge reasoning from human data. Kanjun explicitly rejects that premise, arguing that human reasoning strategies must be generated intentionally and that code serves as a structured curriculum. | |
| Interface Paradigms, Agent Protocols, and Evaluation | 7 | 3 | 2 | 3 | The hosts and Kanjun discuss skeuomorphic chat interfaces, historical computing transitions, agent communication protocols, and recent research like MetaGPT and Voyager. Both sides trade specific literature references comfortably. | |
| Context Memory Limitations and Imbue's Positioning | 7 | 4 | 2 | 6 | Swyx pushes hard with direct competitive questions, asking why RAG is insufficient for complex tasks and confronting Kanjun on why candidates should pick Imbue over Adept. Kanjun provides a structured defense centered on non-leaky abstractions and dogfooding. | |
| Curiosity, Online Learning, and Cultivating Scenius | 6 | 3 | 1 | 2 | Alessio and Swyx explore Kanjun's philosophical website questions and deep involvement in building hacker house ecosystems like The Archive and South Park Commons. Kanjun and Swyx discuss the concept of scenius and idea flux. | |
| Lightning Round, AI Governance, and Safety Takeaways | 4 | 3 | 2 | 2 | In the lightning round, Kanjun deflects Swyx's hypothetical alternative company question by asserting Imbue is the only thing worth building. She concludes with Imbue's multi-pillar approach to AI safety and empirical policy work with the Department of Commerce. |