Nov 16, 2023 · 32m · no-priors

No Priors Ep. 41 | With Imbue Co-Founders Kanjun Qiu and Josh Albrecht

Kanjun Qiu · 12m spoken Josh Albrecht · 10m spoken Sarah Guo · 4m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Imbue co-founders Kanjun Qiu and Josh Albrecht discuss their mission to build autonomous AI agents capable of robust multi-step reasoning and coding. They explain how programmatic reasoning, rigorous internal dogfooding, and ergonomic developer abstractions are redefining software creation and human-computer interaction.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.2% of the talking time here. How this is scored →

The hosts as informed peer 4.7 Guest teaching 4.0 Guest disagreement 1.4 The hosts pushing back 1.4
05100:0010:0020:0030:000:38–4:12 · The hosts as informed peer 2/10 Origins of Imbue and Early Agent Experiments The hosts ask foundational opening questions regarding the genesis of Imbue. Kanjun and Josh provide an overview of their background with Sorceress and how scaling experiments led them to prioritize agent architectures over passive chatbots.4:12–7:40 · The hosts as informed peer 5/10 The Spectrum of Agency and Solving Reliability Through Reasoning Elad frames technological readiness using analogies of Theranos versus 1990s mobile technology. Kanjun reframes the problem away from a binary technology gap to a spectrum of reliability and reasoning error-correction.7:40–10:35 · The hosts as informed peer 6/10 Model Specialization, Inference Costs, and Generalization Elad demonstrates operational expertise by describing the industry workflow of prototyping on GPT-4 and downscaling to fine-tuned open source models to optimize inference costs. Josh and Kanjun explain using general agents to generate specialized code.10:35–14:06 · The hosts as informed peer 5/10 Theoretical Limits of LLMs and Code as a Reasoning Medium Sarah probes on process supervision and multi-step reasoning across frontier labs. Josh clearly educates on the theoretical algorithmic limits of autoregressive LLMs for multi-step arithmetic, requiring an outer control loop.14:07–17:10 · The hosts as informed peer 4/10 Research Methodology: Serious Use and Hierarchical Sub-Agents Sarah inquires about the internal structure of Imbue's research pipeline. The guests articulate their 'serious use' philosophy, showing how targeted sub-agents and broad task-solvers compose into capable systems.17:10–20:12 · The hosts as informed peer 6/10 Evaluation Strategies for Complex and Coding Agents Sarah demonstrates domain knowledge in evaluation design by detailing static analysis, compilation checks, and framework migrations. Josh explains decomposing agent metrics into granular objective signals.20:12–24:31 · The hosts as informed peer 4/10 Imbue's Mission to Build Ergonomic Agent Tooling Sarah asks whether Imbue is primarily a research lab or product organization. Kanjun clarifies their identity as a tooling company, drawing an analogy between current agent development and assembly code.24:31–28:05 · The hosts as informed peer 7/10 Capital Deployment, Compute Scale, and Team Leverage Sarah pushes back with the frontier lab consensus that fewer than 5,000 GPUs prevents competing on frontier reasoning. Josh and Kanjun explain their compute access while emphasizing data quality and small-team automation.28:05–32:11 · The hosts as informed peer 3/10 Transforming Software Quality and Recursive Self-Improvement Elad asks why coding is the central substrate. Josh demystifies recursive self-improvement as practical compounding engineering leverage and automated testing rather than speculative runaway superintelligence.0:38–4:12 · Guest teaching 2/10 Origins of Imbue and Early Agent Experiments The hosts ask foundational opening questions regarding the genesis of Imbue. Kanjun and Josh provide an overview of their background with Sorceress and how scaling experiments led them to prioritize agent architectures over passive chatbots.4:12–7:40 · Guest teaching 5/10 The Spectrum of Agency and Solving Reliability Through Reasoning Elad frames technological readiness using analogies of Theranos versus 1990s mobile technology. Kanjun reframes the problem away from a binary technology gap to a spectrum of reliability and reasoning error-correction.7:40–10:35 · Guest teaching 4/10 Model Specialization, Inference Costs, and Generalization Elad demonstrates operational expertise by describing the industry workflow of prototyping on GPT-4 and downscaling to fine-tuned open source models to optimize inference costs. Josh and Kanjun explain using general agents to generate specialized code.10:35–14:06 · Guest teaching 6/10 Theoretical Limits of LLMs and Code as a Reasoning Medium Sarah probes on process supervision and multi-step reasoning across frontier labs. Josh clearly educates on the theoretical algorithmic limits of autoregressive LLMs for multi-step arithmetic, requiring an outer control loop.14:07–17:10 · Guest teaching 4/10 Research Methodology: Serious Use and Hierarchical Sub-Agents Sarah inquires about the internal structure of Imbue's research pipeline. The guests articulate their 'serious use' philosophy, showing how targeted sub-agents and broad task-solvers compose into capable systems.17:10–20:12 · Guest teaching 3/10 Evaluation Strategies for Complex and Coding Agents Sarah demonstrates domain knowledge in evaluation design by detailing static analysis, compilation checks, and framework migrations. Josh explains decomposing agent metrics into granular objective signals.20:12–24:31 · Guest teaching 3/10 Imbue's Mission to Build Ergonomic Agent Tooling Sarah asks whether Imbue is primarily a research lab or product organization. Kanjun clarifies their identity as a tooling company, drawing an analogy between current agent development and assembly code.24:31–28:05 · Guest teaching 5/10 Capital Deployment, Compute Scale, and Team Leverage Sarah pushes back with the frontier lab consensus that fewer than 5,000 GPUs prevents competing on frontier reasoning. Josh and Kanjun explain their compute access while emphasizing data quality and small-team automation.28:05–32:11 · Guest teaching 4/10 Transforming Software Quality and Recursive Self-Improvement Elad asks why coding is the central substrate. Josh demystifies recursive self-improvement as practical compounding engineering leverage and automated testing rather than speculative runaway superintelligence.0:38–4:12 · Guest disagreement 1/10 Origins of Imbue and Early Agent Experiments The hosts ask foundational opening questions regarding the genesis of Imbue. Kanjun and Josh provide an overview of their background with Sorceress and how scaling experiments led them to prioritize agent architectures over passive chatbots.4:12–7:40 · Guest disagreement 2/10 The Spectrum of Agency and Solving Reliability Through Reasoning Elad frames technological readiness using analogies of Theranos versus 1990s mobile technology. Kanjun reframes the problem away from a binary technology gap to a spectrum of reliability and reasoning error-correction.7:40–10:35 · Guest disagreement 1/10 Model Specialization, Inference Costs, and Generalization Elad demonstrates operational expertise by describing the industry workflow of prototyping on GPT-4 and downscaling to fine-tuned open source models to optimize inference costs. Josh and Kanjun explain using general agents to generate specialized code.10:35–14:06 · Guest disagreement 2/10 Theoretical Limits of LLMs and Code as a Reasoning Medium Sarah probes on process supervision and multi-step reasoning across frontier labs. Josh clearly educates on the theoretical algorithmic limits of autoregressive LLMs for multi-step arithmetic, requiring an outer control loop.14:07–17:10 · Guest disagreement 1/10 Research Methodology: Serious Use and Hierarchical Sub-Agents Sarah inquires about the internal structure of Imbue's research pipeline. The guests articulate their 'serious use' philosophy, showing how targeted sub-agents and broad task-solvers compose into capable systems.17:10–20:12 · Guest disagreement 1/10 Evaluation Strategies for Complex and Coding Agents Sarah demonstrates domain knowledge in evaluation design by detailing static analysis, compilation checks, and framework migrations. Josh explains decomposing agent metrics into granular objective signals.20:12–24:31 · Guest disagreement 2/10 Imbue's Mission to Build Ergonomic Agent Tooling Sarah asks whether Imbue is primarily a research lab or product organization. Kanjun clarifies their identity as a tooling company, drawing an analogy between current agent development and assembly code.24:31–28:05 · Guest disagreement 2/10 Capital Deployment, Compute Scale, and Team Leverage Sarah pushes back with the frontier lab consensus that fewer than 5,000 GPUs prevents competing on frontier reasoning. Josh and Kanjun explain their compute access while emphasizing data quality and small-team automation.28:05–32:11 · Guest disagreement 1/10 Transforming Software Quality and Recursive Self-Improvement Elad asks why coding is the central substrate. Josh demystifies recursive self-improvement as practical compounding engineering leverage and automated testing rather than speculative runaway superintelligence.0:38–4:12 · The hosts pushing back 1/10 Origins of Imbue and Early Agent Experiments The hosts ask foundational opening questions regarding the genesis of Imbue. Kanjun and Josh provide an overview of their background with Sorceress and how scaling experiments led them to prioritize agent architectures over passive chatbots.4:12–7:40 · The hosts pushing back 2/10 The Spectrum of Agency and Solving Reliability Through Reasoning Elad frames technological readiness using analogies of Theranos versus 1990s mobile technology. Kanjun reframes the problem away from a binary technology gap to a spectrum of reliability and reasoning error-correction.7:40–10:35 · The hosts pushing back 1/10 Model Specialization, Inference Costs, and Generalization Elad demonstrates operational expertise by describing the industry workflow of prototyping on GPT-4 and downscaling to fine-tuned open source models to optimize inference costs. Josh and Kanjun explain using general agents to generate specialized code.10:35–14:06 · The hosts pushing back 2/10 Theoretical Limits of LLMs and Code as a Reasoning Medium Sarah probes on process supervision and multi-step reasoning across frontier labs. Josh clearly educates on the theoretical algorithmic limits of autoregressive LLMs for multi-step arithmetic, requiring an outer control loop.14:07–17:10 · The hosts pushing back 1/10 Research Methodology: Serious Use and Hierarchical Sub-Agents Sarah inquires about the internal structure of Imbue's research pipeline. The guests articulate their 'serious use' philosophy, showing how targeted sub-agents and broad task-solvers compose into capable systems.17:10–20:12 · The hosts pushing back 1/10 Evaluation Strategies for Complex and Coding Agents Sarah demonstrates domain knowledge in evaluation design by detailing static analysis, compilation checks, and framework migrations. Josh explains decomposing agent metrics into granular objective signals.20:12–24:31 · The hosts pushing back 1/10 Imbue's Mission to Build Ergonomic Agent Tooling Sarah asks whether Imbue is primarily a research lab or product organization. Kanjun clarifies their identity as a tooling company, drawing an analogy between current agent development and assembly code.24:31–28:05 · The hosts pushing back 3/10 Capital Deployment, Compute Scale, and Team Leverage Sarah pushes back with the frontier lab consensus that fewer than 5,000 GPUs prevents competing on frontier reasoning. Josh and Kanjun explain their compute access while emphasizing data quality and small-team automation.28:05–32:11 · The hosts pushing back 1/10 Transforming Software Quality and Recursive Self-Improvement Elad asks why coding is the central substrate. Josh demystifies recursive self-improvement as practical compounding engineering leverage and automated testing rather than speculative runaway superintelligence.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 25.3% · guest 74.7%0:00 · the hosts 25.3% · guest 74.7%3:00 · the hosts 24.7% · guest 75.3%3:00 · the hosts 24.7% · guest 75.3%6:00 · the hosts 33.7% · guest 66.3%6:00 · the hosts 33.7% · guest 66.3%9:00 · the hosts 18.8% · guest 81.2%9:00 · the hosts 18.8% · guest 81.2%12:00 · the hosts 13.7% · guest 86.3%12:00 · the hosts 13.7% · guest 86.3%15:00 · the hosts 13.4% · guest 86.6%15:00 · the hosts 13.4% · guest 86.6%18:00 · the hosts 36.9% · guest 63.1%18:00 · the hosts 36.9% · guest 63.1%21:00 · the hosts 9.3% · guest 90.7%21:00 · the hosts 9.3% · guest 90.7%24:00 · the hosts 35.1% · guest 64.9%24:00 · the hosts 35.1% · guest 64.9%27:00 · the hosts 2.6% · guest 97.4%27:00 · the hosts 2.6% · guest 97.4%30:00 · the hosts 19% · guest 81%30:00 · the hosts 19% · guest 81%
Sharpest disagreement ▶ 4:57 Reframing binary technological roadblocks

Kanjun gently rejects Elad's premise of a missing binary technological blocker, arguing instead that agents exist along a spectrum of reliability and incremental autonomy.

Hardest push from the hosts ▶ 26:15 Challenging the 5,000 GPU frontier threshold

Sarah directly confronts the guests with the prevailing frontier lab assumption that small teams cannot compete on core reasoning without massive 5,000+ GPU compute commitments.

Biggest teaching moment ▶ 11:10 Formal limits of next-token prediction

Josh educates the audience and hosts on the hard theoretical boundaries of standard LLMs, proving why general arithmetic and reasoning require an external execution wrapper.

The host holds their own ▶ 7:40 Dissecting production inference economics

Elad displays strong technical market expertise by outlining how production teams prototype on GPT-4 before distilling down to fine-tuned open-source models for cost containment.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of Imbue and Early Agent Experiments 2211 The hosts ask foundational opening questions regarding the genesis of Imbue. Kanjun and Josh provide an overview of their background with Sorceress and how scaling experiments led them to prioritize agent architectures over passive chatbots.
The Spectrum of Agency and Solving Reliability Through Reasoning 5522 Elad frames technological readiness using analogies of Theranos versus 1990s mobile technology. Kanjun reframes the problem away from a binary technology gap to a spectrum of reliability and reasoning error-correction.
Model Specialization, Inference Costs, and Generalization 6411 Elad demonstrates operational expertise by describing the industry workflow of prototyping on GPT-4 and downscaling to fine-tuned open source models to optimize inference costs. Josh and Kanjun explain using general agents to generate specialized code.
Theoretical Limits of LLMs and Code as a Reasoning Medium 5622 Sarah probes on process supervision and multi-step reasoning across frontier labs. Josh clearly educates on the theoretical algorithmic limits of autoregressive LLMs for multi-step arithmetic, requiring an outer control loop.
Research Methodology: Serious Use and Hierarchical Sub-Agents 4411 Sarah inquires about the internal structure of Imbue's research pipeline. The guests articulate their 'serious use' philosophy, showing how targeted sub-agents and broad task-solvers compose into capable systems.
Evaluation Strategies for Complex and Coding Agents 6311 Sarah demonstrates domain knowledge in evaluation design by detailing static analysis, compilation checks, and framework migrations. Josh explains decomposing agent metrics into granular objective signals.
Imbue's Mission to Build Ergonomic Agent Tooling 4321 Sarah asks whether Imbue is primarily a research lab or product organization. Kanjun clarifies their identity as a tooling company, drawing an analogy between current agent development and assembly code.
Capital Deployment, Compute Scale, and Team Leverage 7523 Sarah pushes back with the frontier lab consensus that fewer than 5,000 GPUs prevents competing on frontier reasoning. Josh and Kanjun explain their compute access while emphasizing data quality and small-team automation.
Transforming Software Quality and Recursive Self-Improvement 3411 Elad asks why coding is the central substrate. Josh demystifies recursive self-improvement as practical compounding engineering leverage and automated testing rather than speculative runaway superintelligence.

Statements from this episode (23)

Opinion
Qiu: Self-supervised AI may learn representations akin to human cognition
“There's something really interesting here where maybe machines are learning the same kinds of representations or similar representations to what humans are learning. And maybe they can get to a point where they can actually do the types of things that humans a…”
Kanjun Qiu Nov 16, 2023 ▶ 2:09
Insight
Albrecht: The true promise of AI lies in autonomous agents, not chatbots
“Right now, you can ask some kind of chatbot something and it'll give you back a response, but the burden is sort of on you to go do something with that to verify whether it's correct or not. I think the real promise of AI is if we can get systems that can actu…”
Josh Albrecht Nov 16, 2023 ▶ 2:53
Insight
Qiu: Autonomous AI agents represent a calculator-to-computer leap in technology
“The diff between this, where we are today, and that is kind of like the diff between the first calculator and where computers are today.”
Kanjun Qiu Nov 16, 2023 ▶ 3:47
Insight
Qiu: Chain of thought and tree of thought function as error correction
“Reasoning is one big piece of improving reliability, and second chunk of things is like all of this error correction, and I think like chain of thought, tree of thought, these are error correction techniques.”
Kanjun Qiu Nov 16, 2023 ▶ 7:16
Prediction Not checkable as stated
Qiu: Pragmatic, smaller models will eventually address the majority of AI workflows
“And I suspect we're going to see something similar where a lot of use cases are going to be able to be addressed by something pretty pragmatic and relatively small.”
Kanjun Qiu Nov 16, 2023 ▶ 10:19
Assertion Not checkable as stated
Qiu: The AI industry hasn't pushed data limits on small models
“We're definitely not pushing the bounds of what we can do with data today on small models, and so, you know, smaller things can work well.”
Kanjun Qiu Nov 16, 2023 ▶ 10:28
Insight
Albrecht: Pure LLMs theoretically cannot learn general multiplication algorithms due to context
“Like, we know even in theoretical senses, like, they cannot learn to do multiplication in the general sense because it literally doesn't fit in the context window, right? Like, multiplication, they can learn to do addition in a modular sense, and they can lear…”
Josh Albrecht Nov 16, 2023 ▶ 11:24
Insight
Albrecht: Repeatable AI agent workflows must progressively transition into explicit code
“And as you do things that you want to do more robustly and you want to do in a more repeatable way, then you want to move it more towards code, right? And so To the extent that you've never seen this task before, maybe you should be doing it in this more kind …”
Josh Albrecht Nov 16, 2023 ▶ 13:29
Insight
Qiu: Training a giant monolithic model does not magically solve agent reliability
“It's not like, oh, magical, you know, we train a giant model and stick everything into it and then magically it works. Like it does not work. It'll get better at random parts of the agent loop, but that's not what we want.”
Kanjun Qiu Nov 16, 2023 ▶ 15:06
Insight
Qiu: Specialized reasoning planners coordinating sub-agents outperform single monolithic models
“You can kind of have this like more general reasoning layer and also a bunch of sub-agents where it, that, That general reasoning layer is actually very specific. It's a specific planner. It's not that good at, like, browsing the web and things like that, but …”
Kanjun Qiu Nov 16, 2023 ▶ 16:52
Insight
Albrecht: Imbue focuses on coding agents because objective evaluation is significantly easier
“One of the reasons why we work on code is that there are objective answers to a lot of these questions, either the test pass or they don't, either the function is correct or it isn't. Those kinds of things are much easier to evaluate.”
Josh Albrecht Nov 16, 2023 ▶ 18:25
Insight
Qiu: Evaluating AI models solely on binary correctness loses critical evaluation data
“Part of why a lot of teams try to work on just math or code reasoning is because those are the easiest to evaluate and like the clearest Answers, but just relying on, like, is the output correct or not, that loses a lot of information in the evaluation.”
Kanjun Qiu Nov 16, 2023 ▶ 19:04
Insight
Qiu: Developing AI agents today is comparable to writing code in assembly
“So today, like writing agents feels like writing code in assembly, and that really limits the types of agents we can build and also limits the number of people who can build them.”
Kanjun Qiu Nov 16, 2023 ▶ 21:15
Prediction Held up
Albrecht: Narrow AI agents for email triage will work by late 2024
“Yeah, I think a year from now we're going to start to see some of these use cases actually work that today you could you can write these like we have the capabilities you can make some kind of agent to triage your email or to do scheduling or many of these wor…”
Josh Albrecht Nov 16, 2023 ▶ 22:22
Prediction Open · timeframe Nov 2028
Albrecht: Users will create bespoke agents via natural language within five years
“And I think five years from now we're going to have something where it's not just you know, okay we have a scheduling bot we have this other thing but we really have these more general more robust systems where Each of us can individually say, like, I want a t…”
Josh Albrecht Nov 16, 2023 ▶ 22:39
Disclosure
Albrecht: A significant fraction of Imbue's $200M funding round will fund compute
“I mean, I think actually a significant fraction of that is going to go to compute.”
Josh Albrecht Nov 16, 2023 ▶ 24:53
Assertion Not checkable as stated
Qiu: Imbue trains state-of-the-art AI models with only 13 or 14 people
“We're kind of like training state, state of the art models with like 14, 13 people.”
Kanjun Qiu Nov 16, 2023 ▶ 25:54
Assertion Not checkable as stated
Albrecht: Imbue has sufficient compute to train state-of-the-art sized AI models
“We have enough compute to be able to train models that are as large as the largest models have been trained today to date.”
Josh Albrecht Nov 16, 2023 ▶ 27:03
Prediction Not checkable as stated
Albrecht: Imbue will hire fewer recruiting coordinators as internal agents handle scheduling
“Now, you know, I think probably within the next year, we'll probably, you know, not be hiring as many recruiting coordinators because, oh, we're going to do some of the scheduling with the agent that we've built, right?”
Josh Albrecht Nov 16, 2023 ▶ 28:36
Disclosure
Albrecht: Imbue is currently writing unit tests automatically using its AI agents
“We're writing unit tests literally right now automatically.”
Josh Albrecht Nov 16, 2023 ▶ 28:48
Prediction Not checkable as stated
Qiu: Software output will explode as programming democratizes to non-coders
“Software is just dramatically underwritten because it's so hard to write code today. So, you know, as we said in the future, like, computers will be able to be programmed by regular people. What that means is, like, we're gonna write way, way, way more softwar…”
Kanjun Qiu Nov 16, 2023 ▶ 30:31
Prediction Not checkable as stated
Albrecht: AI agents will dramatically improve codebase quality across the industry
“I think there'll just be a huge flourishing of much higher quality, better software as a result, not just more software, but just taking the existing software and making it so much better, which will make it so much nicer and more fun to interact with as progr…”
Josh Albrecht Nov 16, 2023 ▶ 31:11
Prediction Not checkable as stated
Guo: AI advancements will enable 25-person startups to change the world
“We're going to have 25 person companies who can change the world. We're going to have more software, more custom software, and higher quality software for us all to use.”
Sarah Guo Nov 16, 2023 ▶ 32:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.