Aug 5, 2026 · 1h 22m · mad

How to Build Autonomous, Long-Horizon AI Agents | Basis

Mitch Trojanowski · 1h 2m spoken Matt Turck · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Basis co-founder Mitchell Trojanowski about the engineering, architectural, and organizational principles required to build reliable, long-horizon autonomous AI agents for complex domains like accounting. Trojanowski breaks down key concepts including state management, process supervision, open-source behavior specifications, context engineering, and the translation of human organizational workflows into scalable agent architectures.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.4% of the talking time here. How this is scored →

Matt as informed peer 4.4 Guest teaching 5.8 Guest disagreement 1.9 Matt pushing back 1.2
05100:0020:0040:001:00:001:20:001:24–4:22 · Matt as informed peer 4/10 Voice Interfaces and Whispering Context to AI Agents Matt sets the stage by citing a specific scene reported by Stephanie Palazzolo at The Information about Basis employees whispering into mics. Mitch explains that whispering gives high-bandwidth context without summarization loss, which agents prefer over concise human text.4:22–6:24 · Matt as informed peer 3/10 Why Accounting is a Crucial Knowledge Work Domain Matt asks if Mitch chose accounting because of agent technical challenges or market need. Mitch delivers a thoughtful breakdown of accounting as information compression over modern capitalism.6:24–8:45 · Matt as informed peer 3/10 Defining Agency and Long-Horizon Agentic Execution Matt asks for a concrete definition of long-horizon agency. Mitch reframes agency as a spectrum of decision-making autonomy constrained by LLM working memory limits.8:45–11:18 · Matt as informed peer 3/10 End-to-End Tax Return Workflows and Reviewer Trust Matt asks what autonomous execution looks like in Basis tax return workflows. Mitch explains that true autonomy produces reviewable assumptions and split artifacts rather than uninspectable end states.11:18–14:15 · Matt as informed peer 5/10 The History of Agents: ReAct Framework and State Management Matt traces the history of agents back to the 2022 ReAct paper and connects it to Christopher Nolan films. Mitch extends the framework using the movie Memento to illustrate how agents must write notes to regulate their own future inference states.14:15–18:24 · Matt as informed peer 5/10 BabyAGI and the Problem of Compounding Errors Matt brings up BabyAGI and the compounding error problem from 2023. Mitch outlines model history, explaining how Opus 3, o1, and o3 overcame attention degradation over long context windows.18:24–20:32 · Matt as informed peer 6/10 Process Supervision vs. Outcome Supervision in RL Matt cites OpenAI's 2023 'Let's Verify Step by Step' paper and its 800,000 human-labeled reasoning steps. Mitch explains the tradeoff between costly process supervision and outcome supervision scaled via RLVR in DeepSeek R1.20:32–25:15 · Matt as informed peer 5/10 METR Benchmarks and the Unique Properties of Coding Agents Matt asks if METR benchmark doubling rates hold true. Mitch questions METR's small sample size and explains why coding agents succeed due to runtime execution feedback and rich training data rather than purely verifiable rewards.25:15–31:26 · Matt as informed peer 4/10 Why Real-World AI Agents Struggle Outside of Coding Matt asks why agents struggle outside software development. Mitch highlights the lack of runtime compilers, the need to map human organizational checks, and data privacy hurdles in real-world professions.31:26–36:37 · Matt as informed peer 5/10 Translating Human Organizational Processes to Agent Design Matt pushes on why outcome evals are insufficient. Mitch explains that an agent passing 100 evals via Wikipedia is unemployable by an accounting firm because process integrity and primary-source citations are mandatory.36:37–42:17 · Matt as informed peer 4/10 Behavior Specs: Defining and Evaluating Expected Agent Behavior Matt asks about the structure and authorship of behavior specs in markdown. Mitch explains that meta behaviors bridge subjective product standards and machine execution, co-authored by accountants and ML researchers.42:17–50:04 · Matt as informed peer 6/10 Context Engineering, LLM Judges, and Organizational Reliability Matt raises a critical counterargument: does enforcing rigid human processes preclude emergent Move 37 breakthroughs? Mitch responds that enterprise clients purchase predictable reliability and auditability, not Move 37 surprises.50:04–52:42 · Matt as informed peer 4/10 Building Intuition Around Model Capabilities and Behavior Matt asks how developers design systems when underlying LLM internals remain opaque. Mitch explains treating the model as a black-box alien entity with activation states that require uncorrelated review trajectories.52:42–54:57 · Matt as informed peer 3/10 Developing Practical Intuition for Agent Systems Matt asks if agent intuition stems from reading research papers and talking to lab researchers. Mitch flatly dismisses that premise, arguing that real intuition comes exclusively from root-cause debugging one's own workflow automations.54:57–59:45 · Matt as informed peer 5/10 Open-Sourcing the Agent Behavior Standard with Braintrust Matt discusses the open-source collaboration with Braintrust and mentions the 4 AM Sunday Slack exchanges between Basis co-founders. Mitch explains the open-source standard for defining and evaluating agent behavior specs.59:45–1:04:31 · Matt as informed peer 4/10 Designing Ontologies and Company Canon for Agents Matt inquires about the practical design of ontologies in agent systems. Mitch differentiates coding agents from domain agents, noting that domain agents must own and structure their own runtime training context over long horizons.1:04:31–1:06:34 · Matt as informed peer 5/10 Managing Canonical Knowledge in Agent-Native Organizations Matt quotes Mitch on treating documentation with the same rigor as code bases where deleting a paragraph breaks the system. Mitch details how organizations must maintain singular canonical knowledge rather than fragmented Slack or Gong history.1:06:34–1:09:06 · Matt as informed peer 4/10 Emerging Roles: Language Architects and Context Engineers Matt asks about hiring for novel roles like language architects and context engineers. Mitch compares writing durable system prompts to drafting constitutional law that must cleanly withstand thousands of runtime interpretations.1:09:06–1:11:23 · Matt as informed peer 5/10 The Role of the Deployed Intelligence Team Matt asks about the deployed intelligence team concept. Mitch clarifies that DI engineers are neither forward-deployed software engineers nor agent PMs, but workflow transformation specialists embedded with accounting firms.1:11:23–1:14:31 · Matt as informed peer 4/10 System-Level Self-Improvement and Closing the Agent Feedback Loop Matt explores system-level self-improvement loops. Mitch critiques engineers who obsess over pristine code abstraction while leaving prompt context full of slop, stressing that context directly governs runtime performance.1:14:31–1:17:03 · Matt as informed peer 5/10 Reinforcement Learning, Reward Functions, and Model-Level Improvement Matt asks whether Basis plans to transition from harness engineering to fine-tuning model weights via RL. Mitch explains that developing the reward function and signal allocation is the primary challenge, regardless of where signal is applied.1:17:03–1:21:04 · Matt as informed peer 6/10 The Bitter Lesson, Technical Moats, and Business Strategy Matt presses Mitch on whether the Bitter Lesson will swallow harness engineering and if applied AI firms will need proprietary RL moats. Mitch emphatically rejects technical moats, arguing business distribution and workflow embedding determine enterprise value.1:21:04–1:22:30 · Matt as informed peer 3/10 Final Lessons and Advice for AI Builders Matt asks for final parting advice for AI builders. Mitch advises founders to look past daily Twitter hype and base technical strategy on fundamental, enduring paradigm shifts.1:24–4:22 · Guest teaching 4/10 Voice Interfaces and Whispering Context to AI Agents Matt sets the stage by citing a specific scene reported by Stephanie Palazzolo at The Information about Basis employees whispering into mics. Mitch explains that whispering gives high-bandwidth context without summarization loss, which agents prefer over concise human text.4:22–6:24 · Guest teaching 5/10 Why Accounting is a Crucial Knowledge Work Domain Matt asks if Mitch chose accounting because of agent technical challenges or market need. Mitch delivers a thoughtful breakdown of accounting as information compression over modern capitalism.6:24–8:45 · Guest teaching 6/10 Defining Agency and Long-Horizon Agentic Execution Matt asks for a concrete definition of long-horizon agency. Mitch reframes agency as a spectrum of decision-making autonomy constrained by LLM working memory limits.8:45–11:18 · Guest teaching 5/10 End-to-End Tax Return Workflows and Reviewer Trust Matt asks what autonomous execution looks like in Basis tax return workflows. Mitch explains that true autonomy produces reviewable assumptions and split artifacts rather than uninspectable end states.11:18–14:15 · Guest teaching 5/10 The History of Agents: ReAct Framework and State Management Matt traces the history of agents back to the 2022 ReAct paper and connects it to Christopher Nolan films. Mitch extends the framework using the movie Memento to illustrate how agents must write notes to regulate their own future inference states.14:15–18:24 · Guest teaching 5/10 BabyAGI and the Problem of Compounding Errors Matt brings up BabyAGI and the compounding error problem from 2023. Mitch outlines model history, explaining how Opus 3, o1, and o3 overcame attention degradation over long context windows.18:24–20:32 · Guest teaching 6/10 Process Supervision vs. Outcome Supervision in RL Matt cites OpenAI's 2023 'Let's Verify Step by Step' paper and its 800,000 human-labeled reasoning steps. Mitch explains the tradeoff between costly process supervision and outcome supervision scaled via RLVR in DeepSeek R1.20:32–25:15 · Guest teaching 6/10 METR Benchmarks and the Unique Properties of Coding Agents Matt asks if METR benchmark doubling rates hold true. Mitch questions METR's small sample size and explains why coding agents succeed due to runtime execution feedback and rich training data rather than purely verifiable rewards.25:15–31:26 · Guest teaching 6/10 Why Real-World AI Agents Struggle Outside of Coding Matt asks why agents struggle outside software development. Mitch highlights the lack of runtime compilers, the need to map human organizational checks, and data privacy hurdles in real-world professions.31:26–36:37 · Guest teaching 7/10 Translating Human Organizational Processes to Agent Design Matt pushes on why outcome evals are insufficient. Mitch explains that an agent passing 100 evals via Wikipedia is unemployable by an accounting firm because process integrity and primary-source citations are mandatory.36:37–42:17 · Guest teaching 6/10 Behavior Specs: Defining and Evaluating Expected Agent Behavior Matt asks about the structure and authorship of behavior specs in markdown. Mitch explains that meta behaviors bridge subjective product standards and machine execution, co-authored by accountants and ML researchers.42:17–50:04 · Guest teaching 7/10 Context Engineering, LLM Judges, and Organizational Reliability Matt raises a critical counterargument: does enforcing rigid human processes preclude emergent Move 37 breakthroughs? Mitch responds that enterprise clients purchase predictable reliability and auditability, not Move 37 surprises.50:04–52:42 · Guest teaching 6/10 Building Intuition Around Model Capabilities and Behavior Matt asks how developers design systems when underlying LLM internals remain opaque. Mitch explains treating the model as a black-box alien entity with activation states that require uncorrelated review trajectories.52:42–54:57 · Guest teaching 7/10 Developing Practical Intuition for Agent Systems Matt asks if agent intuition stems from reading research papers and talking to lab researchers. Mitch flatly dismisses that premise, arguing that real intuition comes exclusively from root-cause debugging one's own workflow automations.54:57–59:45 · Guest teaching 5/10 Open-Sourcing the Agent Behavior Standard with Braintrust Matt discusses the open-source collaboration with Braintrust and mentions the 4 AM Sunday Slack exchanges between Basis co-founders. Mitch explains the open-source standard for defining and evaluating agent behavior specs.59:45–1:04:31 · Guest teaching 6/10 Designing Ontologies and Company Canon for Agents Matt inquires about the practical design of ontologies in agent systems. Mitch differentiates coding agents from domain agents, noting that domain agents must own and structure their own runtime training context over long horizons.1:04:31–1:06:34 · Guest teaching 6/10 Managing Canonical Knowledge in Agent-Native Organizations Matt quotes Mitch on treating documentation with the same rigor as code bases where deleting a paragraph breaks the system. Mitch details how organizations must maintain singular canonical knowledge rather than fragmented Slack or Gong history.1:06:34–1:09:06 · Guest teaching 6/10 Emerging Roles: Language Architects and Context Engineers Matt asks about hiring for novel roles like language architects and context engineers. Mitch compares writing durable system prompts to drafting constitutional law that must cleanly withstand thousands of runtime interpretations.1:09:06–1:11:23 · Guest teaching 5/10 The Role of the Deployed Intelligence Team Matt asks about the deployed intelligence team concept. Mitch clarifies that DI engineers are neither forward-deployed software engineers nor agent PMs, but workflow transformation specialists embedded with accounting firms.1:11:23–1:14:31 · Guest teaching 7/10 System-Level Self-Improvement and Closing the Agent Feedback Loop Matt explores system-level self-improvement loops. Mitch critiques engineers who obsess over pristine code abstraction while leaving prompt context full of slop, stressing that context directly governs runtime performance.1:14:31–1:17:03 · Guest teaching 6/10 Reinforcement Learning, Reward Functions, and Model-Level Improvement Matt asks whether Basis plans to transition from harness engineering to fine-tuning model weights via RL. Mitch explains that developing the reward function and signal allocation is the primary challenge, regardless of where signal is applied.1:17:03–1:21:04 · Guest teaching 7/10 The Bitter Lesson, Technical Moats, and Business Strategy Matt presses Mitch on whether the Bitter Lesson will swallow harness engineering and if applied AI firms will need proprietary RL moats. Mitch emphatically rejects technical moats, arguing business distribution and workflow embedding determine enterprise value.1:21:04–1:22:30 · Guest teaching 5/10 Final Lessons and Advice for AI Builders Matt asks for final parting advice for AI builders. Mitch advises founders to look past daily Twitter hype and base technical strategy on fundamental, enduring paradigm shifts.1:24–4:22 · Guest disagreement 1/10 Voice Interfaces and Whispering Context to AI Agents Matt sets the stage by citing a specific scene reported by Stephanie Palazzolo at The Information about Basis employees whispering into mics. Mitch explains that whispering gives high-bandwidth context without summarization loss, which agents prefer over concise human text.4:22–6:24 · Guest disagreement 1/10 Why Accounting is a Crucial Knowledge Work Domain Matt asks if Mitch chose accounting because of agent technical challenges or market need. Mitch delivers a thoughtful breakdown of accounting as information compression over modern capitalism.6:24–8:45 · Guest disagreement 2/10 Defining Agency and Long-Horizon Agentic Execution Matt asks for a concrete definition of long-horizon agency. Mitch reframes agency as a spectrum of decision-making autonomy constrained by LLM working memory limits.8:45–11:18 · Guest disagreement 1/10 End-to-End Tax Return Workflows and Reviewer Trust Matt asks what autonomous execution looks like in Basis tax return workflows. Mitch explains that true autonomy produces reviewable assumptions and split artifacts rather than uninspectable end states.11:18–14:15 · Guest disagreement 1/10 The History of Agents: ReAct Framework and State Management Matt traces the history of agents back to the 2022 ReAct paper and connects it to Christopher Nolan films. Mitch extends the framework using the movie Memento to illustrate how agents must write notes to regulate their own future inference states.14:15–18:24 · Guest disagreement 1/10 BabyAGI and the Problem of Compounding Errors Matt brings up BabyAGI and the compounding error problem from 2023. Mitch outlines model history, explaining how Opus 3, o1, and o3 overcame attention degradation over long context windows.18:24–20:32 · Guest disagreement 1/10 Process Supervision vs. Outcome Supervision in RL Matt cites OpenAI's 2023 'Let's Verify Step by Step' paper and its 800,000 human-labeled reasoning steps. Mitch explains the tradeoff between costly process supervision and outcome supervision scaled via RLVR in DeepSeek R1.20:32–25:15 · Guest disagreement 3/10 METR Benchmarks and the Unique Properties of Coding Agents Matt asks if METR benchmark doubling rates hold true. Mitch questions METR's small sample size and explains why coding agents succeed due to runtime execution feedback and rich training data rather than purely verifiable rewards.25:15–31:26 · Guest disagreement 2/10 Why Real-World AI Agents Struggle Outside of Coding Matt asks why agents struggle outside software development. Mitch highlights the lack of runtime compilers, the need to map human organizational checks, and data privacy hurdles in real-world professions.31:26–36:37 · Guest disagreement 3/10 Translating Human Organizational Processes to Agent Design Matt pushes on why outcome evals are insufficient. Mitch explains that an agent passing 100 evals via Wikipedia is unemployable by an accounting firm because process integrity and primary-source citations are mandatory.36:37–42:17 · Guest disagreement 1/10 Behavior Specs: Defining and Evaluating Expected Agent Behavior Matt asks about the structure and authorship of behavior specs in markdown. Mitch explains that meta behaviors bridge subjective product standards and machine execution, co-authored by accountants and ML researchers.42:17–50:04 · Guest disagreement 3/10 Context Engineering, LLM Judges, and Organizational Reliability Matt raises a critical counterargument: does enforcing rigid human processes preclude emergent Move 37 breakthroughs? Mitch responds that enterprise clients purchase predictable reliability and auditability, not Move 37 surprises.50:04–52:42 · Guest disagreement 2/10 Building Intuition Around Model Capabilities and Behavior Matt asks how developers design systems when underlying LLM internals remain opaque. Mitch explains treating the model as a black-box alien entity with activation states that require uncorrelated review trajectories.52:42–54:57 · Guest disagreement 4/10 Developing Practical Intuition for Agent Systems Matt asks if agent intuition stems from reading research papers and talking to lab researchers. Mitch flatly dismisses that premise, arguing that real intuition comes exclusively from root-cause debugging one's own workflow automations.54:57–59:45 · Guest disagreement 1/10 Open-Sourcing the Agent Behavior Standard with Braintrust Matt discusses the open-source collaboration with Braintrust and mentions the 4 AM Sunday Slack exchanges between Basis co-founders. Mitch explains the open-source standard for defining and evaluating agent behavior specs.59:45–1:04:31 · Guest disagreement 2/10 Designing Ontologies and Company Canon for Agents Matt inquires about the practical design of ontologies in agent systems. Mitch differentiates coding agents from domain agents, noting that domain agents must own and structure their own runtime training context over long horizons.1:04:31–1:06:34 · Guest disagreement 1/10 Managing Canonical Knowledge in Agent-Native Organizations Matt quotes Mitch on treating documentation with the same rigor as code bases where deleting a paragraph breaks the system. Mitch details how organizations must maintain singular canonical knowledge rather than fragmented Slack or Gong history.1:06:34–1:09:06 · Guest disagreement 1/10 Emerging Roles: Language Architects and Context Engineers Matt asks about hiring for novel roles like language architects and context engineers. Mitch compares writing durable system prompts to drafting constitutional law that must cleanly withstand thousands of runtime interpretations.1:09:06–1:11:23 · Guest disagreement 2/10 The Role of the Deployed Intelligence Team Matt asks about the deployed intelligence team concept. Mitch clarifies that DI engineers are neither forward-deployed software engineers nor agent PMs, but workflow transformation specialists embedded with accounting firms.1:11:23–1:14:31 · Guest disagreement 4/10 System-Level Self-Improvement and Closing the Agent Feedback Loop Matt explores system-level self-improvement loops. Mitch critiques engineers who obsess over pristine code abstraction while leaving prompt context full of slop, stressing that context directly governs runtime performance.1:14:31–1:17:03 · Guest disagreement 2/10 Reinforcement Learning, Reward Functions, and Model-Level Improvement Matt asks whether Basis plans to transition from harness engineering to fine-tuning model weights via RL. Mitch explains that developing the reward function and signal allocation is the primary challenge, regardless of where signal is applied.1:17:03–1:21:04 · Guest disagreement 4/10 The Bitter Lesson, Technical Moats, and Business Strategy Matt presses Mitch on whether the Bitter Lesson will swallow harness engineering and if applied AI firms will need proprietary RL moats. Mitch emphatically rejects technical moats, arguing business distribution and workflow embedding determine enterprise value.1:21:04–1:22:30 · Guest disagreement 1/10 Final Lessons and Advice for AI Builders Matt asks for final parting advice for AI builders. Mitch advises founders to look past daily Twitter hype and base technical strategy on fundamental, enduring paradigm shifts.1:24–4:22 · Matt pushing back 1/10 Voice Interfaces and Whispering Context to AI Agents Matt sets the stage by citing a specific scene reported by Stephanie Palazzolo at The Information about Basis employees whispering into mics. Mitch explains that whispering gives high-bandwidth context without summarization loss, which agents prefer over concise human text.4:22–6:24 · Matt pushing back 1/10 Why Accounting is a Crucial Knowledge Work Domain Matt asks if Mitch chose accounting because of agent technical challenges or market need. Mitch delivers a thoughtful breakdown of accounting as information compression over modern capitalism.6:24–8:45 · Matt pushing back 1/10 Defining Agency and Long-Horizon Agentic Execution Matt asks for a concrete definition of long-horizon agency. Mitch reframes agency as a spectrum of decision-making autonomy constrained by LLM working memory limits.8:45–11:18 · Matt pushing back 1/10 End-to-End Tax Return Workflows and Reviewer Trust Matt asks what autonomous execution looks like in Basis tax return workflows. Mitch explains that true autonomy produces reviewable assumptions and split artifacts rather than uninspectable end states.11:18–14:15 · Matt pushing back 1/10 The History of Agents: ReAct Framework and State Management Matt traces the history of agents back to the 2022 ReAct paper and connects it to Christopher Nolan films. Mitch extends the framework using the movie Memento to illustrate how agents must write notes to regulate their own future inference states.14:15–18:24 · Matt pushing back 1/10 BabyAGI and the Problem of Compounding Errors Matt brings up BabyAGI and the compounding error problem from 2023. Mitch outlines model history, explaining how Opus 3, o1, and o3 overcame attention degradation over long context windows.18:24–20:32 · Matt pushing back 1/10 Process Supervision vs. Outcome Supervision in RL Matt cites OpenAI's 2023 'Let's Verify Step by Step' paper and its 800,000 human-labeled reasoning steps. Mitch explains the tradeoff between costly process supervision and outcome supervision scaled via RLVR in DeepSeek R1.20:32–25:15 · Matt pushing back 2/10 METR Benchmarks and the Unique Properties of Coding Agents Matt asks if METR benchmark doubling rates hold true. Mitch questions METR's small sample size and explains why coding agents succeed due to runtime execution feedback and rich training data rather than purely verifiable rewards.25:15–31:26 · Matt pushing back 1/10 Why Real-World AI Agents Struggle Outside of Coding Matt asks why agents struggle outside software development. Mitch highlights the lack of runtime compilers, the need to map human organizational checks, and data privacy hurdles in real-world professions.31:26–36:37 · Matt pushing back 2/10 Translating Human Organizational Processes to Agent Design Matt pushes on why outcome evals are insufficient. Mitch explains that an agent passing 100 evals via Wikipedia is unemployable by an accounting firm because process integrity and primary-source citations are mandatory.36:37–42:17 · Matt pushing back 2/10 Behavior Specs: Defining and Evaluating Expected Agent Behavior Matt asks about the structure and authorship of behavior specs in markdown. Mitch explains that meta behaviors bridge subjective product standards and machine execution, co-authored by accountants and ML researchers.42:17–50:04 · Matt pushing back 3/10 Context Engineering, LLM Judges, and Organizational Reliability Matt raises a critical counterargument: does enforcing rigid human processes preclude emergent Move 37 breakthroughs? Mitch responds that enterprise clients purchase predictable reliability and auditability, not Move 37 surprises.50:04–52:42 · Matt pushing back 1/10 Building Intuition Around Model Capabilities and Behavior Matt asks how developers design systems when underlying LLM internals remain opaque. Mitch explains treating the model as a black-box alien entity with activation states that require uncorrelated review trajectories.52:42–54:57 · Matt pushing back 1/10 Developing Practical Intuition for Agent Systems Matt asks if agent intuition stems from reading research papers and talking to lab researchers. Mitch flatly dismisses that premise, arguing that real intuition comes exclusively from root-cause debugging one's own workflow automations.54:57–59:45 · Matt pushing back 1/10 Open-Sourcing the Agent Behavior Standard with Braintrust Matt discusses the open-source collaboration with Braintrust and mentions the 4 AM Sunday Slack exchanges between Basis co-founders. Mitch explains the open-source standard for defining and evaluating agent behavior specs.59:45–1:04:31 · Matt pushing back 1/10 Designing Ontologies and Company Canon for Agents Matt inquires about the practical design of ontologies in agent systems. Mitch differentiates coding agents from domain agents, noting that domain agents must own and structure their own runtime training context over long horizons.1:04:31–1:06:34 · Matt pushing back 1/10 Managing Canonical Knowledge in Agent-Native Organizations Matt quotes Mitch on treating documentation with the same rigor as code bases where deleting a paragraph breaks the system. Mitch details how organizations must maintain singular canonical knowledge rather than fragmented Slack or Gong history.1:06:34–1:09:06 · Matt pushing back 1/10 Emerging Roles: Language Architects and Context Engineers Matt asks about hiring for novel roles like language architects and context engineers. Mitch compares writing durable system prompts to drafting constitutional law that must cleanly withstand thousands of runtime interpretations.1:09:06–1:11:23 · Matt pushing back 1/10 The Role of the Deployed Intelligence Team Matt asks about the deployed intelligence team concept. Mitch clarifies that DI engineers are neither forward-deployed software engineers nor agent PMs, but workflow transformation specialists embedded with accounting firms.1:11:23–1:14:31 · Matt pushing back 1/10 System-Level Self-Improvement and Closing the Agent Feedback Loop Matt explores system-level self-improvement loops. Mitch critiques engineers who obsess over pristine code abstraction while leaving prompt context full of slop, stressing that context directly governs runtime performance.1:14:31–1:17:03 · Matt pushing back 1/10 Reinforcement Learning, Reward Functions, and Model-Level Improvement Matt asks whether Basis plans to transition from harness engineering to fine-tuning model weights via RL. Mitch explains that developing the reward function and signal allocation is the primary challenge, regardless of where signal is applied.1:17:03–1:21:04 · Matt pushing back 2/10 The Bitter Lesson, Technical Moats, and Business Strategy Matt presses Mitch on whether the Bitter Lesson will swallow harness engineering and if applied AI firms will need proprietary RL moats. Mitch emphatically rejects technical moats, arguing business distribution and workflow embedding determine enterprise value.1:21:04–1:22:30 · Matt pushing back 0/10 Final Lessons and Advice for AI Builders Matt asks for final parting advice for AI builders. Mitch advises founders to look past daily Twitter hype and base technical strategy on fundamental, enduring paradigm shifts.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 42.5% · guest 57.5%0:00 · Matt 42.5% · guest 57.5%3:00 · Matt 24.1% · guest 75.9%3:00 · Matt 24.1% · guest 75.9%6:00 · Matt 23.7% · guest 76.3%6:00 · Matt 23.7% · guest 76.3%9:00 · Matt 23.8% · guest 76.2%9:00 · Matt 23.8% · guest 76.2%12:00 · Matt 29% · guest 71%12:00 · Matt 29% · guest 71%15:00 · Matt 8.3% · guest 91.7%15:00 · Matt 8.3% · guest 91.7%18:00 · Matt 17.2% · guest 82.8%18:00 · Matt 17.2% · guest 82.8%21:00 · Matt 17.8% · guest 82.2%21:00 · Matt 17.8% · guest 82.2%24:00 · Matt 17.9% · guest 82.1%24:00 · Matt 17.9% · guest 82.1%27:00 · Matt 18.6% · guest 81.4%27:00 · Matt 18.6% · guest 81.4%30:00 · Matt 10.7% · guest 89.3%30:00 · Matt 10.7% · guest 89.3%33:00 · Matt 15.9% · guest 84.1%33:00 · Matt 15.9% · guest 84.1%36:00 · Matt 10.5% · guest 89.5%36:00 · Matt 10.5% · guest 89.5%39:00 · Matt 11.3% · guest 88.7%39:00 · Matt 11.3% · guest 88.7%42:00 · Matt 15.9% · guest 84.1%42:00 · Matt 15.9% · guest 84.1%45:00 · Matt 19.8% · guest 80.2%45:00 · Matt 19.8% · guest 80.2%48:00 · Matt 22.3% · guest 77.7%48:00 · Matt 22.3% · guest 77.7%51:00 · Matt 7.1% · guest 92.9%51:00 · Matt 7.1% · guest 92.9%54:00 · Matt 23.5% · guest 76.5%54:00 · Matt 23.5% · guest 76.5%57:00 · Matt 13.1% · guest 86.9%57:00 · Matt 13.1% · guest 86.9%1:00:00 · Matt 6.5% · guest 93.5%1:00:00 · Matt 6.5% · guest 93.5%1:03:00 · Matt 12.7% · guest 87.3%1:03:00 · Matt 12.7% · guest 87.3%1:06:00 · Matt 6.8% · guest 93.2%1:06:00 · Matt 6.8% · guest 93.2%1:09:00 · Matt 28% · guest 72%1:09:00 · Matt 28% · guest 72%1:12:00 · Matt 13.4% · guest 86.6%1:12:00 · Matt 13.4% · guest 86.6%1:15:00 · Matt 8.9% · guest 91.1%1:15:00 · Matt 8.9% · guest 91.1%1:18:00 · Matt 8.8% · guest 91.2%1:18:00 · Matt 8.8% · guest 91.2%1:21:00 · Matt 37.7% · guest 62.3%1:21:00 · Matt 37.7% · guest 62.3%
Sharpest disagreement ▶ 1:19:03 Rejecting proprietary technical moats

Mitch forcefully rejects the host's premise that applied AI companies will survive on proprietary RL tricks, stating flatly that technical moats are not real moats.

Hardest push from Matt ▶ 46:45 Challenging process specs against Move 37 innovation

Matt directly challenges Mitch's process-supervision philosophy by arguing it restricts agents from achieving non-human Move 37 efficiencies.

Biggest teaching moment ▶ 1:13:05 English context vs code hygiene schooling

Mitch schools the engineering community on misplaced priorities, explaining that natural language context directly dictates runtime model execution while code formatting does not.

Matt holds his own ▶ 18:24 Citing OpenAI's 800k step-by-step verification paper

Matt demonstrates deep technical domain knowledge by citing OpenAI's specific 2023 paper and its exact volume of 800,000 human-labeled reasoning steps.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Voice Interfaces and Whispering Context to AI Agents 4411 Matt sets the stage by citing a specific scene reported by Stephanie Palazzolo at The Information about Basis employees whispering into mics. Mitch explains that whispering gives high-bandwidth context without summarization loss, which agents prefer over concise human text.
Why Accounting is a Crucial Knowledge Work Domain 3511 Matt asks if Mitch chose accounting because of agent technical challenges or market need. Mitch delivers a thoughtful breakdown of accounting as information compression over modern capitalism.
Defining Agency and Long-Horizon Agentic Execution 3621 Matt asks for a concrete definition of long-horizon agency. Mitch reframes agency as a spectrum of decision-making autonomy constrained by LLM working memory limits.
End-to-End Tax Return Workflows and Reviewer Trust 3511 Matt asks what autonomous execution looks like in Basis tax return workflows. Mitch explains that true autonomy produces reviewable assumptions and split artifacts rather than uninspectable end states.
The History of Agents: ReAct Framework and State Management 5511 Matt traces the history of agents back to the 2022 ReAct paper and connects it to Christopher Nolan films. Mitch extends the framework using the movie Memento to illustrate how agents must write notes to regulate their own future inference states.
BabyAGI and the Problem of Compounding Errors 5511 Matt brings up BabyAGI and the compounding error problem from 2023. Mitch outlines model history, explaining how Opus 3, o1, and o3 overcame attention degradation over long context windows.
Process Supervision vs. Outcome Supervision in RL 6611 Matt cites OpenAI's 2023 'Let's Verify Step by Step' paper and its 800,000 human-labeled reasoning steps. Mitch explains the tradeoff between costly process supervision and outcome supervision scaled via RLVR in DeepSeek R1.
METR Benchmarks and the Unique Properties of Coding Agents 5632 Matt asks if METR benchmark doubling rates hold true. Mitch questions METR's small sample size and explains why coding agents succeed due to runtime execution feedback and rich training data rather than purely verifiable rewards.
Why Real-World AI Agents Struggle Outside of Coding 4621 Matt asks why agents struggle outside software development. Mitch highlights the lack of runtime compilers, the need to map human organizational checks, and data privacy hurdles in real-world professions.
Translating Human Organizational Processes to Agent Design 5732 Matt pushes on why outcome evals are insufficient. Mitch explains that an agent passing 100 evals via Wikipedia is unemployable by an accounting firm because process integrity and primary-source citations are mandatory.
Behavior Specs: Defining and Evaluating Expected Agent Behavior 4612 Matt asks about the structure and authorship of behavior specs in markdown. Mitch explains that meta behaviors bridge subjective product standards and machine execution, co-authored by accountants and ML researchers.
Context Engineering, LLM Judges, and Organizational Reliability 6733 Matt raises a critical counterargument: does enforcing rigid human processes preclude emergent Move 37 breakthroughs? Mitch responds that enterprise clients purchase predictable reliability and auditability, not Move 37 surprises.
Building Intuition Around Model Capabilities and Behavior 4621 Matt asks how developers design systems when underlying LLM internals remain opaque. Mitch explains treating the model as a black-box alien entity with activation states that require uncorrelated review trajectories.
Developing Practical Intuition for Agent Systems 3741 Matt asks if agent intuition stems from reading research papers and talking to lab researchers. Mitch flatly dismisses that premise, arguing that real intuition comes exclusively from root-cause debugging one's own workflow automations.
Open-Sourcing the Agent Behavior Standard with Braintrust 5511 Matt discusses the open-source collaboration with Braintrust and mentions the 4 AM Sunday Slack exchanges between Basis co-founders. Mitch explains the open-source standard for defining and evaluating agent behavior specs.
Designing Ontologies and Company Canon for Agents 4621 Matt inquires about the practical design of ontologies in agent systems. Mitch differentiates coding agents from domain agents, noting that domain agents must own and structure their own runtime training context over long horizons.
Managing Canonical Knowledge in Agent-Native Organizations 5611 Matt quotes Mitch on treating documentation with the same rigor as code bases where deleting a paragraph breaks the system. Mitch details how organizations must maintain singular canonical knowledge rather than fragmented Slack or Gong history.
Emerging Roles: Language Architects and Context Engineers 4611 Matt asks about hiring for novel roles like language architects and context engineers. Mitch compares writing durable system prompts to drafting constitutional law that must cleanly withstand thousands of runtime interpretations.
The Role of the Deployed Intelligence Team 5521 Matt asks about the deployed intelligence team concept. Mitch clarifies that DI engineers are neither forward-deployed software engineers nor agent PMs, but workflow transformation specialists embedded with accounting firms.
System-Level Self-Improvement and Closing the Agent Feedback Loop 4741 Matt explores system-level self-improvement loops. Mitch critiques engineers who obsess over pristine code abstraction while leaving prompt context full of slop, stressing that context directly governs runtime performance.
Reinforcement Learning, Reward Functions, and Model-Level Improvement 5621 Matt asks whether Basis plans to transition from harness engineering to fine-tuning model weights via RL. Mitch explains that developing the reward function and signal allocation is the primary challenge, regardless of where signal is applied.
The Bitter Lesson, Technical Moats, and Business Strategy 6742 Matt presses Mitch on whether the Bitter Lesson will swallow harness engineering and if applied AI firms will need proprietary RL moats. Mitch emphatically rejects technical moats, arguing business distribution and workflow embedding determine enterprise value.
Final Lessons and Advice for AI Builders 3510 Matt asks for final parting advice for AI builders. Mitch advises founders to look past daily Twitter hype and base technical strategy on fundamental, enduring paradigm shifts.

Statements from this episode (38)

Insight
Trojanowski: AI agents prefer raw spoken thoughts over summarized text
“Speaking is just so much faster than writing things down, and in fact, when you try to write things down, you are Actually, essentially trying to summarize all the crazy thoughts in your head, and so that's why it takes a lot of time. And it's useful for your …”
Mitch Trojanowski Aug 5, 2026 ▶ 2:15
Insight
Trojanowski: Accounting Automation Requires Long-Horizon Coherence Across Many Actions
“Accounting is difficult. It's not something that is just purely text in text out. And so it requires the ability for like AIs to be able to, you know, perform lots of actions over long periods of time and actually you know, be coherent over that period of time…”
Mitch Trojanowski Aug 5, 2026 ▶ 3:49
Assertion Supported
Trojanowski: Over Three Million Accountants Work in the US
“It is, one of, if not the largest knowledge work profession in the country. There are over three million, you know, combined kind of accountants in the country.”
Mitch Trojanowski Aug 5, 2026 ▶ 4:31
Insight
Trojanowski: Accounting Acts as Information Compression Over the Economy
“And accounting is actually the art of compressing all of that into something that is structured that now people can look at and understand and make decisions. So something about accounting you could argue in a meta way is kind of like an intelligence over the …”
Mitch Trojanowski Aug 5, 2026 ▶ 5:52
Insight
Trojanowski: Autonomous agents must optimize for reviewer legibility, not just completion
“And you can like optimize, not just for getting the work done, but for making it easy for your reviewer to understand the decisions that you made. And that's obviously very true in software engineering, and it's actually true in, I think, most professions and …”
Mitch Trojanowski Aug 5, 2026 ▶ 10:58
Insight
Trojanowski: Long-Horizon Agents Must Use Reasoning for State Management
“Once you start getting to longer horizons where you're beyond the context window or you're getting to context rot, you need to kind of use in sort of brute force your reasoning to build out your environment. Whether that be with sub agents or compaction, we ta…”
Mitch Trojanowski Aug 5, 2026 ▶ 13:45
Opinion
Trojanowski: Claude 3 Opus Understood Long Contexts That GPT-4 Turbo Failed
“I would say they were Opus III, which I think goes underappreciated, but I think was the first model to truly be able to like actually understand that long context. Before, if you put anything in the, like, 80,000 tokens in the GPT-IV Turbo, it could not under…”
Mitch Trojanowski Aug 5, 2026 ▶ 16:16
Insight
Trojanowski: OpenAI o3 proved post-training improves per-token reasoning quality
“And then I think after a one, it was oh three, because I think oh three helped prove that not only could you scale the amount of reasoning at inference time, but with better training, with more compute, better data, et cetera, in the post-training phase, you c…”
Mitch Trojanowski Aug 5, 2026 ▶ 16:40
Assertion Supported
Trojanowski: DeepSeek-R1 succeeded by scaling outcome supervision over process supervision
“If you look, you know, if you fast forward a bit and you look at the, like the DeepSeq R-one paper where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing in terms of you know, RLVR reasoning f…”
Mitch Trojanowski Aug 5, 2026 ▶ 20:03
Opinion
Trojanowski: METR AI Agent Benchmark Is Inaccurate With a Low Bar
“I think the meter chart is somewhat inaccurate these days because it's so hard to measure. And also the bar is pretty low, right? The bar is pretty low. Yeah, yeah. The bar is pretty low, and I think, I'm not super familiar with their dataset, but my understan…”
Mitch Trojanowski Aug 5, 2026 ▶ 20:48
Insight
Trojanowski: Verifiable rewards cannot solve software engineering because coding is subjective
“At the end of the day, coding is subjective. It's an art and you're not going to solve an art through verifiable rewards.”
Mitch Trojanowski Aug 5, 2026 ▶ 24:26
Assertion Not checkable as stated
Trojanowski: AI coding agents remain worse than junior engineers over two-week projects
“Even now with coding, like, the agents are not yet they're not human level at being coherent over long periods of time. That's obvious because they can't code like a junior engineer on a project for two weeks. So they can't, they're, that's worse than a human …”
Mitch Trojanowski Aug 5, 2026 ▶ 26:35
Insight
Trojanowski: Good AI agent design should replicate how human organizations verify work
“Humans are already used to working with non-deterministic systems. It's just the systems are normally their coworkers, not their computers. And in many ways, like companies and processes, it's all about how do you design a system for non-deterministic entities…”
Mitch Trojanowski Aug 5, 2026 ▶ 27:47
Assertion Supported
Trojanowski: Preparing a Form 1065 tax return takes 20+ hours of human labor
“Performing a 1065 can take a human, you know like 20 plus hours easily of actual work. I don't mean like it took them a day. I meant like literal sitting down work. And it could actually take much longer for very complicated returns.”
Mitch Trojanowski Aug 5, 2026 ▶ 30:54
Opinion
Trojanowski: Applying the Bitter Lesson to Complex Tax Workflows Is a Mistake
“I think it is a mistake to throw out those learnings and say, you know, bitter lesson, throw out those learnings. We're just going to have the agents at runtime develop an entirely new way to do a tax return that is, you know, because bitter lesson, yada, yada…”
Mitch Trojanowski Aug 5, 2026 ▶ 34:40
Insight
Trojanowski: Better to Give AI Agents Principles and Context Than Step-by-Step Instructions
“As people who are good at building agents know, like it's much better at the margins to be able to give principles and the whys and more context and let them figure it out.”
Mitch Trojanowski Aug 5, 2026 ▶ 40:49
Insight
Trojanowski: Context engineering is runtime training with far less data
“I think the framework I like is thinking about it as training data. Except you're just training the model at runtime. It is training data. And because you, the model's learning at inference time, the total amount of training data is far lower, right? Like the …”
Mitch Trojanowski Aug 5, 2026 ▶ 43:17
Disclosure
Trojanowski: Basis uses full AI agents as behavior evaluation judges
“So right now, at least the behavior judging is relatively expensive because it's a pretty advanced judge in that it is also an agent. It's not a judge in the traditional sense. It's literally an agent because it has to look at the trajectory.”
Mitch Trojanowski Aug 5, 2026 ▶ 45:36
Opinion
Trojanowski: Enterprise agents need process reliability, not 'Move 37' breakthroughs
“And so the thing that you know, someone is buying from us is not, this will be the best ever tax return. They're buying that, you know, the confidence... They're buying that it's going to be consistent and reliable and something that they can trust that actual…”
Mitch Trojanowski Aug 5, 2026 ▶ 49:24
Insight
Trojanowski: LLM agent review requires uncorrelated trajectories to avoid bias
“So there's some activation state that by definition is going to be biased to that current trajectory. And so maybe for review, you want an uncorrelated trajectory, right? Where it's like a new box and it's just a smart.”
Mitch Trojanowski Aug 5, 2026 ▶ 52:09
Insight
Trojanowski: Frontier agent architectures combine LLM intuition with org design
“If you build these kind of LLM intuitions and you combine them with, you know, maybe basic principles of, like, organizational design and management I think you start to get to maybe what is, like, the frontier of agent building.”
Mitch Trojanowski Aug 5, 2026 ▶ 52:31
Opinion
Trojanowski: Nothing Paradigm-Shifting Has Changed Since o3
“Things have really not changed since oh three, I would say. Almost everything since oh three has been relatively on, I don't want to say on trend and then like I knew this exact trend, but I would say it's all within the same paradigm. Like nothing paradigm sh…”
Mitch Trojanowski Aug 5, 2026 ▶ 53:22
Insight
Trojanowski: Best Agent Intuition Comes From Automating Your Own Work
“A lot of the people with the best agent intuition, actually, yes, a lot of people come from ML backgrounds, but people who don't, a lot of them are ones who are just really good at automating their own work.”
Mitch Trojanowski Aug 5, 2026 ▶ 54:21
Insight
Trojanowski: Agent Behavior Specs Align Human Teams, Not Just Models
“It is actually both a spec and a rubric. We call this specs and there were some people who asked, isn't this a rubric? And it is, it's both. The reason it's both is because it is not just used to grade or potentially reward the agent. It's also used to align t…”
Mitch Trojanowski Aug 5, 2026 ▶ 55:12
Insight
Trojanowski: 10-Hour Agent Runs Are Not Black Boxes
“Not thinking that an agent operating over 10 hours is a black box. It's not. It has a lot of data, and you're probably doing a disservice to your customers if you don't understand, like, how it's going about the work.”
Mitch Trojanowski Aug 5, 2026 ▶ 58:07
Insight
Trojanowski: Codebase ontology matters as much for AI agents as for engineers
“The ontology of your code base, it always mattered for engineers. It matters just as much, if not more, for really good agents over time.”
Mitch Trojanowski Aug 5, 2026 ▶ 1:01:12
Insight
Trojanowski: Cheaper inference is replacing graphs and embeddings for agent ontologies
“As models are getting cheaper and cheaper, more and more of that actually can just be done using inference instead of using determined, like things like graphs or things like embeddings.”
Mitch Trojanowski Aug 5, 2026 ▶ 1:03:30
Insight
Trojanowski: Maintaining company canon will be a key human role in agent-native firms
“And that's why I think, you know, for true agent native companies especially in the future today, I think it's still quite early, but especially in the future, having a clear understanding of what your company canon is and organizing that in an ontology that m…”
Mitch Trojanowski Aug 5, 2026 ▶ 1:06:09
Insight
Trojanowski: Key skill for agent roles is systems thinking across abstractions
“I think one thing, one skill that really matters is good systems thinking. And where does good systems thinking come from? It comes from people who have had to think about some abstraction, some system, something, and design it in such a way that it performs i…”
Mitch Trojanowski Aug 5, 2026 ▶ 1:07:03
Opinion
Trojanowski: Most engineering historically was execution-oriented, not systems thinking
“Now, I think the majority of engineering historically has not been really systems thinking based. You know, it's been a little bit more execution oriented, but if you think about like the hardest engineering, like, Hey, I'm trying to design like what this is g…”
Mitch Trojanowski Aug 5, 2026 ▶ 1:07:36
Insight
Trojanowski: Legal drafting requires the same skills as AI context engineering
“I think law actually is kind of like that. You know, in many ways I think about, I think maybe the founding fathers would have been really good context engineers or agent managers, because you had to write, you know, a piece of English that was going to be, yo…”
Mitch Trojanowski Aug 5, 2026 ▶ 1:07:52
Prediction Not checkable as stated
Trojanowski: Closing the agent self-improvement loop will be close by year-end
“I think that closing the loop is gonna happen pretty fast. I think you'll have, like, I don't know about the entire loop being closed, but I think you'll be Relatively close by end of year.”
Mitch Trojanowski Aug 5, 2026 ▶ 1:12:02
Assertion Not checkable as stated
Trojanowski: AI models are far worse at agent engineering than software
“Right now they are far, far, far worse at engineering agent systems than they are at engineering most software. Far worse. because by definition, like that kind of work, which is so novel, has not seen a large amount in their training data. and so they have …”
Mitch Trojanowski Aug 5, 2026 ▶ 1:12:35
Insight
Trojanowski: English prompt context matters more than code quality in agents
“A lot of engineers, they treat the code as more precious than the English, when actually the English is more precious because the English affects the performance. The code does not affect the performance, right? If the logic, assuming the logic's the same, it …”
Mitch Trojanowski Aug 5, 2026 ▶ 1:13:18
Disclosure
Trojanowski: Basis Prioritizes Agent Orchestration Over Model Weight Updates
“Today we don't go directly into the weights. And primarily the reason we don't do that is because A lot of the advancements the models are having when it comes to orchestrating themselves yield far more performance gains than benefits you would have of like up…”
Mitch Trojanowski Aug 5, 2026 ▶ 1:15:49
Prediction Not checkable as stated
Trojanowski: Scaling Verifiable Rewards Won't Reliably Automate Tax Returns
“I don't believe that if you were to train a model, you know, and you scale up the amount of pre-training computing, you scale up the amount of Like post training from like perfectly verifiable rewards that suddenly will output a model that will do a tax return…”
Mitch Trojanowski Aug 5, 2026 ▶ 1:16:37
Prediction Not checkable as stated
Trojanowski: AI models will subsume agent harness engineering in 2-5 years
“I think that how long it will take to get swallowed up. I don't know exactly. You know, I think it's probably sub five years. I don't think it's sub two years. I think it's probably sub five years.”
Mitch Trojanowski Aug 5, 2026 ▶ 1:17:48
Insight
Trojanowski: Technical AI moats are not real moats
“Technical moats are not real moats. Like, I, there's no portion of basis's long-term terminal value that stems from some, you know, secret RL trick we found that nobody else found.”
Mitch Trojanowski Aug 5, 2026 ▶ 1:19:08
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.