Aug 5, 2026 · 1h 22m · mad
How to Build Autonomous, Long-Horizon AI Agents | Basis
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews Basis co-founder Mitchell Trojanowski about the engineering, architectural, and organizational principles required to build reliable, long-horizon autonomous AI agents for complex domains like accounting. Trojanowski breaks down key concepts including state management, process supervision, open-source behavior specifications, context engineering, and the translation of human organizational workflows into scalable agent architectures.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.4% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Mitch forcefully rejects the host's premise that applied AI companies will survive on proprietary RL tricks, stating flatly that technical moats are not real moats.
Hardest push from Matt ▶ 46:45 Challenging process specs against Move 37 innovationMatt directly challenges Mitch's process-supervision philosophy by arguing it restricts agents from achieving non-human Move 37 efficiencies.
Biggest teaching moment ▶ 1:13:05 English context vs code hygiene schoolingMitch schools the engineering community on misplaced priorities, explaining that natural language context directly dictates runtime model execution while code formatting does not.
Matt holds his own ▶ 18:24 Citing OpenAI's 800k step-by-step verification paperMatt demonstrates deep technical domain knowledge by citing OpenAI's specific 2023 paper and its exact volume of 800,000 human-labeled reasoning steps.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Voice Interfaces and Whispering Context to AI Agents | 4 | 4 | 1 | 1 | Matt sets the stage by citing a specific scene reported by Stephanie Palazzolo at The Information about Basis employees whispering into mics. Mitch explains that whispering gives high-bandwidth context without summarization loss, which agents prefer over concise human text. | |
| Why Accounting is a Crucial Knowledge Work Domain | 3 | 5 | 1 | 1 | Matt asks if Mitch chose accounting because of agent technical challenges or market need. Mitch delivers a thoughtful breakdown of accounting as information compression over modern capitalism. | |
| Defining Agency and Long-Horizon Agentic Execution | 3 | 6 | 2 | 1 | Matt asks for a concrete definition of long-horizon agency. Mitch reframes agency as a spectrum of decision-making autonomy constrained by LLM working memory limits. | |
| End-to-End Tax Return Workflows and Reviewer Trust | 3 | 5 | 1 | 1 | Matt asks what autonomous execution looks like in Basis tax return workflows. Mitch explains that true autonomy produces reviewable assumptions and split artifacts rather than uninspectable end states. | |
| The History of Agents: ReAct Framework and State Management | 5 | 5 | 1 | 1 | Matt traces the history of agents back to the 2022 ReAct paper and connects it to Christopher Nolan films. Mitch extends the framework using the movie Memento to illustrate how agents must write notes to regulate their own future inference states. | |
| BabyAGI and the Problem of Compounding Errors | 5 | 5 | 1 | 1 | Matt brings up BabyAGI and the compounding error problem from 2023. Mitch outlines model history, explaining how Opus 3, o1, and o3 overcame attention degradation over long context windows. | |
| Process Supervision vs. Outcome Supervision in RL | 6 | 6 | 1 | 1 | Matt cites OpenAI's 2023 'Let's Verify Step by Step' paper and its 800,000 human-labeled reasoning steps. Mitch explains the tradeoff between costly process supervision and outcome supervision scaled via RLVR in DeepSeek R1. | |
| METR Benchmarks and the Unique Properties of Coding Agents | 5 | 6 | 3 | 2 | Matt asks if METR benchmark doubling rates hold true. Mitch questions METR's small sample size and explains why coding agents succeed due to runtime execution feedback and rich training data rather than purely verifiable rewards. | |
| Why Real-World AI Agents Struggle Outside of Coding | 4 | 6 | 2 | 1 | Matt asks why agents struggle outside software development. Mitch highlights the lack of runtime compilers, the need to map human organizational checks, and data privacy hurdles in real-world professions. | |
| Translating Human Organizational Processes to Agent Design | 5 | 7 | 3 | 2 | Matt pushes on why outcome evals are insufficient. Mitch explains that an agent passing 100 evals via Wikipedia is unemployable by an accounting firm because process integrity and primary-source citations are mandatory. | |
| Behavior Specs: Defining and Evaluating Expected Agent Behavior | 4 | 6 | 1 | 2 | Matt asks about the structure and authorship of behavior specs in markdown. Mitch explains that meta behaviors bridge subjective product standards and machine execution, co-authored by accountants and ML researchers. | |
| Context Engineering, LLM Judges, and Organizational Reliability | 6 | 7 | 3 | 3 | Matt raises a critical counterargument: does enforcing rigid human processes preclude emergent Move 37 breakthroughs? Mitch responds that enterprise clients purchase predictable reliability and auditability, not Move 37 surprises. | |
| Building Intuition Around Model Capabilities and Behavior | 4 | 6 | 2 | 1 | Matt asks how developers design systems when underlying LLM internals remain opaque. Mitch explains treating the model as a black-box alien entity with activation states that require uncorrelated review trajectories. | |
| Developing Practical Intuition for Agent Systems | 3 | 7 | 4 | 1 | Matt asks if agent intuition stems from reading research papers and talking to lab researchers. Mitch flatly dismisses that premise, arguing that real intuition comes exclusively from root-cause debugging one's own workflow automations. | |
| Open-Sourcing the Agent Behavior Standard with Braintrust | 5 | 5 | 1 | 1 | Matt discusses the open-source collaboration with Braintrust and mentions the 4 AM Sunday Slack exchanges between Basis co-founders. Mitch explains the open-source standard for defining and evaluating agent behavior specs. | |
| Designing Ontologies and Company Canon for Agents | 4 | 6 | 2 | 1 | Matt inquires about the practical design of ontologies in agent systems. Mitch differentiates coding agents from domain agents, noting that domain agents must own and structure their own runtime training context over long horizons. | |
| Managing Canonical Knowledge in Agent-Native Organizations | 5 | 6 | 1 | 1 | Matt quotes Mitch on treating documentation with the same rigor as code bases where deleting a paragraph breaks the system. Mitch details how organizations must maintain singular canonical knowledge rather than fragmented Slack or Gong history. | |
| Emerging Roles: Language Architects and Context Engineers | 4 | 6 | 1 | 1 | Matt asks about hiring for novel roles like language architects and context engineers. Mitch compares writing durable system prompts to drafting constitutional law that must cleanly withstand thousands of runtime interpretations. | |
| The Role of the Deployed Intelligence Team | 5 | 5 | 2 | 1 | Matt asks about the deployed intelligence team concept. Mitch clarifies that DI engineers are neither forward-deployed software engineers nor agent PMs, but workflow transformation specialists embedded with accounting firms. | |
| System-Level Self-Improvement and Closing the Agent Feedback Loop | 4 | 7 | 4 | 1 | Matt explores system-level self-improvement loops. Mitch critiques engineers who obsess over pristine code abstraction while leaving prompt context full of slop, stressing that context directly governs runtime performance. | |
| Reinforcement Learning, Reward Functions, and Model-Level Improvement | 5 | 6 | 2 | 1 | Matt asks whether Basis plans to transition from harness engineering to fine-tuning model weights via RL. Mitch explains that developing the reward function and signal allocation is the primary challenge, regardless of where signal is applied. | |
| The Bitter Lesson, Technical Moats, and Business Strategy | 6 | 7 | 4 | 2 | Matt presses Mitch on whether the Bitter Lesson will swallow harness engineering and if applied AI firms will need proprietary RL moats. Mitch emphatically rejects technical moats, arguing business distribution and workflow embedding determine enterprise value. | |
| Final Lessons and Advice for AI Builders | 3 | 5 | 1 | 0 | Matt asks for final parting advice for AI builders. Mitch advises founders to look past daily Twitter hype and base technical strategy on fundamental, enduring paradigm shifts. |