Apr 15, 2026 · 1h 25m · latent-space

Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work

Sarah Sachs · 39m spoken Simon Last · 24m spoken Shawn Wang · 6m spoken Alessio Fanelli · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Notion co-founder Simon Last and Head of AI Sarah Sachs join the Latent Space podcast to discuss the engineering architecture, evaluation rigor, and product philosophy behind Notion Custom Agents. They explore how horizontal primitives, autonomous software factories, and robust tool-calling infrastructure are shaping Notion's role as the enterprise system of record.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.1% of the talking time here. How this is scored →

The hosts as informed peer 5.3 Guest teaching 5.9 Guest disagreement 1.5 The hosts pushing back 2.0
05100:0020:0040:001:00:001:20:000:42–5:21 · The hosts as informed peer 4/10 Latent Space Message: Why We Ask for Subscriptions Swyx opens with an ad-free subscriber plea and then presses the guests on why building background agents failed in earlier iterations. Simon and Sarah explain how early context limits and the lack of native tool calling forced multiple redesigns before Sonnet unlocked reliable background execution.5:22–9:34 · The hosts as informed peer 5/10 Balancing AGI Ambitions and the Software Factory Concept Alessio and Swyx ask how Notion plans product roadmaps when frontier model capabilities shift rapidly. Sarah reframes the problem around sensing which direction the technological current is moving rather than fighting upstream against model limitations.9:34–11:57 · The hosts as informed peer 6/10 Horizontal Platform Primitives vs. Vertical SaaS Expertise Alessio contrasts Datadog's vertical persona with Notion's horizontal user base, asking how product expertise is encoded at scale. Sarah and Simon explain that horizontal primitives must stay anchored in concrete user journeys rather than building novel tools for their own sake.11:57–14:44 · The hosts as informed peer 4/10 Fostering an Egoless Engineering Culture of Code Deletion Swyx asks Sarah about her management philosophy on the AI team. Sarah details an egoless engineering culture where engineers are encouraged to throw away old prototypes and rewrite harnesses without tying code ownership to career advancement.14:44–18:07 · The hosts as informed peer 5/10 The 'Simon Vortex' and Hackathons as Capability Uplifters Swyx asks whether regular hackathons drive Notion's product momentum. Sarah and Simon explain that company-wide hackathons serve to uplift baseline technical literacy, while core skunkworks innovation happens continuously in rapid prototyping cycles.18:07–21:53 · The hosts as informed peer 4/10 Organizational Structure and the Design Playground Prototype Repo Alessio asks about prototype evaluation bars across engineering and design. Simon and Sarah describe their isolated design playground repository where designers ship functional code rather than static mockups.21:53–24:03 · The hosts as informed peer 5/10 Scaling Agent Dev Velocity and Decentralized Team Evals Sarah pushes back on Simon's simplified take on prototyping by highlighting the heavy eval and infrastructure burden required before production. She explains that every feature team owns its domain evals while the platform team builds shared harness tooling.24:03–27:24 · The hosts as informed peer 6/10 Detecting Model Provider Regressions and Silent Degradation Swyx asks whether Notion catches silent quality degradation or hidden quantization during peak hours. Sarah shares that enterprise-focused evaluation suites frequently detect regressions that model providers fail to catch with standard coding benchmarks.27:24–30:08 · The hosts as informed peer 5/10 The Rise of Model Behavior Engineers (MBEs) Swyx asks about the Model Behavior Engineer role. Sarah outlines how the position evolved from manual spreadsheet audits by linguistics graduates into engineers who direct coding agents and design LLM judges.30:09–33:24 · The hosts as informed peer 6/10 The Evolution of Software Engineering and Autonomous Workflows Swyx asks if traditional software engineering roles will disappear. Simon and Sarah argue that human engineers are transitioning from syntax implementation to system architecture, verification loops, and delegating to agent fleets.33:24–36:34 · The hosts as informed peer 6/10 Demoing Custom Agents: Automating Tenant Lead Ingestion Alessio demos a live custom agent that automates coworking applicant ingestion, enrichment, and database tracking. Sarah explains how automating tedious internal processes creates high operational leverage across teams.36:34–39:42 · The hosts as informed peer 5/10 Agent Coordination, Manager Agents, and Native Memory Alessio and Swyx ask how agents coordinate without creating infinite recursive loops. Simon explains that Notion relies on database primitives and dedicated manager agents rather than complex bespoke messaging architectures.39:43–43:40 · The hosts as informed peer 6/10 Native Feature Optimizations vs. Third-Party Integrations Swyx asks for Simon's perspective on MCP versus CLI agent interfaces. Simon outlines CLI strengths in progressive disclosure and self-healing while acknowledging MCP's clean sandboxing and permission boundaries.43:40–48:14 · The hosts as informed peer 5/10 Deterministic Code Execution and Cost-Effective Pricing Models Sarah emphasizes that using LLMs to orchestrate deterministic workflows wastes tokens and drives up user costs. She argues that executing deterministic code sandboxes delivers superior reliability and margin efficiency.48:14–52:36 · The hosts as informed peer 7/10 Evolution of Notion's Internal Representation: XML to Markdown and SQLite Swyx asks Simon to walk through the technical iterations of Notion's agent architecture. Simon details the journey from brittle custom XML blocks to markdown and SQLite databases that natively align with model pre-training distributions.52:36–56:25 · The hosts as informed peer 5/10 Decentralizing Tool Ownership and Progressive Disclosure Sarah details how moving away from massive few-shot system prompts to discrete tool definitions allowed Notion to decentralize tool development across independent product squads.56:25–1:00:22 · The hosts as informed peer 5/10 Demystifying System Prompts and Building for Power Users Swyx asks whether Notion protects system prompts as intellectual property. Sarah and Simon push back, explaining that making prompts and tools fully transparent enables power users to build sophisticated workflows.1:00:22–1:02:40 · The hosts as informed peer 5/10 The 'Flippy' Interface: Unifying Settings and Chat Workflows Alessio praises the unified editing and execution canvas. Simon and Sarah share the behind-the-scenes decision to delay launch in order to build the 'Flippy' interface where conversational agents modify their own configurations.1:02:40–1:06:50 · The hosts as informed peer 6/10 Designing Credit-Based Pricing and Usage-Based Economics Alessio asks how Notion structures value-based pricing when individual agent tasks vary wildly in business impact. Sarah explains why usage-based credit abstractions protect unit economics without imposing artificial complexity.1:06:50–1:10:22 · The hosts as informed peer 5/10 The 'Auto' Model Router and Filling the Intelligence-Cost Gap Sarah breaks down the 'Auto' model router, explaining how it steers tasks to the most cost-efficient models while highlighting the mid-tier capability vacuum left by frontier model pricing strategies.1:10:22–1:13:09 · The hosts as informed peer 6/10 Contextual Enterprise Fine-Tuning and Running Overnight Agent Loops Swyx asks if Notion plans to train proprietary foundation models. Simon and Sarah clarify that they prioritize enterprise-specific contextual fine-tuning and overnight autonomous coding agent loops over building generic foundation models.1:13:09–1:15:35 · The hosts as informed peer 5/10 Why Tool Velocity and Outer Loop Robustness Beat Model Training Simon argues that engineering attention is better invested in tool interfaces and harness reliability than fine-tuning models on rapidly changing internal APIs.1:15:35–1:18:25 · The hosts as informed peer 7/10 Adapting Search and Retrieval for Agent-Driven Workloads Sarah details how agent-generated search queries differ fundamentally from human queries, requiring a complete redesign of retrieval architectures, top-K ranking, and query expansion models.1:18:25–1:22:46 · The hosts as informed peer 5/10 Meeting Notes as an Agentic Flywheel and Capture Primitive Swyx asks about Notion's meeting notes and wearable devices. Simon and Sarah conclude by positioning meeting notes as a foundational context-capture primitive that feeds Notion's broader collaboration ecosystem.0:42–5:21 · Guest teaching 5/10 Latent Space Message: Why We Ask for Subscriptions Swyx opens with an ad-free subscriber plea and then presses the guests on why building background agents failed in earlier iterations. Simon and Sarah explain how early context limits and the lack of native tool calling forced multiple redesigns before Sonnet unlocked reliable background execution.5:22–9:34 · Guest teaching 6/10 Balancing AGI Ambitions and the Software Factory Concept Alessio and Swyx ask how Notion plans product roadmaps when frontier model capabilities shift rapidly. Sarah reframes the problem around sensing which direction the technological current is moving rather than fighting upstream against model limitations.9:34–11:57 · Guest teaching 5/10 Horizontal Platform Primitives vs. Vertical SaaS Expertise Alessio contrasts Datadog's vertical persona with Notion's horizontal user base, asking how product expertise is encoded at scale. Sarah and Simon explain that horizontal primitives must stay anchored in concrete user journeys rather than building novel tools for their own sake.11:57–14:44 · Guest teaching 6/10 Fostering an Egoless Engineering Culture of Code Deletion Swyx asks Sarah about her management philosophy on the AI team. Sarah details an egoless engineering culture where engineers are encouraged to throw away old prototypes and rewrite harnesses without tying code ownership to career advancement.14:44–18:07 · Guest teaching 4/10 The 'Simon Vortex' and Hackathons as Capability Uplifters Swyx asks whether regular hackathons drive Notion's product momentum. Sarah and Simon explain that company-wide hackathons serve to uplift baseline technical literacy, while core skunkworks innovation happens continuously in rapid prototyping cycles.18:07–21:53 · Guest teaching 6/10 Organizational Structure and the Design Playground Prototype Repo Alessio asks about prototype evaluation bars across engineering and design. Simon and Sarah describe their isolated design playground repository where designers ship functional code rather than static mockups.21:53–24:03 · Guest teaching 6/10 Scaling Agent Dev Velocity and Decentralized Team Evals Sarah pushes back on Simon's simplified take on prototyping by highlighting the heavy eval and infrastructure burden required before production. She explains that every feature team owns its domain evals while the platform team builds shared harness tooling.24:03–27:24 · Guest teaching 5/10 Detecting Model Provider Regressions and Silent Degradation Swyx asks whether Notion catches silent quality degradation or hidden quantization during peak hours. Sarah shares that enterprise-focused evaluation suites frequently detect regressions that model providers fail to catch with standard coding benchmarks.27:24–30:08 · Guest teaching 7/10 The Rise of Model Behavior Engineers (MBEs) Swyx asks about the Model Behavior Engineer role. Sarah outlines how the position evolved from manual spreadsheet audits by linguistics graduates into engineers who direct coding agents and design LLM judges.30:09–33:24 · Guest teaching 5/10 The Evolution of Software Engineering and Autonomous Workflows Swyx asks if traditional software engineering roles will disappear. Simon and Sarah argue that human engineers are transitioning from syntax implementation to system architecture, verification loops, and delegating to agent fleets.33:24–36:34 · Guest teaching 5/10 Demoing Custom Agents: Automating Tenant Lead Ingestion Alessio demos a live custom agent that automates coworking applicant ingestion, enrichment, and database tracking. Sarah explains how automating tedious internal processes creates high operational leverage across teams.36:34–39:42 · Guest teaching 6/10 Agent Coordination, Manager Agents, and Native Memory Alessio and Swyx ask how agents coordinate without creating infinite recursive loops. Simon explains that Notion relies on database primitives and dedicated manager agents rather than complex bespoke messaging architectures.39:43–43:40 · Guest teaching 6/10 Native Feature Optimizations vs. Third-Party Integrations Swyx asks for Simon's perspective on MCP versus CLI agent interfaces. Simon outlines CLI strengths in progressive disclosure and self-healing while acknowledging MCP's clean sandboxing and permission boundaries.43:40–48:14 · Guest teaching 6/10 Deterministic Code Execution and Cost-Effective Pricing Models Sarah emphasizes that using LLMs to orchestrate deterministic workflows wastes tokens and drives up user costs. She argues that executing deterministic code sandboxes delivers superior reliability and margin efficiency.48:14–52:36 · Guest teaching 7/10 Evolution of Notion's Internal Representation: XML to Markdown and SQLite Swyx asks Simon to walk through the technical iterations of Notion's agent architecture. Simon details the journey from brittle custom XML blocks to markdown and SQLite databases that natively align with model pre-training distributions.52:36–56:25 · Guest teaching 7/10 Decentralizing Tool Ownership and Progressive Disclosure Sarah details how moving away from massive few-shot system prompts to discrete tool definitions allowed Notion to decentralize tool development across independent product squads.56:25–1:00:22 · Guest teaching 6/10 Demystifying System Prompts and Building for Power Users Swyx asks whether Notion protects system prompts as intellectual property. Sarah and Simon push back, explaining that making prompts and tools fully transparent enables power users to build sophisticated workflows.1:00:22–1:02:40 · Guest teaching 6/10 The 'Flippy' Interface: Unifying Settings and Chat Workflows Alessio praises the unified editing and execution canvas. Simon and Sarah share the behind-the-scenes decision to delay launch in order to build the 'Flippy' interface where conversational agents modify their own configurations.1:02:40–1:06:50 · Guest teaching 6/10 Designing Credit-Based Pricing and Usage-Based Economics Alessio asks how Notion structures value-based pricing when individual agent tasks vary wildly in business impact. Sarah explains why usage-based credit abstractions protect unit economics without imposing artificial complexity.1:06:50–1:10:22 · Guest teaching 7/10 The 'Auto' Model Router and Filling the Intelligence-Cost Gap Sarah breaks down the 'Auto' model router, explaining how it steers tasks to the most cost-efficient models while highlighting the mid-tier capability vacuum left by frontier model pricing strategies.1:10:22–1:13:09 · Guest teaching 5/10 Contextual Enterprise Fine-Tuning and Running Overnight Agent Loops Swyx asks if Notion plans to train proprietary foundation models. Simon and Sarah clarify that they prioritize enterprise-specific contextual fine-tuning and overnight autonomous coding agent loops over building generic foundation models.1:13:09–1:15:35 · Guest teaching 7/10 Why Tool Velocity and Outer Loop Robustness Beat Model Training Simon argues that engineering attention is better invested in tool interfaces and harness reliability than fine-tuning models on rapidly changing internal APIs.1:15:35–1:18:25 · Guest teaching 6/10 Adapting Search and Retrieval for Agent-Driven Workloads Sarah details how agent-generated search queries differ fundamentally from human queries, requiring a complete redesign of retrieval architectures, top-K ranking, and query expansion models.1:18:25–1:22:46 · Guest teaching 6/10 Meeting Notes as an Agentic Flywheel and Capture Primitive Swyx asks about Notion's meeting notes and wearable devices. Simon and Sarah conclude by positioning meeting notes as a foundational context-capture primitive that feeds Notion's broader collaboration ecosystem.0:42–5:21 · Guest disagreement 1/10 Latent Space Message: Why We Ask for Subscriptions Swyx opens with an ad-free subscriber plea and then presses the guests on why building background agents failed in earlier iterations. Simon and Sarah explain how early context limits and the lack of native tool calling forced multiple redesigns before Sonnet unlocked reliable background execution.5:22–9:34 · Guest disagreement 2/10 Balancing AGI Ambitions and the Software Factory Concept Alessio and Swyx ask how Notion plans product roadmaps when frontier model capabilities shift rapidly. Sarah reframes the problem around sensing which direction the technological current is moving rather than fighting upstream against model limitations.9:34–11:57 · Guest disagreement 2/10 Horizontal Platform Primitives vs. Vertical SaaS Expertise Alessio contrasts Datadog's vertical persona with Notion's horizontal user base, asking how product expertise is encoded at scale. Sarah and Simon explain that horizontal primitives must stay anchored in concrete user journeys rather than building novel tools for their own sake.11:57–14:44 · Guest disagreement 1/10 Fostering an Egoless Engineering Culture of Code Deletion Swyx asks Sarah about her management philosophy on the AI team. Sarah details an egoless engineering culture where engineers are encouraged to throw away old prototypes and rewrite harnesses without tying code ownership to career advancement.14:44–18:07 · Guest disagreement 2/10 The 'Simon Vortex' and Hackathons as Capability Uplifters Swyx asks whether regular hackathons drive Notion's product momentum. Sarah and Simon explain that company-wide hackathons serve to uplift baseline technical literacy, while core skunkworks innovation happens continuously in rapid prototyping cycles.18:07–21:53 · Guest disagreement 2/10 Organizational Structure and the Design Playground Prototype Repo Alessio asks about prototype evaluation bars across engineering and design. Simon and Sarah describe their isolated design playground repository where designers ship functional code rather than static mockups.21:53–24:03 · Guest disagreement 2/10 Scaling Agent Dev Velocity and Decentralized Team Evals Sarah pushes back on Simon's simplified take on prototyping by highlighting the heavy eval and infrastructure burden required before production. She explains that every feature team owns its domain evals while the platform team builds shared harness tooling.24:03–27:24 · Guest disagreement 1/10 Detecting Model Provider Regressions and Silent Degradation Swyx asks whether Notion catches silent quality degradation or hidden quantization during peak hours. Sarah shares that enterprise-focused evaluation suites frequently detect regressions that model providers fail to catch with standard coding benchmarks.27:24–30:08 · Guest disagreement 1/10 The Rise of Model Behavior Engineers (MBEs) Swyx asks about the Model Behavior Engineer role. Sarah outlines how the position evolved from manual spreadsheet audits by linguistics graduates into engineers who direct coding agents and design LLM judges.30:09–33:24 · Guest disagreement 2/10 The Evolution of Software Engineering and Autonomous Workflows Swyx asks if traditional software engineering roles will disappear. Simon and Sarah argue that human engineers are transitioning from syntax implementation to system architecture, verification loops, and delegating to agent fleets.33:24–36:34 · Guest disagreement 1/10 Demoing Custom Agents: Automating Tenant Lead Ingestion Alessio demos a live custom agent that automates coworking applicant ingestion, enrichment, and database tracking. Sarah explains how automating tedious internal processes creates high operational leverage across teams.36:34–39:42 · Guest disagreement 1/10 Agent Coordination, Manager Agents, and Native Memory Alessio and Swyx ask how agents coordinate without creating infinite recursive loops. Simon explains that Notion relies on database primitives and dedicated manager agents rather than complex bespoke messaging architectures.39:43–43:40 · Guest disagreement 1/10 Native Feature Optimizations vs. Third-Party Integrations Swyx asks for Simon's perspective on MCP versus CLI agent interfaces. Simon outlines CLI strengths in progressive disclosure and self-healing while acknowledging MCP's clean sandboxing and permission boundaries.43:40–48:14 · Guest disagreement 2/10 Deterministic Code Execution and Cost-Effective Pricing Models Sarah emphasizes that using LLMs to orchestrate deterministic workflows wastes tokens and drives up user costs. She argues that executing deterministic code sandboxes delivers superior reliability and margin efficiency.48:14–52:36 · Guest disagreement 1/10 Evolution of Notion's Internal Representation: XML to Markdown and SQLite Swyx asks Simon to walk through the technical iterations of Notion's agent architecture. Simon details the journey from brittle custom XML blocks to markdown and SQLite databases that natively align with model pre-training distributions.52:36–56:25 · Guest disagreement 2/10 Decentralizing Tool Ownership and Progressive Disclosure Sarah details how moving away from massive few-shot system prompts to discrete tool definitions allowed Notion to decentralize tool development across independent product squads.56:25–1:00:22 · Guest disagreement 2/10 Demystifying System Prompts and Building for Power Users Swyx asks whether Notion protects system prompts as intellectual property. Sarah and Simon push back, explaining that making prompts and tools fully transparent enables power users to build sophisticated workflows.1:00:22–1:02:40 · Guest disagreement 1/10 The 'Flippy' Interface: Unifying Settings and Chat Workflows Alessio praises the unified editing and execution canvas. Simon and Sarah share the behind-the-scenes decision to delay launch in order to build the 'Flippy' interface where conversational agents modify their own configurations.1:02:40–1:06:50 · Guest disagreement 1/10 Designing Credit-Based Pricing and Usage-Based Economics Alessio asks how Notion structures value-based pricing when individual agent tasks vary wildly in business impact. Sarah explains why usage-based credit abstractions protect unit economics without imposing artificial complexity.1:06:50–1:10:22 · Guest disagreement 2/10 The 'Auto' Model Router and Filling the Intelligence-Cost Gap Sarah breaks down the 'Auto' model router, explaining how it steers tasks to the most cost-efficient models while highlighting the mid-tier capability vacuum left by frontier model pricing strategies.1:10:22–1:13:09 · Guest disagreement 2/10 Contextual Enterprise Fine-Tuning and Running Overnight Agent Loops Swyx asks if Notion plans to train proprietary foundation models. Simon and Sarah clarify that they prioritize enterprise-specific contextual fine-tuning and overnight autonomous coding agent loops over building generic foundation models.1:13:09–1:15:35 · Guest disagreement 2/10 Why Tool Velocity and Outer Loop Robustness Beat Model Training Simon argues that engineering attention is better invested in tool interfaces and harness reliability than fine-tuning models on rapidly changing internal APIs.1:15:35–1:18:25 · Guest disagreement 1/10 Adapting Search and Retrieval for Agent-Driven Workloads Sarah details how agent-generated search queries differ fundamentally from human queries, requiring a complete redesign of retrieval architectures, top-K ranking, and query expansion models.1:18:25–1:22:46 · Guest disagreement 2/10 Meeting Notes as an Agentic Flywheel and Capture Primitive Swyx asks about Notion's meeting notes and wearable devices. Simon and Sarah conclude by positioning meeting notes as a foundational context-capture primitive that feeds Notion's broader collaboration ecosystem.0:42–5:21 · The hosts pushing back 2/10 Latent Space Message: Why We Ask for Subscriptions Swyx opens with an ad-free subscriber plea and then presses the guests on why building background agents failed in earlier iterations. Simon and Sarah explain how early context limits and the lack of native tool calling forced multiple redesigns before Sonnet unlocked reliable background execution.5:22–9:34 · The hosts pushing back 2/10 Balancing AGI Ambitions and the Software Factory Concept Alessio and Swyx ask how Notion plans product roadmaps when frontier model capabilities shift rapidly. Sarah reframes the problem around sensing which direction the technological current is moving rather than fighting upstream against model limitations.9:34–11:57 · The hosts pushing back 3/10 Horizontal Platform Primitives vs. Vertical SaaS Expertise Alessio contrasts Datadog's vertical persona with Notion's horizontal user base, asking how product expertise is encoded at scale. Sarah and Simon explain that horizontal primitives must stay anchored in concrete user journeys rather than building novel tools for their own sake.11:57–14:44 · The hosts pushing back 1/10 Fostering an Egoless Engineering Culture of Code Deletion Swyx asks Sarah about her management philosophy on the AI team. Sarah details an egoless engineering culture where engineers are encouraged to throw away old prototypes and rewrite harnesses without tying code ownership to career advancement.14:44–18:07 · The hosts pushing back 2/10 The 'Simon Vortex' and Hackathons as Capability Uplifters Swyx asks whether regular hackathons drive Notion's product momentum. Sarah and Simon explain that company-wide hackathons serve to uplift baseline technical literacy, while core skunkworks innovation happens continuously in rapid prototyping cycles.18:07–21:53 · The hosts pushing back 2/10 Organizational Structure and the Design Playground Prototype Repo Alessio asks about prototype evaluation bars across engineering and design. Simon and Sarah describe their isolated design playground repository where designers ship functional code rather than static mockups.21:53–24:03 · The hosts pushing back 2/10 Scaling Agent Dev Velocity and Decentralized Team Evals Sarah pushes back on Simon's simplified take on prototyping by highlighting the heavy eval and infrastructure burden required before production. She explains that every feature team owns its domain evals while the platform team builds shared harness tooling.24:03–27:24 · The hosts pushing back 3/10 Detecting Model Provider Regressions and Silent Degradation Swyx asks whether Notion catches silent quality degradation or hidden quantization during peak hours. Sarah shares that enterprise-focused evaluation suites frequently detect regressions that model providers fail to catch with standard coding benchmarks.27:24–30:08 · The hosts pushing back 2/10 The Rise of Model Behavior Engineers (MBEs) Swyx asks about the Model Behavior Engineer role. Sarah outlines how the position evolved from manual spreadsheet audits by linguistics graduates into engineers who direct coding agents and design LLM judges.30:09–33:24 · The hosts pushing back 3/10 The Evolution of Software Engineering and Autonomous Workflows Swyx asks if traditional software engineering roles will disappear. Simon and Sarah argue that human engineers are transitioning from syntax implementation to system architecture, verification loops, and delegating to agent fleets.33:24–36:34 · The hosts pushing back 1/10 Demoing Custom Agents: Automating Tenant Lead Ingestion Alessio demos a live custom agent that automates coworking applicant ingestion, enrichment, and database tracking. Sarah explains how automating tedious internal processes creates high operational leverage across teams.36:34–39:42 · The hosts pushing back 2/10 Agent Coordination, Manager Agents, and Native Memory Alessio and Swyx ask how agents coordinate without creating infinite recursive loops. Simon explains that Notion relies on database primitives and dedicated manager agents rather than complex bespoke messaging architectures.39:43–43:40 · The hosts pushing back 2/10 Native Feature Optimizations vs. Third-Party Integrations Swyx asks for Simon's perspective on MCP versus CLI agent interfaces. Simon outlines CLI strengths in progressive disclosure and self-healing while acknowledging MCP's clean sandboxing and permission boundaries.43:40–48:14 · The hosts pushing back 2/10 Deterministic Code Execution and Cost-Effective Pricing Models Sarah emphasizes that using LLMs to orchestrate deterministic workflows wastes tokens and drives up user costs. She argues that executing deterministic code sandboxes delivers superior reliability and margin efficiency.48:14–52:36 · The hosts pushing back 2/10 Evolution of Notion's Internal Representation: XML to Markdown and SQLite Swyx asks Simon to walk through the technical iterations of Notion's agent architecture. Simon details the journey from brittle custom XML blocks to markdown and SQLite databases that natively align with model pre-training distributions.52:36–56:25 · The hosts pushing back 2/10 Decentralizing Tool Ownership and Progressive Disclosure Sarah details how moving away from massive few-shot system prompts to discrete tool definitions allowed Notion to decentralize tool development across independent product squads.56:25–1:00:22 · The hosts pushing back 2/10 Demystifying System Prompts and Building for Power Users Swyx asks whether Notion protects system prompts as intellectual property. Sarah and Simon push back, explaining that making prompts and tools fully transparent enables power users to build sophisticated workflows.1:00:22–1:02:40 · The hosts pushing back 1/10 The 'Flippy' Interface: Unifying Settings and Chat Workflows Alessio praises the unified editing and execution canvas. Simon and Sarah share the behind-the-scenes decision to delay launch in order to build the 'Flippy' interface where conversational agents modify their own configurations.1:02:40–1:06:50 · The hosts pushing back 2/10 Designing Credit-Based Pricing and Usage-Based Economics Alessio asks how Notion structures value-based pricing when individual agent tasks vary wildly in business impact. Sarah explains why usage-based credit abstractions protect unit economics without imposing artificial complexity.1:06:50–1:10:22 · The hosts pushing back 2/10 The 'Auto' Model Router and Filling the Intelligence-Cost Gap Sarah breaks down the 'Auto' model router, explaining how it steers tasks to the most cost-efficient models while highlighting the mid-tier capability vacuum left by frontier model pricing strategies.1:10:22–1:13:09 · The hosts pushing back 2/10 Contextual Enterprise Fine-Tuning and Running Overnight Agent Loops Swyx asks if Notion plans to train proprietary foundation models. Simon and Sarah clarify that they prioritize enterprise-specific contextual fine-tuning and overnight autonomous coding agent loops over building generic foundation models.1:13:09–1:15:35 · The hosts pushing back 1/10 Why Tool Velocity and Outer Loop Robustness Beat Model Training Simon argues that engineering attention is better invested in tool interfaces and harness reliability than fine-tuning models on rapidly changing internal APIs.1:15:35–1:18:25 · The hosts pushing back 2/10 Adapting Search and Retrieval for Agent-Driven Workloads Sarah details how agent-generated search queries differ fundamentally from human queries, requiring a complete redesign of retrieval architectures, top-K ranking, and query expansion models.1:18:25–1:22:46 · The hosts pushing back 2/10 Meeting Notes as an Agentic Flywheel and Capture Primitive Swyx asks about Notion's meeting notes and wearable devices. Simon and Sarah conclude by positioning meeting notes as a foundational context-capture primitive that feeds Notion's broader collaboration ecosystem.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 40.5% · guest 59.5%0:00 · the hosts 40.5% · guest 59.5%3:00 · the hosts 17.8% · guest 82.2%3:00 · the hosts 17.8% · guest 82.2%6:00 · the hosts 12.7% · guest 87.3%6:00 · the hosts 12.7% · guest 87.3%9:00 · the hosts 26.1% · guest 73.9%9:00 · the hosts 26.1% · guest 73.9%12:00 · the hosts 15.5% · guest 84.5%12:00 · the hosts 15.5% · guest 84.5%15:00 · the hosts 3.2% · guest 96.8%15:00 · the hosts 3.2% · guest 96.8%18:00 · the hosts 12.3% · guest 87.7%18:00 · the hosts 12.3% · guest 87.7%21:00 · the hosts 3.4% · guest 96.6%21:00 · the hosts 3.4% · guest 96.6%24:00 · the hosts 14.6% · guest 85.4%24:00 · the hosts 14.6% · guest 85.4%27:00 · the hosts 2.4% · guest 97.6%27:00 · the hosts 2.4% · guest 97.6%30:00 · the hosts 10.8% · guest 89.2%30:00 · the hosts 10.8% · guest 89.2%33:00 · the hosts 58% · guest 42%33:00 · the hosts 58% · guest 42%36:00 · the hosts 24.4% · guest 75.6%36:00 · the hosts 24.4% · guest 75.6%39:00 · the hosts 17.7% · guest 82.3%39:00 · the hosts 17.7% · guest 82.3%42:00 · the hosts 1.7% · guest 98.3%42:00 · the hosts 1.7% · guest 98.3%45:00 · the hosts 23.9% · guest 76.1%45:00 · the hosts 23.9% · guest 76.1%48:00 · the hosts 16.9% · guest 83.1%48:00 · the hosts 16.9% · guest 83.1%51:00 · the hosts 17.8% · guest 82.2%51:00 · the hosts 17.8% · guest 82.2%54:00 · the hosts 3.9% · guest 96.1%54:00 · the hosts 3.9% · guest 96.1%57:00 · the hosts 15.4% · guest 84.6%57:00 · the hosts 15.4% · guest 84.6%1:00:00 · the hosts 23.6% · guest 76.4%1:00:00 · the hosts 23.6% · guest 76.4%1:03:00 · the hosts 20.6% · guest 79.4%1:03:00 · the hosts 20.6% · guest 79.4%1:06:00 · the hosts 8.1% · guest 91.9%1:06:00 · the hosts 8.1% · guest 91.9%1:09:00 · the hosts 6.6% · guest 93.4%1:09:00 · the hosts 6.6% · guest 93.4%1:12:00 · the hosts 4.2% · guest 95.8%1:12:00 · the hosts 4.2% · guest 95.8%1:15:00 · the hosts 14.5% · guest 85.5%1:15:00 · the hosts 14.5% · guest 85.5%1:18:00 · the hosts 21.2% · guest 78.8%1:18:00 · the hosts 21.2% · guest 78.8%1:21:00 · the hosts 9.4% · guest 90.6%1:21:00 · the hosts 9.4% · guest 90.6%1:24:00 · the hosts 26.5% · guest 73.5%1:24:00 · the hosts 26.5% · guest 73.5%
Sharpest disagreement ▶ 57:28 Rejecting the accessible UX paradigm

Sarah firmly rejects the premise that AI tools should be simplified for the lowest common denominator, arguing that dumbing down interfaces destroys agent capability.

Hardest push from the hosts ▶ 24:03 Interrogating silent model degradation

Swyx presses the guests on whether major frontier labs secretly quantize or throttle model quality during high-traffic enterprise working hours.

Biggest teaching moment ▶ 49:41 Designing systems for model training distributions

Simon breaks down why Notion abandoned its internal XML formats for standard Markdown and SQLite after learning models perform best on representations mirroring pre-training data.

The host holds their own ▶ 1:16:00 Challenging retrieval optimization metrics

Swyx challenges Sarah on top-K retrieval nuances, prompting a detailed technical breakdown of loss functions and agentic query distribution patterns.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Latent Space Message: Why We Ask for Subscriptions 4512 Swyx opens with an ad-free subscriber plea and then presses the guests on why building background agents failed in earlier iterations. Simon and Sarah explain how early context limits and the lack of native tool calling forced multiple redesigns before Sonnet unlocked reliable background execution.
Balancing AGI Ambitions and the Software Factory Concept 5622 Alessio and Swyx ask how Notion plans product roadmaps when frontier model capabilities shift rapidly. Sarah reframes the problem around sensing which direction the technological current is moving rather than fighting upstream against model limitations.
Horizontal Platform Primitives vs. Vertical SaaS Expertise 6523 Alessio contrasts Datadog's vertical persona with Notion's horizontal user base, asking how product expertise is encoded at scale. Sarah and Simon explain that horizontal primitives must stay anchored in concrete user journeys rather than building novel tools for their own sake.
Fostering an Egoless Engineering Culture of Code Deletion 4611 Swyx asks Sarah about her management philosophy on the AI team. Sarah details an egoless engineering culture where engineers are encouraged to throw away old prototypes and rewrite harnesses without tying code ownership to career advancement.
The 'Simon Vortex' and Hackathons as Capability Uplifters 5422 Swyx asks whether regular hackathons drive Notion's product momentum. Sarah and Simon explain that company-wide hackathons serve to uplift baseline technical literacy, while core skunkworks innovation happens continuously in rapid prototyping cycles.
Organizational Structure and the Design Playground Prototype Repo 4622 Alessio asks about prototype evaluation bars across engineering and design. Simon and Sarah describe their isolated design playground repository where designers ship functional code rather than static mockups.
Scaling Agent Dev Velocity and Decentralized Team Evals 5622 Sarah pushes back on Simon's simplified take on prototyping by highlighting the heavy eval and infrastructure burden required before production. She explains that every feature team owns its domain evals while the platform team builds shared harness tooling.
Detecting Model Provider Regressions and Silent Degradation 6513 Swyx asks whether Notion catches silent quality degradation or hidden quantization during peak hours. Sarah shares that enterprise-focused evaluation suites frequently detect regressions that model providers fail to catch with standard coding benchmarks.
The Rise of Model Behavior Engineers (MBEs) 5712 Swyx asks about the Model Behavior Engineer role. Sarah outlines how the position evolved from manual spreadsheet audits by linguistics graduates into engineers who direct coding agents and design LLM judges.
The Evolution of Software Engineering and Autonomous Workflows 6523 Swyx asks if traditional software engineering roles will disappear. Simon and Sarah argue that human engineers are transitioning from syntax implementation to system architecture, verification loops, and delegating to agent fleets.
Demoing Custom Agents: Automating Tenant Lead Ingestion 6511 Alessio demos a live custom agent that automates coworking applicant ingestion, enrichment, and database tracking. Sarah explains how automating tedious internal processes creates high operational leverage across teams.
Agent Coordination, Manager Agents, and Native Memory 5612 Alessio and Swyx ask how agents coordinate without creating infinite recursive loops. Simon explains that Notion relies on database primitives and dedicated manager agents rather than complex bespoke messaging architectures.
Native Feature Optimizations vs. Third-Party Integrations 6612 Swyx asks for Simon's perspective on MCP versus CLI agent interfaces. Simon outlines CLI strengths in progressive disclosure and self-healing while acknowledging MCP's clean sandboxing and permission boundaries.
Deterministic Code Execution and Cost-Effective Pricing Models 5622 Sarah emphasizes that using LLMs to orchestrate deterministic workflows wastes tokens and drives up user costs. She argues that executing deterministic code sandboxes delivers superior reliability and margin efficiency.
Evolution of Notion's Internal Representation: XML to Markdown and SQLite 7712 Swyx asks Simon to walk through the technical iterations of Notion's agent architecture. Simon details the journey from brittle custom XML blocks to markdown and SQLite databases that natively align with model pre-training distributions.
Decentralizing Tool Ownership and Progressive Disclosure 5722 Sarah details how moving away from massive few-shot system prompts to discrete tool definitions allowed Notion to decentralize tool development across independent product squads.
Demystifying System Prompts and Building for Power Users 5622 Swyx asks whether Notion protects system prompts as intellectual property. Sarah and Simon push back, explaining that making prompts and tools fully transparent enables power users to build sophisticated workflows.
The 'Flippy' Interface: Unifying Settings and Chat Workflows 5611 Alessio praises the unified editing and execution canvas. Simon and Sarah share the behind-the-scenes decision to delay launch in order to build the 'Flippy' interface where conversational agents modify their own configurations.
Designing Credit-Based Pricing and Usage-Based Economics 6612 Alessio asks how Notion structures value-based pricing when individual agent tasks vary wildly in business impact. Sarah explains why usage-based credit abstractions protect unit economics without imposing artificial complexity.
The 'Auto' Model Router and Filling the Intelligence-Cost Gap 5722 Sarah breaks down the 'Auto' model router, explaining how it steers tasks to the most cost-efficient models while highlighting the mid-tier capability vacuum left by frontier model pricing strategies.
Contextual Enterprise Fine-Tuning and Running Overnight Agent Loops 6522 Swyx asks if Notion plans to train proprietary foundation models. Simon and Sarah clarify that they prioritize enterprise-specific contextual fine-tuning and overnight autonomous coding agent loops over building generic foundation models.
Why Tool Velocity and Outer Loop Robustness Beat Model Training 5721 Simon argues that engineering attention is better invested in tool interfaces and harness reliability than fine-tuning models on rapidly changing internal APIs.
Adapting Search and Retrieval for Agent-Driven Workloads 7612 Sarah details how agent-generated search queries differ fundamentally from human queries, requiring a complete redesign of retrieval architectures, top-K ranking, and query expansion models.
Meeting Notes as an Agentic Flywheel and Capture Primitive 5622 Swyx asks about Notion's meeting notes and wearable devices. Simon and Sarah conclude by positioning meeting notes as a foundational context-capture primitive that feeds Notion's broader collaboration ecosystem.

Statements from this episode (51)

Assertion Not checkable as stated
Sachs: Custom Agents Was Notion's Most Successful Launch by Conversions
“It was our most successful launch in terms of free trials and converting people and things like that.”
Sarah Sachs Apr 15, 2026 ▶ 2:30
Disclosure
Notion Rebuilt Its Agent Framework Four or Five Times Before Launch
“It's probably the fourth or fifth time that we rebuilt that.”
Simon Last Apr 15, 2026 ▶ 2:48
Disclosure
Notion Fine-Tuned Custom Function-Calling Models Before Standard Tool Calling Existed
“Before function calling came out, we were trying to fine tune with the frontier labs and with fireworks, like a function calling model on notion functions.”
Sarah Sachs Apr 15, 2026 ▶ 3:21
Insight
Claude Sonnet Was the Breakthrough That Unlocked Notion's Custom Agents
“The big unlock was Probably like Sonic 3.6 or seven early last year. And that's when we started working on our agent, which we shipped last year.”
Simon Last Apr 15, 2026 ▶ 4:20
Insight
Simon Last Argues That Coding Agents Are the Kernel of AGI
“I think one thing that's becoming more clear is I think like coding agents are the kernel VGI, sort of everything is a coding agent.”
Simon Last Apr 15, 2026 ▶ 6:09
Insight
Last: Coding agents allow software to bootstrap and debug itself
“The exciting thing about that is sort of your agent can sort of bootstrap its own software and capabilities and actually debug and maintain them.”
Simon Last Apr 15, 2026 ▶ 6:19
Insight
Frontier AI Product Development Requires Knowing When to Stop Fighting Model Limits
“Is really what will set notion apart for every new capability is we have like two skills that are crucial when it comes to frontier capabilities. One is not letting yourself swim upstream. So like quickly realizing if you're just pressing against model capabil…”
Sarah Sachs Apr 15, 2026 ▶ 7:15
Opinion
Notion Will Beat Frontier Labs in Collaboration Like Datadog Beat AWS
“Datadog could not exist without cloud storage, right? That it's kind of fundamental that that works. And AWS has like a cloud watch product, but Datadog is an expert on understanding how people want observability on the products they launch. And we're experts …”
Sarah Sachs Apr 15, 2026 ▶ 9:14
Insight
AI Teams Lose Velocity When Building Cool Tools Without User Journeys
“I actually think the failure of our team is when we focus too much on what are tools that are cool tools. I actually think that's when we make have the least velocity because you still need some sort of focus on a user journey.”
Sarah Sachs Apr 15, 2026 ▶ 10:54
Disclosure
Notion's AI Team Reviews p99 Token-Exhaustive Agent Transcripts Every Friday
“We'll all sit down every Friday and look at the P. 99 of like the most token exhaustive custom agent transcript and just look at why it didn't do well and cut a bunch of tasks.”
Sarah Sachs Apr 15, 2026 ▶ 11:07
Disclosure
Notion Dedicates 50 Engineers to Core AI and 40 to Packaging
“I manage, ah, the team that's what we'll call like core AI capabilities and infrastructure. That's about 50 people. But then we have, I partner teams that do packaging, so how it shows up in the corner chat versus custom agents versus meeting notes, that's ano…”
Sarah Sachs Apr 15, 2026 ▶ 18:11
Disclosure
Notion's Design Team Builds Live Code Prototypes Instead of Static Mocks
“Our design team made a whole separate GitHub repo called the design playground. And it's basically just to create a bunch of like helper components and for quickly throwing together UIs and it's become like actually quite sophisticated. Like it has like an age…”
Simon Last Apr 15, 2026 ▶ 19:29
Insight
Frictionless AI Prototyping Requires Strong Product Conviction to Avoid Feature Bloat
“I think it's put a big pressure point on us to have really strong product conviction, because if anything can be demoed, You really need a strong filter of making sure that if, you know, you're doing X amount of work, you're making the, you're focusing on one …”
Sarah Sachs Apr 15, 2026 ▶ 20:52
Disclosure
Notion Feature Teams Own Their AI Evaluations Running Nightly in CI
“We maintain the eval framework. Every team owns their own evals and a lot of them we've integrated to opt into CI or we run them nightly and we have a team, a custom agent that triggers to a team to look at the major failures.”
Sarah Sachs Apr 15, 2026 ▶ 23:08
Assertion Not checkable as stated
Notion Observes Increased AI API Latency During Standard Business Hours
“We've definitely noticed, particularly for some providers that things are slower during working hours.”
Sarah Sachs Apr 15, 2026 ▶ 24:20
Assertion Supported
Sachs: AI Model Quality Varies Between First-Party APIs and Cloud Providers
“Companies that say they're selling the same model through different vendors, whether it be through first party or Bedrock, Azure, et cetera, we do see different qualities sometimes, and that's not necessarily what's advertised.”
Sarah Sachs Apr 15, 2026 ▶ 24:37
Assertion Not checkable as stated
Frontier AI Labs Altered Final Model Snapshots Based on Notion's Enterprise Feedback
“Like they have shipped models that were not the snapshots that we wanted, and they have changed the snapshots that they shipped based on the feedback that we give, because our feedback tends to be more enterprise work focused and not coding agent focused.”
Sarah Sachs Apr 15, 2026 ▶ 25:23
Disclosure
Notion Partnered With Anthropic and OpenAI to Build 30% Pass Rate Evals
“And then what we have, what we call Frontier Headroom evals, where we actively want to be at 30% pass rate. And that's actually been a effort that we took in partnership with Anthropic and OpenAI in the past maybe two or three months, because we actually hit a…”
Sarah Sachs Apr 15, 2026 ▶ 26:28
Opinion
Model Behavior Engineers Do Not Need Traditional Engineering Backgrounds
“And so, we have, like, very firm conviction And we have had for a number of years now that that is its own career path, and we have always welcomed the misfits, so to speak. So we really firmly believe that you don't need an engineering background to be the be…”
Sarah Sachs Apr 15, 2026 ▶ 28:49
Opinion
AI Evaluation Harnesses Should Remain Compatible With General CLI Tools
“Yeah, I think it'll be a mistake to like fix it on any particular coding agent. At the end of the day, it's just like CLI tools.”
Simon Last Apr 15, 2026 ▶ 29:43
Insight
Coding Agents Should Write AI Model Evaluations Like Standard Unit Tests
“It's like the same way that you would have a coding agent write the unit test. You should have a coding agent write the eval.”
Sarah Sachs Apr 15, 2026 ▶ 29:49
Insight
Notion Engineers Faced an Identity Crisis Shifting From Coding to AI Delegation
“Every software engineer in Notion this summer went through like this sheer one of our engineering leads at the company called it like, every software engineer is going through the identity crisis that every manager goes through, where all of a sudden they real…”
Sarah Sachs Apr 15, 2026 ▶ 31:17
Insight
Supervising AI Agents Is a Deeply Technical Systems Problem, Unlike Managing Humans
“There's a critical difference To being a manager, which is that like, it is actually very deeply technical. The problem of, you know, humans are very like fuzzy and you can't like treat a team of humans like a rigorous system where like, you know, PRs like fle…”
Simon Last Apr 15, 2026 ▶ 31:41
Disclosure
Notion's Internal Bug-Triaging Custom Agent Completely Changed Company Operations
“One of our biggest use cases is bug triaging. So if someone posts something in Slack and you just have a custom agent that lives there that has its own routing constitution of what team this belongs to. Creates a task in your task database and then posts in th…”
Sarah Sachs Apr 15, 2026 ▶ 35:43
Disclosure
Notion Will Soon Release Direct Agent-to-Agent Invocation Settings
“So I think it's actually not released yet. Releasing it like next week is in the settings for an agent, you can give it access to invoke any other agent. So you can have them just talk directly.”
Simon Last Apr 15, 2026 ▶ 36:55
Disclosure
Notion Custom Agents Use Standard Pages and Databases as Their Memory
“Another example of this is we have no built-in memory concept. Memory is, is just pages and databases. And so if you want to give it memory, just give it a page and give it access to that page.”
Simon Last Apr 15, 2026 ▶ 39:23
Insight
CLI Environments Let AI Agents Autonomously Debug and Fix Broken Tools
“And then I think the most important thing that's super cool is that there it's also inherently a bootstrapped. So if there's an issue of the agent can debug and fix itself within the same environment that it uses the tool, right?”
Simon Last Apr 15, 2026 ▶ 41:45
Insight
MCP Provides a Safer Permission Model Than CLIs by Preventing Token Exfiltration
“MCP inherently has a really strong permission model. Like, all you can do is call the tools. A CLI is a little bit murkier. It's like, can I access the API token? Are you, like, properly sort of, like, re-encrypting the token so it can't, like, exfiltrate it? …”
Simon Last Apr 15, 2026 ▶ 43:20
Disclosure
Notion Commits to Supporting MCP Integrations as Long as Users Demand Them
“So we will always support our MCP insofar as other people are using MCPs, right?”
Sarah Sachs Apr 15, 2026 ▶ 43:50
Insight
Executing Deterministic Code via CLI Avoids Recurring LLM Token Fees From MCP
“And so if we can have our agents properly execute code that calls on CLI deterministically, it's a one time cost, right? Versus constantly having a language model integrate with an MCP over and over and over and paying those like repeated token fees.”
Sarah Sachs Apr 15, 2026 ▶ 44:42
Disclosure
Notion Bypasses Third-Party MCP Search to Maintain Internal Quality Control
“There are a lot of things where we choose not to use MCP because we want to add more high touch to quality. I think search and agentic find is like the largest instance of that, where we have slack and linear and JIRA search and notion that is not using necess…”
Sarah Sachs Apr 15, 2026 ▶ 46:03
Assertion Supported
MCP Lacks a Trigger Protocol, Forcing Notion to Build Custom Trigger Infrastructure
“MCP is a really great way to gain access to tools. Works really well. But you just looked at the trigger UI, for example, there's no trigger protocol. And so those are things we had to build ourselves.”
Simon Last Apr 15, 2026 ▶ 47:30
Insight
Agent Environments Must Favor Model Preferences Over Internal System Architecture
“I mean, that was, I would say that was a big learning is just, you know, really be savvy and really careful thinking about what the model wants in terms of, you know, its environment and cater around that and really try so hard not to expose it to any complexi…”
Simon Last Apr 15, 2026 ▶ 51:22
Insight
Goal-Driven Tool Definitions Allowed Notion to Distribute Agent Tool Ownership
“When we went away from few shots to describing the goal of the tool and like goal driven, basically moving from a DAG to like a true system with feedback, that's when we could distribute tool ownership to the teams much better.”
Sarah Sachs Apr 15, 2026 ▶ 54:33
Assertion Not checkable as stated
Claude Sonnet Crashed on Duplicate Tool Names While OpenAI Handled the Error
“Sonic couldn't handle two tools with the same name in OpenAI, GPT, 5.2. It was like, ah, I can figure this out. So that was an interesting one that we learned by accident through a SEV.”
Sarah Sachs Apr 15, 2026 ▶ 55:35
Assertion Not publicly verifiable
Notion's Custom Agent Operates With Over 100 Distinct Internal Feature Tools
“So there's now like over a hundred tools just for all, all the crazy notion stuff.”
Simon Last Apr 15, 2026 ▶ 56:33
Insight
Over-Simplifying AI Agent Interfaces Abstracts Interpretability and Nerfs Agent Capabilities
“I'd actually say we don't try and make it as easy as possible to use because the more we do that, the more we abstract away that interpretability that Simon's talking about that basically nerfs the model or nerfs the agent from being super capable.”
Sarah Sachs Apr 15, 2026 ▶ 57:29
Disclosure
Self-Healing Agent Capabilities Are Next on Notion's Product Roadmap
“Obviously we should build product of self-healing. That's next on our roadmap.”
Sarah Sachs Apr 15, 2026 ▶ 59:01
Assertion Supported
Notion Custom Agents Are Deployed With Zero Permissions by Default
“The thing about custom agents that is that by default, it has no permission to do anything. And then you have to explicitly granted all the permissions.”
Simon Last Apr 15, 2026 ▶ 59:32
Disclosure
Notion's Database Autofill Will Soon Become Agentic With Usage-Based Pricing
“The database autofill feature will soon be agentic. That will be associated with usage based pricing”
Sarah Sachs Apr 15, 2026 ▶ 1:05:23
Insight
Users Do Not Care About Inference Speed for Asynchronous Custom Agents
“People don't care about speed and custom agents. And so the incentive of like haiku being faster, people don't care when it's asynchronous.”
Sarah Sachs Apr 15, 2026 ▶ 1:06:30
Assertion Not checkable as stated
The Majority of Notion AI Tasks Do Not Require Opus-Level Intelligence
“It's not Opus, that's for sure, because a majority of the tasks people are doing aren't Opus level intelligence.”
Sarah Sachs Apr 15, 2026 ▶ 1:07:55
Opinion
AI Labs Cluster at Extremes and Ignore Intermediate Intelligence-Latency Tiers
“Everyone's clustered in capability, where everyone's clustered, I mean, Haiku's not that much cheaper. No one's really in the middle. Like, people really tend to cluster around two. Like, this is really capable, and it's really fast, but it's really expensive,…”
Sarah Sachs Apr 15, 2026 ▶ 1:09:58
Assertion Not checkable as stated
Notion Fine-Tunes Its AI Capabilities Entirely Without Relying on User Data
“There are certain aspects of Notion where we do fine tune and do reinforce and fine tuning on our own capabilities. But that's not necessarily trained on user data. You don't need that, that much data in the first place.”
Sarah Sachs Apr 15, 2026 ▶ 1:11:44
Disclosure
Simon Last Runs Autonomous Coding Agents Overnight Before Going to Sleep
“So now every night before I go to bed, I'm like, okay, did I start enough agents to get them done? I get everything done.”
Simon Last Apr 15, 2026 ▶ 1:12:47
Insight
Fine-Tuning Models on Internal Tools Unnecessarily Slows Down Rapid Product Development
“It would actually really slow us down to have a model that was fine tuned on our tools because we'd have to retrain it and cut a new model every time we did that.”
Sarah Sachs Apr 15, 2026 ▶ 1:14:18
Insight
Simon Last: 99% of Agent Failures Are Tool Bugs, Not Model Flaws
“And actually, 99% of the time, it's a bug in one of the tools. Right. And so, just fix the bug.”
Simon Last Apr 15, 2026 ▶ 1:15:20
Assertion Not checkable as stated
The Majority of Notion Enterprise Search Traffic Now Comes From AI Agents
“We're at a point right now in our business and enterprise or AI enabled plans where the search load and the search traffic, a majority of it's coming from agents, not humans.”
Sarah Sachs Apr 15, 2026 ▶ 1:15:41
Insight
Agent Search Prioritizes Top-K Retrieval Depth Over Precise Positional Ranking
“Positional ranking matters less, but top care retrieval mode matters more.”
Sarah Sachs Apr 15, 2026 ▶ 1:15:59
Insight
Optimizing Vector Embedding Models Is No Longer the Right RAG Retrieval Focus
“We don't spend a lot of time trying to optimize what vector embedding we use anymore. That was a period of time, but that's just not the right level of optimization, right?”
Sarah Sachs Apr 15, 2026 ▶ 1:18:12
Disclosure
Notion's Meeting Notes Automatically Mention Individuals Referenced in Generated Summaries
“One thing that the meeting notes team added recently that was But I've been blowing my mind is they, we have, they made it so it actually, when it makes the summary will actually mention the people that were referenced in it. So I now get notifications wheneve…”
Simon Last Apr 15, 2026 ▶ 1:21:00
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.