Nov 15, 2024 · 1h 8m · latent-space

Agents @ Work: Lindy.ai (with live demo!)

Florent Crivello · 38m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

On the Latent Space Podcast, Lindy.ai founder and CEO Florent Crivello explores the evolution of AI agents from fragile prompts to deterministic no-code workflows, delivers comprehensive live product demonstrations, and shares strategic insights on startup building, API-first automation, and AI governance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.9 Guest teaching 4.9 Guest disagreement 2.5 The hosts pushing back 2.2
05100:0015:0030:0045:001:00:000:55–2:59 · The hosts as informed peer 4/10 Lindy on Rails: Evolution of the AI Agent Architecture The hosts set up the conversation referencing Florent's previous talks and Andrew Wilkinson's adoption of Lindy. Florent explains the architectural transition to 'Lindy on Rails', acknowledging past over-reliance on pure prompting.3:00–5:57 · The hosts as informed peer 6/10 Deterministic Workflows vs. Pure Prompting Co-host frames the philosophy of putting 'Shoggoth in a box' and prioritizing deterministic glue code over pure natural language prompts. Florent explains why GUI and guardrails outperform text boxes for real users.5:58–8:51 · The hosts as informed peer 5/10 Managing Permissions, Scopes, and Agent Accounts Alessio queries how Lindy handles broad OAuth scopes and enterprise permissions. Florent details incremental permissions and explains the emerging hack of users provisioning dedicated Google Workspace accounts for their agents.8:51–13:21 · The hosts as informed peer 4/10 Live Demo: Automated Rental Logging and Structured Nodes Florent runs a live demo showing automated rental logging and his personal meeting recorder Lindy. The hosts ask clarifying architectural questions regarding structured output nodes and context restoration.13:21–16:13 · The hosts as informed peer 5/10 Agent Memory Architecture and Management Co-host presses on why Florent's Lindy has so few saved memories despite years of use. Florent explains that agent architectures degrade with excessive memory and cites conversations with Meta's Llama team on prompt engineering limits.16:13–19:38 · The hosts as informed peer 4/10 Modular Agents, Health Tracking, and Nested Workflows Florent showcases health logging workflows and explains the modular boundary between general personal Lindies and specialized sub-agents. Alessio clarifies how multiple Lindies invoke and chain one another.19:38–24:03 · The hosts as informed peer 5/10 Specialized Tools: Calendar Scheduling and PR Reviewers Alessio asks about vertical vs. horizontal SaaS boundaries. Florent presents his thesis that agents aggregate horizontally like Google Search across verticals, while the co-host playfully attempts to fish for a spicy hot take.24:03–26:15 · The hosts as informed peer 5/10 Building a No-Code Agent Community and User Base Co-host advises Florent on framing the user base around high-budget executive assistants rather than 'no-code tinkerers'. Florent notes that practical CEO behavior doesn't support self-serve setup without tinkering.26:16–30:29 · The hosts as informed peer 5/10 Customer Support Evals and the Rickroll Incident Florent recounts an incident where their customer support agent hallucinated a Rickroll link. Alessio and the co-host explore eval generation strategies, structured outputs, and the necessity of domain-specific test harnesses.30:30–34:24 · The hosts as informed peer 6/10 Model Advancements, Prompt Caching, and Poor Man's RLHF Florent explains 'poor man's RLHF' via user approval loops and vector database caching, highlighting Claude 3.5 Sonnet's superiority over GPT-4o. Alessio contributes his own prompt refactoring experience around prompt caching.34:25–37:30 · The hosts as informed peer 6/10 The Bitter Lesson and Cognitive Architectures Florent argues for the Bitter Lesson, dismissing complex cognitive architectures and multi-agent debate wrappers. The co-host points to Replit's opposing architecture and benchmark results on o1 vs. Sonnet.37:31–40:38 · The hosts as informed peer 5/10 Computer Use vs. API-Driven Execution Addressing Anthropic's new Computer Use release, Florent forcefully argues that APIs will dominate due to latency and reliability, using the early Android JVM/garbage-collection vs. iOS performance gap as a cautionary analogy.40:38–43:00 · The hosts as informed peer 4/10 Operational Scaling: General Managers and the Factorio Analogy Florent explains hiring general managers to scale horizontal Lindy templates, drawing on his Uber experience. The conversation shifts to how running thousands of autonomous agents resembles Factorio factory optimization.43:00–46:50 · The hosts as informed peer 5/10 Bottom-Up Execution vs. Top-Down Legibility Co-host reflects on breaking large complex workflows into small steps. Florent references 'Seeing Like a State', illustrating how optimizing for top-down legibility inevitably sacrifices bottom-up operational performance.46:51–51:07 · The hosts as informed peer 5/10 The Shift from Remote Work to In-Person Collaboration Florent does an about-face on remote work despite previously founding a virtual office startup. When the co-host defends remote companies like GitLab and Automattic, Florent dismisses Automattic as a commercial failure compared to centralized alternatives.51:08–58:27 · The hosts as informed peer 5/10 Direction vs Magnitude in Creative Engineering Florent claims tech workers outside San Francisco lack either judgment or ambition. Alessio joins in detailing his relocation from Europe, and the group critiques European venture appetite and regulatory stagnation.58:27–1:03:18 · The hosts as informed peer 5/10 Overton Windows, Contrarianism, and AI Regulation Florent discusses pushing the Overton window on Twitter and supporting AI safety bill SB 1047, drawing a historical parallel to French resistance during WWII. Co-host calls out an inconsistency between imminent AGI timelines and claims of low stakes.1:03:18–1:05:33 · The hosts as informed peer 5/10 P(Doom), Utopia, and Model-Layer Safety Alessio asks how safety concerns square with building agentic tooling. Florent estimates a 10% P(Doom) versus a 90% chance of post-scarcity utopia, arguing catastrophic downside risk resides at the model layer rather than the application layer.1:05:33–1:07:48 · The hosts as informed peer 4/10 Hiring and the Primacy of Product Design The conversation concludes with hiring announcements. Co-host and Florent align on the importance of product design and end-to-end user journeys as key defensible moats in the agent ecosystem.0:55–2:59 · Guest teaching 4/10 Lindy on Rails: Evolution of the AI Agent Architecture The hosts set up the conversation referencing Florent's previous talks and Andrew Wilkinson's adoption of Lindy. Florent explains the architectural transition to 'Lindy on Rails', acknowledging past over-reliance on pure prompting.3:00–5:57 · Guest teaching 5/10 Deterministic Workflows vs. Pure Prompting Co-host frames the philosophy of putting 'Shoggoth in a box' and prioritizing deterministic glue code over pure natural language prompts. Florent explains why GUI and guardrails outperform text boxes for real users.5:58–8:51 · Guest teaching 4/10 Managing Permissions, Scopes, and Agent Accounts Alessio queries how Lindy handles broad OAuth scopes and enterprise permissions. Florent details incremental permissions and explains the emerging hack of users provisioning dedicated Google Workspace accounts for their agents.8:51–13:21 · Guest teaching 5/10 Live Demo: Automated Rental Logging and Structured Nodes Florent runs a live demo showing automated rental logging and his personal meeting recorder Lindy. The hosts ask clarifying architectural questions regarding structured output nodes and context restoration.13:21–16:13 · Guest teaching 5/10 Agent Memory Architecture and Management Co-host presses on why Florent's Lindy has so few saved memories despite years of use. Florent explains that agent architectures degrade with excessive memory and cites conversations with Meta's Llama team on prompt engineering limits.16:13–19:38 · Guest teaching 4/10 Modular Agents, Health Tracking, and Nested Workflows Florent showcases health logging workflows and explains the modular boundary between general personal Lindies and specialized sub-agents. Alessio clarifies how multiple Lindies invoke and chain one another.19:38–24:03 · Guest teaching 6/10 Specialized Tools: Calendar Scheduling and PR Reviewers Alessio asks about vertical vs. horizontal SaaS boundaries. Florent presents his thesis that agents aggregate horizontally like Google Search across verticals, while the co-host playfully attempts to fish for a spicy hot take.24:03–26:15 · Guest teaching 3/10 Building a No-Code Agent Community and User Base Co-host advises Florent on framing the user base around high-budget executive assistants rather than 'no-code tinkerers'. Florent notes that practical CEO behavior doesn't support self-serve setup without tinkering.26:16–30:29 · Guest teaching 5/10 Customer Support Evals and the Rickroll Incident Florent recounts an incident where their customer support agent hallucinated a Rickroll link. Alessio and the co-host explore eval generation strategies, structured outputs, and the necessity of domain-specific test harnesses.30:30–34:24 · Guest teaching 5/10 Model Advancements, Prompt Caching, and Poor Man's RLHF Florent explains 'poor man's RLHF' via user approval loops and vector database caching, highlighting Claude 3.5 Sonnet's superiority over GPT-4o. Alessio contributes his own prompt refactoring experience around prompt caching.34:25–37:30 · Guest teaching 6/10 The Bitter Lesson and Cognitive Architectures Florent argues for the Bitter Lesson, dismissing complex cognitive architectures and multi-agent debate wrappers. The co-host points to Replit's opposing architecture and benchmark results on o1 vs. Sonnet.37:31–40:38 · Guest teaching 7/10 Computer Use vs. API-Driven Execution Addressing Anthropic's new Computer Use release, Florent forcefully argues that APIs will dominate due to latency and reliability, using the early Android JVM/garbage-collection vs. iOS performance gap as a cautionary analogy.40:38–43:00 · Guest teaching 4/10 Operational Scaling: General Managers and the Factorio Analogy Florent explains hiring general managers to scale horizontal Lindy templates, drawing on his Uber experience. The conversation shifts to how running thousands of autonomous agents resembles Factorio factory optimization.43:00–46:50 · Guest teaching 6/10 Bottom-Up Execution vs. Top-Down Legibility Co-host reflects on breaking large complex workflows into small steps. Florent references 'Seeing Like a State', illustrating how optimizing for top-down legibility inevitably sacrifices bottom-up operational performance.46:51–51:07 · Guest teaching 6/10 The Shift from Remote Work to In-Person Collaboration Florent does an about-face on remote work despite previously founding a virtual office startup. When the co-host defends remote companies like GitLab and Automattic, Florent dismisses Automattic as a commercial failure compared to centralized alternatives.51:08–58:27 · Guest teaching 5/10 Direction vs Magnitude in Creative Engineering Florent claims tech workers outside San Francisco lack either judgment or ambition. Alessio joins in detailing his relocation from Europe, and the group critiques European venture appetite and regulatory stagnation.58:27–1:03:18 · Guest teaching 4/10 Overton Windows, Contrarianism, and AI Regulation Florent discusses pushing the Overton window on Twitter and supporting AI safety bill SB 1047, drawing a historical parallel to French resistance during WWII. Co-host calls out an inconsistency between imminent AGI timelines and claims of low stakes.1:03:18–1:05:33 · Guest teaching 5/10 P(Doom), Utopia, and Model-Layer Safety Alessio asks how safety concerns square with building agentic tooling. Florent estimates a 10% P(Doom) versus a 90% chance of post-scarcity utopia, arguing catastrophic downside risk resides at the model layer rather than the application layer.1:05:33–1:07:48 · Guest teaching 4/10 Hiring and the Primacy of Product Design The conversation concludes with hiring announcements. Co-host and Florent align on the importance of product design and end-to-end user journeys as key defensible moats in the agent ecosystem.0:55–2:59 · Guest disagreement 2/10 Lindy on Rails: Evolution of the AI Agent Architecture The hosts set up the conversation referencing Florent's previous talks and Andrew Wilkinson's adoption of Lindy. Florent explains the architectural transition to 'Lindy on Rails', acknowledging past over-reliance on pure prompting.3:00–5:57 · Guest disagreement 3/10 Deterministic Workflows vs. Pure Prompting Co-host frames the philosophy of putting 'Shoggoth in a box' and prioritizing deterministic glue code over pure natural language prompts. Florent explains why GUI and guardrails outperform text boxes for real users.5:58–8:51 · Guest disagreement 1/10 Managing Permissions, Scopes, and Agent Accounts Alessio queries how Lindy handles broad OAuth scopes and enterprise permissions. Florent details incremental permissions and explains the emerging hack of users provisioning dedicated Google Workspace accounts for their agents.8:51–13:21 · Guest disagreement 1/10 Live Demo: Automated Rental Logging and Structured Nodes Florent runs a live demo showing automated rental logging and his personal meeting recorder Lindy. The hosts ask clarifying architectural questions regarding structured output nodes and context restoration.13:21–16:13 · Guest disagreement 2/10 Agent Memory Architecture and Management Co-host presses on why Florent's Lindy has so few saved memories despite years of use. Florent explains that agent architectures degrade with excessive memory and cites conversations with Meta's Llama team on prompt engineering limits.16:13–19:38 · Guest disagreement 1/10 Modular Agents, Health Tracking, and Nested Workflows Florent showcases health logging workflows and explains the modular boundary between general personal Lindies and specialized sub-agents. Alessio clarifies how multiple Lindies invoke and chain one another.19:38–24:03 · Guest disagreement 3/10 Specialized Tools: Calendar Scheduling and PR Reviewers Alessio asks about vertical vs. horizontal SaaS boundaries. Florent presents his thesis that agents aggregate horizontally like Google Search across verticals, while the co-host playfully attempts to fish for a spicy hot take.24:03–26:15 · Guest disagreement 2/10 Building a No-Code Agent Community and User Base Co-host advises Florent on framing the user base around high-budget executive assistants rather than 'no-code tinkerers'. Florent notes that practical CEO behavior doesn't support self-serve setup without tinkering.26:16–30:29 · Guest disagreement 1/10 Customer Support Evals and the Rickroll Incident Florent recounts an incident where their customer support agent hallucinated a Rickroll link. Alessio and the co-host explore eval generation strategies, structured outputs, and the necessity of domain-specific test harnesses.30:30–34:24 · Guest disagreement 2/10 Model Advancements, Prompt Caching, and Poor Man's RLHF Florent explains 'poor man's RLHF' via user approval loops and vector database caching, highlighting Claude 3.5 Sonnet's superiority over GPT-4o. Alessio contributes his own prompt refactoring experience around prompt caching.34:25–37:30 · Guest disagreement 3/10 The Bitter Lesson and Cognitive Architectures Florent argues for the Bitter Lesson, dismissing complex cognitive architectures and multi-agent debate wrappers. The co-host points to Replit's opposing architecture and benchmark results on o1 vs. Sonnet.37:31–40:38 · Guest disagreement 3/10 Computer Use vs. API-Driven Execution Addressing Anthropic's new Computer Use release, Florent forcefully argues that APIs will dominate due to latency and reliability, using the early Android JVM/garbage-collection vs. iOS performance gap as a cautionary analogy.40:38–43:00 · Guest disagreement 1/10 Operational Scaling: General Managers and the Factorio Analogy Florent explains hiring general managers to scale horizontal Lindy templates, drawing on his Uber experience. The conversation shifts to how running thousands of autonomous agents resembles Factorio factory optimization.43:00–46:50 · Guest disagreement 2/10 Bottom-Up Execution vs. Top-Down Legibility Co-host reflects on breaking large complex workflows into small steps. Florent references 'Seeing Like a State', illustrating how optimizing for top-down legibility inevitably sacrifices bottom-up operational performance.46:51–51:07 · Guest disagreement 5/10 The Shift from Remote Work to In-Person Collaboration Florent does an about-face on remote work despite previously founding a virtual office startup. When the co-host defends remote companies like GitLab and Automattic, Florent dismisses Automattic as a commercial failure compared to centralized alternatives.51:08–58:27 · Guest disagreement 6/10 Direction vs Magnitude in Creative Engineering Florent claims tech workers outside San Francisco lack either judgment or ambition. Alessio joins in detailing his relocation from Europe, and the group critiques European venture appetite and regulatory stagnation.58:27–1:03:18 · Guest disagreement 6/10 Overton Windows, Contrarianism, and AI Regulation Florent discusses pushing the Overton window on Twitter and supporting AI safety bill SB 1047, drawing a historical parallel to French resistance during WWII. Co-host calls out an inconsistency between imminent AGI timelines and claims of low stakes.1:03:18–1:05:33 · Guest disagreement 3/10 P(Doom), Utopia, and Model-Layer Safety Alessio asks how safety concerns square with building agentic tooling. Florent estimates a 10% P(Doom) versus a 90% chance of post-scarcity utopia, arguing catastrophic downside risk resides at the model layer rather than the application layer.1:05:33–1:07:48 · Guest disagreement 1/10 Hiring and the Primacy of Product Design The conversation concludes with hiring announcements. Co-host and Florent align on the importance of product design and end-to-end user journeys as key defensible moats in the agent ecosystem.0:55–2:59 · The hosts pushing back 1/10 Lindy on Rails: Evolution of the AI Agent Architecture The hosts set up the conversation referencing Florent's previous talks and Andrew Wilkinson's adoption of Lindy. Florent explains the architectural transition to 'Lindy on Rails', acknowledging past over-reliance on pure prompting.3:00–5:57 · The hosts pushing back 2/10 Deterministic Workflows vs. Pure Prompting Co-host frames the philosophy of putting 'Shoggoth in a box' and prioritizing deterministic glue code over pure natural language prompts. Florent explains why GUI and guardrails outperform text boxes for real users.5:58–8:51 · The hosts pushing back 2/10 Managing Permissions, Scopes, and Agent Accounts Alessio queries how Lindy handles broad OAuth scopes and enterprise permissions. Florent details incremental permissions and explains the emerging hack of users provisioning dedicated Google Workspace accounts for their agents.8:51–13:21 · The hosts pushing back 1/10 Live Demo: Automated Rental Logging and Structured Nodes Florent runs a live demo showing automated rental logging and his personal meeting recorder Lindy. The hosts ask clarifying architectural questions regarding structured output nodes and context restoration.13:21–16:13 · The hosts pushing back 3/10 Agent Memory Architecture and Management Co-host presses on why Florent's Lindy has so few saved memories despite years of use. Florent explains that agent architectures degrade with excessive memory and cites conversations with Meta's Llama team on prompt engineering limits.16:13–19:38 · The hosts pushing back 1/10 Modular Agents, Health Tracking, and Nested Workflows Florent showcases health logging workflows and explains the modular boundary between general personal Lindies and specialized sub-agents. Alessio clarifies how multiple Lindies invoke and chain one another.19:38–24:03 · The hosts pushing back 3/10 Specialized Tools: Calendar Scheduling and PR Reviewers Alessio asks about vertical vs. horizontal SaaS boundaries. Florent presents his thesis that agents aggregate horizontally like Google Search across verticals, while the co-host playfully attempts to fish for a spicy hot take.24:03–26:15 · The hosts pushing back 2/10 Building a No-Code Agent Community and User Base Co-host advises Florent on framing the user base around high-budget executive assistants rather than 'no-code tinkerers'. Florent notes that practical CEO behavior doesn't support self-serve setup without tinkering.26:16–30:29 · The hosts pushing back 2/10 Customer Support Evals and the Rickroll Incident Florent recounts an incident where their customer support agent hallucinated a Rickroll link. Alessio and the co-host explore eval generation strategies, structured outputs, and the necessity of domain-specific test harnesses.30:30–34:24 · The hosts pushing back 2/10 Model Advancements, Prompt Caching, and Poor Man's RLHF Florent explains 'poor man's RLHF' via user approval loops and vector database caching, highlighting Claude 3.5 Sonnet's superiority over GPT-4o. Alessio contributes his own prompt refactoring experience around prompt caching.34:25–37:30 · The hosts pushing back 3/10 The Bitter Lesson and Cognitive Architectures Florent argues for the Bitter Lesson, dismissing complex cognitive architectures and multi-agent debate wrappers. The co-host points to Replit's opposing architecture and benchmark results on o1 vs. Sonnet.37:31–40:38 · The hosts pushing back 2/10 Computer Use vs. API-Driven Execution Addressing Anthropic's new Computer Use release, Florent forcefully argues that APIs will dominate due to latency and reliability, using the early Android JVM/garbage-collection vs. iOS performance gap as a cautionary analogy.40:38–43:00 · The hosts pushing back 1/10 Operational Scaling: General Managers and the Factorio Analogy Florent explains hiring general managers to scale horizontal Lindy templates, drawing on his Uber experience. The conversation shifts to how running thousands of autonomous agents resembles Factorio factory optimization.43:00–46:50 · The hosts pushing back 2/10 Bottom-Up Execution vs. Top-Down Legibility Co-host reflects on breaking large complex workflows into small steps. Florent references 'Seeing Like a State', illustrating how optimizing for top-down legibility inevitably sacrifices bottom-up operational performance.46:51–51:07 · The hosts pushing back 3/10 The Shift from Remote Work to In-Person Collaboration Florent does an about-face on remote work despite previously founding a virtual office startup. When the co-host defends remote companies like GitLab and Automattic, Florent dismisses Automattic as a commercial failure compared to centralized alternatives.51:08–58:27 · The hosts pushing back 2/10 Direction vs Magnitude in Creative Engineering Florent claims tech workers outside San Francisco lack either judgment or ambition. Alessio joins in detailing his relocation from Europe, and the group critiques European venture appetite and regulatory stagnation.58:27–1:03:18 · The hosts pushing back 6/10 Overton Windows, Contrarianism, and AI Regulation Florent discusses pushing the Overton window on Twitter and supporting AI safety bill SB 1047, drawing a historical parallel to French resistance during WWII. Co-host calls out an inconsistency between imminent AGI timelines and claims of low stakes.1:03:18–1:05:33 · The hosts pushing back 3/10 P(Doom), Utopia, and Model-Layer Safety Alessio asks how safety concerns square with building agentic tooling. Florent estimates a 10% P(Doom) versus a 90% chance of post-scarcity utopia, arguing catastrophic downside risk resides at the model layer rather than the application layer.1:05:33–1:07:48 · The hosts pushing back 1/10 Hiring and the Primacy of Product Design The conversation concludes with hiring announcements. Co-host and Florent align on the importance of product design and end-to-end user journeys as key defensible moats in the agent ecosystem.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 53:40 Lack of judgment or ambition outside SF

Florent bluntly asserts that anyone in tech or AI who chooses not to base themselves in San Francisco lacks either judgment or ambition.

Hardest push from the hosts ▶ 1:02:52 Pushback on AGI timeline vs stakes inconsistency

Co-host directly challenges Florent, noting an irreconcilable inconsistency between believing AGI is right around the corner and simultaneously claiming the stakes are low.

Biggest teaching moment ▶ 38:25 Android JVM vs iOS analogy for Computer Use

Florent educates the room on why API integration will beat raw computer use for years, drawing on the historical CPU and garbage collection bottlenecks between early Android and iOS.

The host holds their own ▶ 4:05 Framing Shoggoth in a minimal viable box

Co-host synthesizes the core AI engineering tension between frontier labs pushing pure prompting and developers enforcing deterministic software constraints, earning immediate praise from the guest.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Lindy on Rails: Evolution of the AI Agent Architecture 4421 The hosts set up the conversation referencing Florent's previous talks and Andrew Wilkinson's adoption of Lindy. Florent explains the architectural transition to 'Lindy on Rails', acknowledging past over-reliance on pure prompting.
Deterministic Workflows vs. Pure Prompting 6532 Co-host frames the philosophy of putting 'Shoggoth in a box' and prioritizing deterministic glue code over pure natural language prompts. Florent explains why GUI and guardrails outperform text boxes for real users.
Managing Permissions, Scopes, and Agent Accounts 5412 Alessio queries how Lindy handles broad OAuth scopes and enterprise permissions. Florent details incremental permissions and explains the emerging hack of users provisioning dedicated Google Workspace accounts for their agents.
Live Demo: Automated Rental Logging and Structured Nodes 4511 Florent runs a live demo showing automated rental logging and his personal meeting recorder Lindy. The hosts ask clarifying architectural questions regarding structured output nodes and context restoration.
Agent Memory Architecture and Management 5523 Co-host presses on why Florent's Lindy has so few saved memories despite years of use. Florent explains that agent architectures degrade with excessive memory and cites conversations with Meta's Llama team on prompt engineering limits.
Modular Agents, Health Tracking, and Nested Workflows 4411 Florent showcases health logging workflows and explains the modular boundary between general personal Lindies and specialized sub-agents. Alessio clarifies how multiple Lindies invoke and chain one another.
Specialized Tools: Calendar Scheduling and PR Reviewers 5633 Alessio asks about vertical vs. horizontal SaaS boundaries. Florent presents his thesis that agents aggregate horizontally like Google Search across verticals, while the co-host playfully attempts to fish for a spicy hot take.
Building a No-Code Agent Community and User Base 5322 Co-host advises Florent on framing the user base around high-budget executive assistants rather than 'no-code tinkerers'. Florent notes that practical CEO behavior doesn't support self-serve setup without tinkering.
Customer Support Evals and the Rickroll Incident 5512 Florent recounts an incident where their customer support agent hallucinated a Rickroll link. Alessio and the co-host explore eval generation strategies, structured outputs, and the necessity of domain-specific test harnesses.
Model Advancements, Prompt Caching, and Poor Man's RLHF 6522 Florent explains 'poor man's RLHF' via user approval loops and vector database caching, highlighting Claude 3.5 Sonnet's superiority over GPT-4o. Alessio contributes his own prompt refactoring experience around prompt caching.
The Bitter Lesson and Cognitive Architectures 6633 Florent argues for the Bitter Lesson, dismissing complex cognitive architectures and multi-agent debate wrappers. The co-host points to Replit's opposing architecture and benchmark results on o1 vs. Sonnet.
Computer Use vs. API-Driven Execution 5732 Addressing Anthropic's new Computer Use release, Florent forcefully argues that APIs will dominate due to latency and reliability, using the early Android JVM/garbage-collection vs. iOS performance gap as a cautionary analogy.
Operational Scaling: General Managers and the Factorio Analogy 4411 Florent explains hiring general managers to scale horizontal Lindy templates, drawing on his Uber experience. The conversation shifts to how running thousands of autonomous agents resembles Factorio factory optimization.
Bottom-Up Execution vs. Top-Down Legibility 5622 Co-host reflects on breaking large complex workflows into small steps. Florent references 'Seeing Like a State', illustrating how optimizing for top-down legibility inevitably sacrifices bottom-up operational performance.
The Shift from Remote Work to In-Person Collaboration 5653 Florent does an about-face on remote work despite previously founding a virtual office startup. When the co-host defends remote companies like GitLab and Automattic, Florent dismisses Automattic as a commercial failure compared to centralized alternatives.
Direction vs Magnitude in Creative Engineering 5562 Florent claims tech workers outside San Francisco lack either judgment or ambition. Alessio joins in detailing his relocation from Europe, and the group critiques European venture appetite and regulatory stagnation.
Overton Windows, Contrarianism, and AI Regulation 5466 Florent discusses pushing the Overton window on Twitter and supporting AI safety bill SB 1047, drawing a historical parallel to French resistance during WWII. Co-host calls out an inconsistency between imminent AGI timelines and claims of low stakes.
P(Doom), Utopia, and Model-Layer Safety 5533 Alessio asks how safety concerns square with building agentic tooling. Florent estimates a 10% P(Doom) versus a 90% chance of post-scarcity utopia, arguing catastrophic downside risk resides at the model layer rather than the application layer.
Hiring and the Primacy of Product Design 4411 The conversation concludes with hiring announcements. Co-host and Florent align on the importance of product design and end-to-end user journeys as key defensible moats in the agent ecosystem.

Statements from this episode (42)

Disclosure
Florent Crivello: Lindy is to LangChain as Airtable is to MySQL
“We are a no-code platform letting you build your own AI agents easily. So you can think of, we are to Langchain as Airtable to MySQL. Like you can just pin up AI agents super easily by clicking around and no code required. You didn't have to be an engineer and…”
Florent Crivello Nov 15, 2024 ▶ 0:36
Insight
Crivello: Early AI Agent Builders Overestimated LLM Capabilities
“I think many people, when they started working on agents, they were very LLM-peeled and ChatGPT-peeled, right? They got ahead of themselves in a way, and as included, and they thought that agents were actually and LLMs were actually more advanced than they act…”
Florent Crivello Nov 15, 2024 ▶ 1:55
Insight
Crivello: Putting AI Agents on Rails Maximizes Reliability and Usability
“The more you can put your agent on rails, one, the more reliable it's going to be, obviously, but two, it's also going to be easier to use for the user, because you can really, as a user, you get, instead of just getting this, like, big, giant, intimidating te…”
Florent Crivello Nov 15, 2024 ▶ 2:14
Assertion Not checkable as stated
Crivello: 60% to 70% of user agent prompts are meaningless
“Some, I would say, 60 or 70% of the prompts that people type don't mean anything. Me, as a human, as AGI, I don't understand what they mean. I don't know what they mean.”
Florent Crivello Nov 15, 2024 ▶ 5:44
Insight
Crivello: A GUI is always better than pure text for agent workflows
“It is actually, I think, whenever you can have a GUI, it is better than to have just a pure text interface.”
Florent Crivello Nov 15, 2024 ▶ 5:52
Assertion Not checkable as stated
Crivello: Users Add Delays to Fake Google Accounts to Disguise AI
“One pattern we've started to see is people provisioning accounts for their AI agents. And so in particular, Google Workspace accounts. So for example, Lindy can be used as a scheduling assistant. And so you can just CC her To your emails when you're trying to …”
Florent Crivello Nov 15, 2024 ▶ 7:38
Prediction Not checkable as stated
Crivello: Tech industry will not fix OAuth for AI agents before AGI
“I'm not optimistic On us actually patching OAuth, because I agree with you ultimately, like we would want to patch OAuth, because the new account thing is kind of a cludge. It's really a hack. You would want to patch OAuth to have more granular access control …”
Florent Crivello Nov 15, 2024 ▶ 8:15
Assertion Supported
Crivello: Lindy lets users configure different AI models per workflow node
“And you can see here for every node, you can configure which model you want to power the node.”
Florent Crivello Nov 15, 2024 ▶ 10:42
Disclosure
Crivello Often Skips Meetings and Queries His Lindy Agent Instead
“Because I know I can have this sort of interactive Q&A with these meetings, it means that very often now I don't go to meetings anymore. I just send my Lindy. And instead of going to like a six-minute meeting, I have like a five-minute chat with my Lindy after…”
Florent Crivello Nov 15, 2024 ▶ 12:44
Assertion Not checkable as stated
Crivello: Top Lindy Use Cases Are Sales, Support, and Personal Assistance
“There's a lot of like sales automations, customer support automations, and a lot of this, which is basically personal assistance automations, like meeting scheduling and so forth.”
Florent Crivello Nov 15, 2024 ▶ 13:14
Insight
Crivello: AI agents get confused with too many memories
“Agents still get confused if they have too many memories, to my point earlier about that.”
Florent Crivello Nov 15, 2024 ▶ 14:23
Insight
Crivello: Separate agents by job to be done and target audience
“I think of it in terms of like jobs to be done, and I think of it in terms of who is the Lindy serving.”
Florent Crivello Nov 15, 2024 ▶ 18:58
Assertion Not checkable as stated
Crivello: Lindy's calendar availability action requires roughly 1,000 lines of code
“So for example, one that is surprisingly hard is this find available times action. You would not believe this is like a thousand lines of code or something. It's just a very beefy action.”
Florent Crivello Nov 15, 2024 ▶ 20:30
Disclosure
Crivello: Lindy uses an AI PR reviewer reading 40-page bug guidelines
“We have like a Lindy PR reviewer. And it's really handy because anytime any bug happens, so the Lindy reads our guidelines on Google Docs. By now the guidelines are like 40 pages long or something. And so every time any new kind of bug happens, we just go to t…”
Florent Crivello Nov 15, 2024 ▶ 20:58
Insight
Crivello: Horizontal Platforms Win Because Core Infrastructure Is Universal Across Verticals
“And I think that the reason for that is because search in each vertical has more in common with search than it does with each vertical. And search is so expensive to get right. Like Google is a big company. That it makes a lot of sense to aggregate all of thes…”
Florent Crivello Nov 15, 2024 ▶ 22:09
Prediction Not checkable as stated
Crivello: A Few Very Big Horizontal AI Agent Platforms Will Dominate
“I think AI is going to completely penetrate every category of software, but then I also think there are going to be a few very, very, very big horizontal agents that serve a lot of functions for people.”
Florent Crivello Nov 15, 2024 ▶ 23:04
Insight
Crivello: Vertical AI Tools Emerge First Because Horizontal Platforms Are Harder
“I also think it is natural for the vertical solutions to emerge first, because they're just easier to build. It's just much, much, much harder to build something horizontal.”
Florent Crivello Nov 15, 2024 ▶ 23:55
Opinion
Crivello: AI Agent Community Will Mirror Webflow, Zapier, and Airtable Users
“I think the no-code tinkerers is the community. Yeah. It is going to be the same sort of community as what, yeah, Webflow, Zapier, Airtable, Notion, to some extent.”
Florent Crivello Nov 15, 2024 ▶ 25:35
Insight
Crivello: Selling AI EAs to CEOs Fails Because CEOs Won't Tinker
“The problem with EA is, like, the CEO has no willingness to actually tinker and play with the platform.”
Florent Crivello Nov 15, 2024 ▶ 25:59
Disclosure
Crivello: Lindy Uses Its Own AI Agents for Customer Support
“So we have Lindy obviously doing our customer support and we do check after the Lindy.”
Florent Crivello Nov 15, 2024 ▶ 26:30
Assertion Not checkable as stated
Crivello: Lindy Found 3–4 Instances of Agent Rickrolling Customers
“And I checked afterwards in, in the logs, we did like a database query and we found like, I think like three or four other instances of it.”
Florent Crivello Nov 15, 2024 ▶ 27:05
Disclosure
Crivello: Lindy will likely switch to Braintrust for AI evaluations
“We're most likely going to switch to Braintrust.”
Florent Crivello Nov 15, 2024 ▶ 28:44
Insight
Crivello: Do not build custom eval tools; use existing solutions instead
“I wouldn't recommend it to build your own eval tool. There's better solutions out there, and our eval tool breaks all the time, and it's a nightmare to maintain, and that's not something we want to be spending our time on.”
Florent Crivello Nov 15, 2024 ▶ 28:56
Assertion Not checkable as stated
Crivello: Lindy AI agents perform better than humans for most use cases
“I think the bar is it needs to be better than a human. And for most use cases we serve today, it is better than the human, especially if you put it on rails.”
Florent Crivello Nov 15, 2024 ▶ 30:22
Opinion
Crivello: OpenAI's GPT-4o Is Overhyped and Poor for Agents
“I think four O is overhyped. Frankly, we don't use four O. I don't think it's good for agentic behavior.”
Florent Crivello Nov 15, 2024 ▶ 32:07
Prediction Not checkable as stated
Crivello: AI Agent Software Will Create Trillions Replacing Human Labor
“It's going to sound like an exaggeration, but it is a fact that it's going to create trillions of dollars of value in a few years, right? It's going to, for the first time, we're actually having software directly replace human labor.”
Florent Crivello Nov 15, 2024 ▶ 33:52
Insight
Crivello: The Bitter Lesson applies to agent cognitive architectures
“I think people fail to truly, and me included, they fail to truly internalize the bitter lesson. So for the listeners out there who don't know about it, it's basically like you just scale the model, like GPUs go brrr, it's all that matters. I think it also hol…”
Florent Crivello Nov 15, 2024 ▶ 34:34
Prediction Not checkable as stated
Crivello: Infinite and cheap context windows will arrive within 18 months
“Now we just assume that infinite context windows are going to be here in a year or something, a year and a half and infinitely cheap as well. And dynamic compute is going to be here. Like we just assume all of these things are going to happen.”
Florent Crivello Nov 15, 2024 ▶ 36:16
Insight
Crivello: Tasks capable of using APIs must stay API-driven over computer use
“My philosophy about it is anything that can be done with an API must be done by an API or should be done by an API for a very long time.”
Florent Crivello Nov 15, 2024 ▶ 38:16
Prediction Not checkable as stated
Crivello: Future businesses will operate like Factorio with thousands of AI agents
“We actually very often talk about how the business of the future is like a game of factorial. It's like you just wake up in the morning and you've got your Lindy instance. It's like Slack and you've got like 5000 Lindys in the sidebar and your job is to someho…”
Florent Crivello Nov 15, 2024 ▶ 41:59
Insight
Crivello: Moonshots require greedy search execution rather than rigid planning
“As long as you have some inductive bias about like some loose idea about where you want to go, I think it makes sense to follow a sort of greedy search along that path.”
Florent Crivello Nov 15, 2024 ▶ 43:52
Insight
Crivello: Top-down legibility inherently degrades bottom-up system performance
“Anytime you make a system more understandable from the top down, it performs less well from the bottom up. And it's fine if that's what you want, but you should at least make this trade off with your eyes wide open. You should know I am sacrificing performance…”
Florent Crivello Nov 15, 2024 ▶ 45:24
Assertion Not checkable as stated
Crivello: Every major remote success story lost to an in-person competitor
“In every one of these examples, you have a co-located counterfactual that is sometimes orders of magnitude bigger.”
Florent Crivello Nov 15, 2024 ▶ 49:48
Opinion
Crivello: WordPress is a commercial failure despite powering 60% of the web
“WordPress is a commercial failure. They run 60% of the internet, and they're, like, a fraction of the size of even Substack, right?”
Florent Crivello Nov 15, 2024 ▶ 49:58
Insight
Crivello: Remote work optimizes for cost; building creative software requires in-person collaboration
“If you're optimizing for cost, absolutely be remote. If you're optimizing for creativity, which I think that software and product building is a creative endeavor, if you're optimizing for creativity, it's kind of like composing an album. You can't do it on the…”
Florent Crivello Nov 15, 2024 ▶ 50:36
Insight
Crivello: Creative work defines direction, while execution focuses on magnitude
“By creativity, I mean, is it about direction or magnitude? If it is about direction, like decide what to do, then it's a creative endeavor. If it is about magnitude and just do it as fast as possible, as cheap as possible, then it's magnitude.”
Florent Crivello Nov 15, 2024 ▶ 51:46
Opinion
Crivello: Linear executes on a known category rather than novel software
“Linear. I look up to a huge amount, like such amazing product builders, but they know what they're building. They're building a task tracker.”
Florent Crivello Nov 15, 2024 ▶ 52:04
Opinion
Crivello: Tech workers outside SF lack judgment or ambition
“I think at this point, if you are in tech, especially in AI, but if you're in tech and you're not in San Francisco, you either lack judgment or you lack ambition. It's one of the two.”
Florent Crivello Nov 15, 2024 ▶ 53:34
Insight
Crivello: San Francisco's tech network effects cannot be broken
“Once there is a network effects that are just so incredibly powerful, they can't be broken, really. And we tried with San Francisco, I tried with San Francisco, like during COVID, there was a movement of people moving to Miami. Co-Host (Founder of Smol AI): Yo…”
Florent Crivello Nov 15, 2024 ▶ 56:17
Assertion Not checkable as stated
Crivello: Marc Andreessen Blocked Him on Twitter Over AI Safety Views
“Like at some point, Marc Andreessen blocked me on Twitter and I, it hurt, frankly, I really look up to Marc Andreessen and I knew he would block me.”
Florent Crivello Nov 15, 2024 ▶ 1:01:05
Prediction Not checkable as stated
Crivello: AI has 10% p(doom) and 90% chance of disease-free utopia
“My PDOOM, insofar as I can quantify it, which I cannot, But if I had to, like, my vibe is like, 10% or something like that? And so there's also like a 90% chance that we live in like a pure utopia, right? And that's awesome, right? So like, let's go after the …”
Florent Crivello Nov 15, 2024 ▶ 1:04:31
Opinion
Crivello: AI risks exist at model layer, not application layer
“I know that it's very self-serving to say, oh, you know, like the downside doesn't exist at my layer, it exists at like the model layer. But truly, look at Lindy, look at the Apple building. I struggle to see exactly how it would like get up and start doing cr…”
Florent Crivello Nov 15, 2024 ▶ 1:05:06
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.