Nov 15, 2024 · 1h 8m · latent-space
Agents @ Work: Lindy.ai (with live demo!)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
On the Latent Space Podcast, Lindy.ai founder and CEO Florent Crivello explores the evolution of AI agents from fragile prompts to deterministic no-code workflows, delivers comprehensive live product demonstrations, and shares strategic insights on startup building, API-first automation, and AI governance.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Florent bluntly asserts that anyone in tech or AI who chooses not to base themselves in San Francisco lacks either judgment or ambition.
Hardest push from the hosts ▶ 1:02:52 Pushback on AGI timeline vs stakes inconsistencyCo-host directly challenges Florent, noting an irreconcilable inconsistency between believing AGI is right around the corner and simultaneously claiming the stakes are low.
Biggest teaching moment ▶ 38:25 Android JVM vs iOS analogy for Computer UseFlorent educates the room on why API integration will beat raw computer use for years, drawing on the historical CPU and garbage collection bottlenecks between early Android and iOS.
The host holds their own ▶ 4:05 Framing Shoggoth in a minimal viable boxCo-host synthesizes the core AI engineering tension between frontier labs pushing pure prompting and developers enforcing deterministic software constraints, earning immediate praise from the guest.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Lindy on Rails: Evolution of the AI Agent Architecture | 4 | 4 | 2 | 1 | The hosts set up the conversation referencing Florent's previous talks and Andrew Wilkinson's adoption of Lindy. Florent explains the architectural transition to 'Lindy on Rails', acknowledging past over-reliance on pure prompting. | |
| Deterministic Workflows vs. Pure Prompting | 6 | 5 | 3 | 2 | Co-host frames the philosophy of putting 'Shoggoth in a box' and prioritizing deterministic glue code over pure natural language prompts. Florent explains why GUI and guardrails outperform text boxes for real users. | |
| Managing Permissions, Scopes, and Agent Accounts | 5 | 4 | 1 | 2 | Alessio queries how Lindy handles broad OAuth scopes and enterprise permissions. Florent details incremental permissions and explains the emerging hack of users provisioning dedicated Google Workspace accounts for their agents. | |
| Live Demo: Automated Rental Logging and Structured Nodes | 4 | 5 | 1 | 1 | Florent runs a live demo showing automated rental logging and his personal meeting recorder Lindy. The hosts ask clarifying architectural questions regarding structured output nodes and context restoration. | |
| Agent Memory Architecture and Management | 5 | 5 | 2 | 3 | Co-host presses on why Florent's Lindy has so few saved memories despite years of use. Florent explains that agent architectures degrade with excessive memory and cites conversations with Meta's Llama team on prompt engineering limits. | |
| Modular Agents, Health Tracking, and Nested Workflows | 4 | 4 | 1 | 1 | Florent showcases health logging workflows and explains the modular boundary between general personal Lindies and specialized sub-agents. Alessio clarifies how multiple Lindies invoke and chain one another. | |
| Specialized Tools: Calendar Scheduling and PR Reviewers | 5 | 6 | 3 | 3 | Alessio asks about vertical vs. horizontal SaaS boundaries. Florent presents his thesis that agents aggregate horizontally like Google Search across verticals, while the co-host playfully attempts to fish for a spicy hot take. | |
| Building a No-Code Agent Community and User Base | 5 | 3 | 2 | 2 | Co-host advises Florent on framing the user base around high-budget executive assistants rather than 'no-code tinkerers'. Florent notes that practical CEO behavior doesn't support self-serve setup without tinkering. | |
| Customer Support Evals and the Rickroll Incident | 5 | 5 | 1 | 2 | Florent recounts an incident where their customer support agent hallucinated a Rickroll link. Alessio and the co-host explore eval generation strategies, structured outputs, and the necessity of domain-specific test harnesses. | |
| Model Advancements, Prompt Caching, and Poor Man's RLHF | 6 | 5 | 2 | 2 | Florent explains 'poor man's RLHF' via user approval loops and vector database caching, highlighting Claude 3.5 Sonnet's superiority over GPT-4o. Alessio contributes his own prompt refactoring experience around prompt caching. | |
| The Bitter Lesson and Cognitive Architectures | 6 | 6 | 3 | 3 | Florent argues for the Bitter Lesson, dismissing complex cognitive architectures and multi-agent debate wrappers. The co-host points to Replit's opposing architecture and benchmark results on o1 vs. Sonnet. | |
| Computer Use vs. API-Driven Execution | 5 | 7 | 3 | 2 | Addressing Anthropic's new Computer Use release, Florent forcefully argues that APIs will dominate due to latency and reliability, using the early Android JVM/garbage-collection vs. iOS performance gap as a cautionary analogy. | |
| Operational Scaling: General Managers and the Factorio Analogy | 4 | 4 | 1 | 1 | Florent explains hiring general managers to scale horizontal Lindy templates, drawing on his Uber experience. The conversation shifts to how running thousands of autonomous agents resembles Factorio factory optimization. | |
| Bottom-Up Execution vs. Top-Down Legibility | 5 | 6 | 2 | 2 | Co-host reflects on breaking large complex workflows into small steps. Florent references 'Seeing Like a State', illustrating how optimizing for top-down legibility inevitably sacrifices bottom-up operational performance. | |
| The Shift from Remote Work to In-Person Collaboration | 5 | 6 | 5 | 3 | Florent does an about-face on remote work despite previously founding a virtual office startup. When the co-host defends remote companies like GitLab and Automattic, Florent dismisses Automattic as a commercial failure compared to centralized alternatives. | |
| Direction vs Magnitude in Creative Engineering | 5 | 5 | 6 | 2 | Florent claims tech workers outside San Francisco lack either judgment or ambition. Alessio joins in detailing his relocation from Europe, and the group critiques European venture appetite and regulatory stagnation. | |
| Overton Windows, Contrarianism, and AI Regulation | 5 | 4 | 6 | 6 | Florent discusses pushing the Overton window on Twitter and supporting AI safety bill SB 1047, drawing a historical parallel to French resistance during WWII. Co-host calls out an inconsistency between imminent AGI timelines and claims of low stakes. | |
| P(Doom), Utopia, and Model-Layer Safety | 5 | 5 | 3 | 3 | Alessio asks how safety concerns square with building agentic tooling. Florent estimates a 10% P(Doom) versus a 90% chance of post-scarcity utopia, arguing catastrophic downside risk resides at the model layer rather than the application layer. | |
| Hiring and the Primacy of Product Design | 4 | 4 | 1 | 1 | The conversation concludes with hiring announcements. Co-host and Florent align on the importance of product design and end-to-end user journeys as key defensible moats in the agent ecosystem. |