May 29, 2025 · 48m · mad
AI That Ends Busy Work — Hebbia CEO on “Agent Employees”
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast recorded live at Data-Driven NYC, Hebbia Founder and CEO George Sivulka joins host Matt Turck to explore the future of enterprise AI, detailing how AI is evolving from basic chatbots into autonomous agent employees that reshape organizational workflows across finance, law, and corporate knowledge work.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 19.2% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
George forcefully rejects the common venture capital thesis of vertical AI, calling verticalization a massive fallacy and insisting generalization wins every time.
Hardest push from Matt ▶ 13:44 Reality check on GDP agent replacement claimsMatt refuses to accept George's aggressive timeline for agent enterprise deployment without pushing for a reality check on overhyped capabilities versus real-world adoption.
Biggest teaching moment ▶ 30:31 Abandoning a world-record re-ranker for LLM computeGeorge educates Matt on why classical search metrics fail for complex reasoning, revealing that Hebbia shelved its record-breaking re-ranker model in favor of heavy LLM compute loops.
Matt holds his own ▶ 30:22 Probing classical RAG re-ranker componentsMatt shows strong technical grounding in AI retrieval infrastructure by explicitly probing whether Hebbia still relies on classical re-rankers like ColBERT.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Stage Reconnection and Hebbia's Elevator Pitch | 3 | 3 | 1 | 1 | Matt warmly reconnects with George, recalling their interview two years prior while asking for an updated elevator pitch. George explains Hebbia's evolution into a horizontal workplace AI platform beyond initial financial and legal verticals. | |
| Leaving Stanford PhD for GPT-3 | 2 | 4 | 2 | 1 | Matt playfully teases George about his young age and early Stanford graduation before asking about leaving his PhD. George shares his lightbulb moment using GPT-3 in June 2020, realizing it was a meta-learner that surpassed his research trajectory. | |
| Hebbia's Trajectory, Scale, and Page Processing | 2 | 5 | 2 | 1 | Matt asks for company metrics and scale. George jokingly references VC confidentiality constraints before revealing that Hebbia went from processing 100 million pages last year to an expected 4 to 5 billion pages this year. | |
| The Concept of the Agent Employee | 3 | 6 | 1 | 1 | Matt introduces George's recent blog post on the 'agent employee'. George articulates how organizational design is shifting to treat AI agents as distinct nodes in org charts with emails, Slacks, and dedicated workflows. | |
| Prompting as Management and Enterprise Adoption Curves | 4 | 5 | 2 | 4 | Matt presses for a reality check on bold claims that AI agents will contribute more to GDP than humans within a decade. George acknowledges enterprise adoption latency using the transition from cash to credit cards as an analogy. | |
| Test-Time Compute and Extending Context Windows | 5 | 6 | 1 | 3 | Matt drills down into current AI research, specifically double-clicking on how to artificially extend LLM context windows. George explains inference-time compute scaling and how Hebbia pioneered multi-call runtime infrastructure. | |
| Matrix Architecture and Multi-Agent Grid Workflows | 4 | 7 | 6 | 2 | George directly attacks standard venture consensus by declaring vertical AI a 'massive fallacy'. He argues forcefully that generalization always beats specialization because top experts pull insights from outside their core domains. | |
| Spreadsheets as the True Paradigm for Knowledge Work | 3 | 6 | 2 | 2 | Matt highlights George's skepticism toward chatbot UIs. George compares chatbots to TI-84 calculators, asserting that flexible spreadsheets are the true paradigm through which knowledge workers operate. | |
| Technical Deep Dive: Ingestion, Indexing, and ISD | 6 | 7 | 2 | 3 | Matt initiates a technical deep dive into ingestion, indexing, and ISD vs standard RAG architecture. George reveals that Hebbia developed a world-record accurate re-ranker but discarded it because keyword/vector retrieval fails for deep agent reasoning. | |
| Multi-Model Strategy and Maximizer Router | 4 | 6 | 2 | 2 | Matt asks about multi-model usage and system scaling tricks. George explains 'Maximizer', an air-traffic-controller system that routes requests to achieve theoretical maximum throughput across 250 billion monthly LLM calls. | |
| Addressing Hallucinations and the Cost of Accuracy | 4 | 6 | 6 | 3 | Matt brings up model hallucinations in high-stakes industries like finance and law. George forcefully rejects the concern as 'old news' and 'fugazi', claiming models are already vastly superior to human workers when provided proper context. | |
| The State of AI Innovation and Startup Alpha | 5 | 6 | 5 | 3 | Matt prompts George on whether fundamental AI research is slowing down. George gives a contrarian response that 'alpha is gone' for starting new AI companies and criticizes traditional SaaS GTM playbooks in favor of hiring domain consultants. | |
| Building Value Cases and Calculating ROI for Enterprise AI | 4 | 5 | 1 | 2 | Matt asks how buyers evaluate ROI and price justifications for enterprise AI. George explains that while cost savings exist, the primary value driver is revenue expansion by giving workers unlimited expert analysis capacity. | |
| The Impact of AI on Junior Roles and Knowledge Work | 4 | 6 | 2 | 2 | Matt asks if clients are reducing junior analyst hires and inquires how George leads as a young founder. George uses Morgan Stanley's historical creation of the analyst role during the computer era to argue AI will evolve junior work rather than destroy it. | |
| Future Outlook: From Chatbots to Agentic AI Applications | 2 | 3 | 0 | 0 | Matt asks for a 3-year outlook to conclude the interview. George shares his core aspiration of transitioning users worldwide from simple chatbots to agentic, value-creating applications. |