Jun 21, 2023 · 31m · mad
AI and the Future of Knowledge Work with Hebbia’s CEO George Sivulka
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Data Driven NYC presentation, Hebbia Founder and CEO George Sivulka joins FirstMark Partner Matt Turck to discuss Hebbia's emergence as an LLM-native productivity tool, its technical architecture, and its potential to revolutionize complex document workflows across knowledge work.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.4% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
George explicitly pushes back against the audience member's suggestion that AI should autonomously surface unprompted insights, insisting that AI requires a human-defined principal component of meaning.
Hardest push from Matt ▶ 12:04 Host Probing Job Replacement vs. CopilotMatt presses George on whether Hebbia is merely aiding analysts as a copilot or actually enabling enterprise customers to eliminate analyst headcount.
Biggest teaching moment ▶ 19:25 Math Breakdown of Multi-Step LLM Compound ErrorsGeorge educates the audience on LLM reliability math, demonstrating how chaining steps with 90% accuracy results in compound error rates that destroy utility on complex tasks.
Matt holds his own ▶ 16:20 Host Demonstrating Specific Product Feature KnowledgeMatt showcases thorough prep by specifically naming technical platform capabilities including question lists, table extraction, and semantic alerts.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| George Sivulka's Academic Background and the Genesis of Hebbia | 2 | 2 | 1 | 0 | Matt opens with a warm, well-researched introduction citing Hebbia's funding history and notable investors before asking George about his journey. George shares a narrative about observing Stanford students at investment banks doing tedious labor and building early prototypes using DEFM-FORTEENA filings. | |
| Defining Hebbia as an LLM-Native Productivity Tool and the Excel Analogy | 3 | 3 | 1 | 1 | Matt asks pragmatic questions about product delivery, model selection, and enterprise customization. George re-frames Hebbia's mission around the 1985 Excel analogy, explaining how LLM-native productivity tools allow non-technical users to programmatically wrangle models. | |
| Target Markets: Private Equity Due Diligence and Human-in-the-Loop AI | 3 | 2 | 2 | 3 | Matt probes whether Hebbia acts purely as a copilot or if it actively replaces human analysts in private equity due diligence. George clarifies his product philosophy, arguing that human-in-the-loop tools amplify human productivity rather than replace jobs. | |
| Technical Architecture: Engineering Input, Output, and Reliability | 2 | 4 | 1 | 1 | Matt asks about Hebbia's underlying technical architecture and defensive moat. George provides commentary on OpenAI's function calling API released that same day, explaining how Hebbia focuses R&D on input, output, and workflow flexibility to avoid model obsolescence. | |
| Key Platform Features: Question Lists, Document Comparison, and Semantic Alerts | 4 | 2 | 0 | 0 | Matt demonstrates prep work by citing specific platform capabilities including question lists, table extraction, and semantic alerts. George elaborates on how these features serve specific workflow needs like CEO-level contract tracking. | |
| Addressing AI Bottlenecks: Multi-Step Reasoning and Compound Errors | 3 | 5 | 2 | 1 | Matt inquires about AI limitations and technical bottlenecks. George breaks down why simple prompts and naive orchestration frameworks fail, illustrating how multi-step LLM operations suffer from compounding errors across nth-order tasks. | |
| Hebbia's Long-Term Roadmap and Vision Beyond Finance | 2 | 4 | 1 | 0 | Matt asks about Hebbia's long-term roadmap before opening to audience questions. An audience member asks about building user trust in multi-step AI outputs, prompting George to explain why raw cosine similarity scores confused users and required hybrid retrieval UX. | |
| Audience Q&A: UI/UX Innovation and Proactive AI vs Human-in-the-Loop | 0 | 4 | 3 | 0 | In response to an audience member asking if AI will proactively reveal unprompted insights, George rejects the premise, asserting that unprompted outlier detection is unhelpful without human-defined intent. | |
| Audience Q&A: Differentiating Hebbia from Enterprise Search Solutions Like Glean | 1 | 3 | 2 | 0 | An audience member asks how Hebbia differentiates from enterprise search tools like Glean. George offers respect for Glean's search focus while clarifying that Hebbia is built as a general-purpose analytical productivity tool for research. |