Jun 21, 2023 · 31m · mad

AI and the Future of Knowledge Work with Hebbia’s CEO George Sivulka

George Sivulka · 21m spoken Matt Turck · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC presentation, Hebbia Founder and CEO George Sivulka joins FirstMark Partner Matt Turck to discuss Hebbia's emergence as an LLM-native productivity tool, its technical architecture, and its potential to revolutionize complex document workflows across knowledge work.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.4% of the talking time here. How this is scored →

Matt as informed peer 2.2 Guest teaching 3.2 Guest disagreement 1.4 Matt pushing back 0.7
05100:0010:0020:0030:000:11–5:18 · Matt as informed peer 2/10 George Sivulka's Academic Background and the Genesis of Hebbia Matt opens with a warm, well-researched introduction citing Hebbia's funding history and notable investors before asking George about his journey. George shares a narrative about observing Stanford students at investment banks doing tedious labor and building early prototypes using DEFM-FORTEENA filings.5:18–9:33 · Matt as informed peer 3/10 Defining Hebbia as an LLM-Native Productivity Tool and the Excel Analogy Matt asks pragmatic questions about product delivery, model selection, and enterprise customization. George re-frames Hebbia's mission around the 1985 Excel analogy, explaining how LLM-native productivity tools allow non-technical users to programmatically wrangle models.9:33–14:04 · Matt as informed peer 3/10 Target Markets: Private Equity Due Diligence and Human-in-the-Loop AI Matt probes whether Hebbia acts purely as a copilot or if it actively replaces human analysts in private equity due diligence. George clarifies his product philosophy, arguing that human-in-the-loop tools amplify human productivity rather than replace jobs.14:04–16:20 · Matt as informed peer 2/10 Technical Architecture: Engineering Input, Output, and Reliability Matt asks about Hebbia's underlying technical architecture and defensive moat. George provides commentary on OpenAI's function calling API released that same day, explaining how Hebbia focuses R&D on input, output, and workflow flexibility to avoid model obsolescence.16:20–18:27 · Matt as informed peer 4/10 Key Platform Features: Question Lists, Document Comparison, and Semantic Alerts Matt demonstrates prep work by citing specific platform capabilities including question lists, table extraction, and semantic alerts. George elaborates on how these features serve specific workflow needs like CEO-level contract tracking.18:27–21:03 · Matt as informed peer 3/10 Addressing AI Bottlenecks: Multi-Step Reasoning and Compound Errors Matt inquires about AI limitations and technical bottlenecks. George breaks down why simple prompts and naive orchestration frameworks fail, illustrating how multi-step LLM operations suffer from compounding errors across nth-order tasks.21:03–25:36 · Matt as informed peer 2/10 Hebbia's Long-Term Roadmap and Vision Beyond Finance Matt asks about Hebbia's long-term roadmap before opening to audience questions. An audience member asks about building user trust in multi-step AI outputs, prompting George to explain why raw cosine similarity scores confused users and required hybrid retrieval UX.25:36–29:48 · Matt as informed peer 0/10 Audience Q&A: UI/UX Innovation and Proactive AI vs Human-in-the-Loop In response to an audience member asking if AI will proactively reveal unprompted insights, George rejects the premise, asserting that unprompted outlier detection is unhelpful without human-defined intent.29:48–31:24 · Matt as informed peer 1/10 Audience Q&A: Differentiating Hebbia from Enterprise Search Solutions Like Glean An audience member asks how Hebbia differentiates from enterprise search tools like Glean. George offers respect for Glean's search focus while clarifying that Hebbia is built as a general-purpose analytical productivity tool for research.0:11–5:18 · Guest teaching 2/10 George Sivulka's Academic Background and the Genesis of Hebbia Matt opens with a warm, well-researched introduction citing Hebbia's funding history and notable investors before asking George about his journey. George shares a narrative about observing Stanford students at investment banks doing tedious labor and building early prototypes using DEFM-FORTEENA filings.5:18–9:33 · Guest teaching 3/10 Defining Hebbia as an LLM-Native Productivity Tool and the Excel Analogy Matt asks pragmatic questions about product delivery, model selection, and enterprise customization. George re-frames Hebbia's mission around the 1985 Excel analogy, explaining how LLM-native productivity tools allow non-technical users to programmatically wrangle models.9:33–14:04 · Guest teaching 2/10 Target Markets: Private Equity Due Diligence and Human-in-the-Loop AI Matt probes whether Hebbia acts purely as a copilot or if it actively replaces human analysts in private equity due diligence. George clarifies his product philosophy, arguing that human-in-the-loop tools amplify human productivity rather than replace jobs.14:04–16:20 · Guest teaching 4/10 Technical Architecture: Engineering Input, Output, and Reliability Matt asks about Hebbia's underlying technical architecture and defensive moat. George provides commentary on OpenAI's function calling API released that same day, explaining how Hebbia focuses R&D on input, output, and workflow flexibility to avoid model obsolescence.16:20–18:27 · Guest teaching 2/10 Key Platform Features: Question Lists, Document Comparison, and Semantic Alerts Matt demonstrates prep work by citing specific platform capabilities including question lists, table extraction, and semantic alerts. George elaborates on how these features serve specific workflow needs like CEO-level contract tracking.18:27–21:03 · Guest teaching 5/10 Addressing AI Bottlenecks: Multi-Step Reasoning and Compound Errors Matt inquires about AI limitations and technical bottlenecks. George breaks down why simple prompts and naive orchestration frameworks fail, illustrating how multi-step LLM operations suffer from compounding errors across nth-order tasks.21:03–25:36 · Guest teaching 4/10 Hebbia's Long-Term Roadmap and Vision Beyond Finance Matt asks about Hebbia's long-term roadmap before opening to audience questions. An audience member asks about building user trust in multi-step AI outputs, prompting George to explain why raw cosine similarity scores confused users and required hybrid retrieval UX.25:36–29:48 · Guest teaching 4/10 Audience Q&A: UI/UX Innovation and Proactive AI vs Human-in-the-Loop In response to an audience member asking if AI will proactively reveal unprompted insights, George rejects the premise, asserting that unprompted outlier detection is unhelpful without human-defined intent.29:48–31:24 · Guest teaching 3/10 Audience Q&A: Differentiating Hebbia from Enterprise Search Solutions Like Glean An audience member asks how Hebbia differentiates from enterprise search tools like Glean. George offers respect for Glean's search focus while clarifying that Hebbia is built as a general-purpose analytical productivity tool for research.0:11–5:18 · Guest disagreement 1/10 George Sivulka's Academic Background and the Genesis of Hebbia Matt opens with a warm, well-researched introduction citing Hebbia's funding history and notable investors before asking George about his journey. George shares a narrative about observing Stanford students at investment banks doing tedious labor and building early prototypes using DEFM-FORTEENA filings.5:18–9:33 · Guest disagreement 1/10 Defining Hebbia as an LLM-Native Productivity Tool and the Excel Analogy Matt asks pragmatic questions about product delivery, model selection, and enterprise customization. George re-frames Hebbia's mission around the 1985 Excel analogy, explaining how LLM-native productivity tools allow non-technical users to programmatically wrangle models.9:33–14:04 · Guest disagreement 2/10 Target Markets: Private Equity Due Diligence and Human-in-the-Loop AI Matt probes whether Hebbia acts purely as a copilot or if it actively replaces human analysts in private equity due diligence. George clarifies his product philosophy, arguing that human-in-the-loop tools amplify human productivity rather than replace jobs.14:04–16:20 · Guest disagreement 1/10 Technical Architecture: Engineering Input, Output, and Reliability Matt asks about Hebbia's underlying technical architecture and defensive moat. George provides commentary on OpenAI's function calling API released that same day, explaining how Hebbia focuses R&D on input, output, and workflow flexibility to avoid model obsolescence.16:20–18:27 · Guest disagreement 0/10 Key Platform Features: Question Lists, Document Comparison, and Semantic Alerts Matt demonstrates prep work by citing specific platform capabilities including question lists, table extraction, and semantic alerts. George elaborates on how these features serve specific workflow needs like CEO-level contract tracking.18:27–21:03 · Guest disagreement 2/10 Addressing AI Bottlenecks: Multi-Step Reasoning and Compound Errors Matt inquires about AI limitations and technical bottlenecks. George breaks down why simple prompts and naive orchestration frameworks fail, illustrating how multi-step LLM operations suffer from compounding errors across nth-order tasks.21:03–25:36 · Guest disagreement 1/10 Hebbia's Long-Term Roadmap and Vision Beyond Finance Matt asks about Hebbia's long-term roadmap before opening to audience questions. An audience member asks about building user trust in multi-step AI outputs, prompting George to explain why raw cosine similarity scores confused users and required hybrid retrieval UX.25:36–29:48 · Guest disagreement 3/10 Audience Q&A: UI/UX Innovation and Proactive AI vs Human-in-the-Loop In response to an audience member asking if AI will proactively reveal unprompted insights, George rejects the premise, asserting that unprompted outlier detection is unhelpful without human-defined intent.29:48–31:24 · Guest disagreement 2/10 Audience Q&A: Differentiating Hebbia from Enterprise Search Solutions Like Glean An audience member asks how Hebbia differentiates from enterprise search tools like Glean. George offers respect for Glean's search focus while clarifying that Hebbia is built as a general-purpose analytical productivity tool for research.0:11–5:18 · Matt pushing back 0/10 George Sivulka's Academic Background and the Genesis of Hebbia Matt opens with a warm, well-researched introduction citing Hebbia's funding history and notable investors before asking George about his journey. George shares a narrative about observing Stanford students at investment banks doing tedious labor and building early prototypes using DEFM-FORTEENA filings.5:18–9:33 · Matt pushing back 1/10 Defining Hebbia as an LLM-Native Productivity Tool and the Excel Analogy Matt asks pragmatic questions about product delivery, model selection, and enterprise customization. George re-frames Hebbia's mission around the 1985 Excel analogy, explaining how LLM-native productivity tools allow non-technical users to programmatically wrangle models.9:33–14:04 · Matt pushing back 3/10 Target Markets: Private Equity Due Diligence and Human-in-the-Loop AI Matt probes whether Hebbia acts purely as a copilot or if it actively replaces human analysts in private equity due diligence. George clarifies his product philosophy, arguing that human-in-the-loop tools amplify human productivity rather than replace jobs.14:04–16:20 · Matt pushing back 1/10 Technical Architecture: Engineering Input, Output, and Reliability Matt asks about Hebbia's underlying technical architecture and defensive moat. George provides commentary on OpenAI's function calling API released that same day, explaining how Hebbia focuses R&D on input, output, and workflow flexibility to avoid model obsolescence.16:20–18:27 · Matt pushing back 0/10 Key Platform Features: Question Lists, Document Comparison, and Semantic Alerts Matt demonstrates prep work by citing specific platform capabilities including question lists, table extraction, and semantic alerts. George elaborates on how these features serve specific workflow needs like CEO-level contract tracking.18:27–21:03 · Matt pushing back 1/10 Addressing AI Bottlenecks: Multi-Step Reasoning and Compound Errors Matt inquires about AI limitations and technical bottlenecks. George breaks down why simple prompts and naive orchestration frameworks fail, illustrating how multi-step LLM operations suffer from compounding errors across nth-order tasks.21:03–25:36 · Matt pushing back 0/10 Hebbia's Long-Term Roadmap and Vision Beyond Finance Matt asks about Hebbia's long-term roadmap before opening to audience questions. An audience member asks about building user trust in multi-step AI outputs, prompting George to explain why raw cosine similarity scores confused users and required hybrid retrieval UX.25:36–29:48 · Matt pushing back 0/10 Audience Q&A: UI/UX Innovation and Proactive AI vs Human-in-the-Loop In response to an audience member asking if AI will proactively reveal unprompted insights, George rejects the premise, asserting that unprompted outlier detection is unhelpful without human-defined intent.29:48–31:24 · Matt pushing back 0/10 Audience Q&A: Differentiating Hebbia from Enterprise Search Solutions Like Glean An audience member asks how Hebbia differentiates from enterprise search tools like Glean. George offers respect for Glean's search focus while clarifying that Hebbia is built as a general-purpose analytical productivity tool for research.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 30.5% · guest 69.5%0:00 · Matt 30.5% · guest 69.5%3:00 · Matt 4.2% · guest 95.8%3:00 · Matt 4.2% · guest 95.8%6:00 · Matt 18.4% · guest 81.6%6:00 · Matt 18.4% · guest 81.6%9:00 · Matt 12.9% · guest 87.1%9:00 · Matt 12.9% · guest 87.1%12:00 · Matt 21.3% · guest 78.7%12:00 · Matt 21.3% · guest 78.7%15:00 · Matt 13.6% · guest 86.4%15:00 · Matt 13.6% · guest 86.4%18:00 · Matt 14.7% · guest 85.3%18:00 · Matt 14.7% · guest 85.3%21:00 · Matt 8.2% · guest 91.8%21:00 · Matt 8.2% · guest 91.8%24:00 · Matt 0% · guest 100%24:00 · Matt 0% · guest 100%27:00 · Matt 6.7% · guest 93.3%27:00 · Matt 6.7% · guest 93.3%30:00 · Matt 0% · guest 100%30:00 · Matt 0% · guest 100%
Sharpest disagreement ▶ 28:40 Rejection of Autonomous Proactive AI Premise

George explicitly pushes back against the audience member's suggestion that AI should autonomously surface unprompted insights, insisting that AI requires a human-defined principal component of meaning.

Hardest push from Matt ▶ 12:04 Host Probing Job Replacement vs. Copilot

Matt presses George on whether Hebbia is merely aiding analysts as a copilot or actually enabling enterprise customers to eliminate analyst headcount.

Biggest teaching moment ▶ 19:25 Math Breakdown of Multi-Step LLM Compound Errors

George educates the audience on LLM reliability math, demonstrating how chaining steps with 90% accuracy results in compound error rates that destroy utility on complex tasks.

Matt holds his own ▶ 16:20 Host Demonstrating Specific Product Feature Knowledge

Matt showcases thorough prep by specifically naming technical platform capabilities including question lists, table extraction, and semantic alerts.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
George Sivulka's Academic Background and the Genesis of Hebbia 2210 Matt opens with a warm, well-researched introduction citing Hebbia's funding history and notable investors before asking George about his journey. George shares a narrative about observing Stanford students at investment banks doing tedious labor and building early prototypes using DEFM-FORTEENA filings.
Defining Hebbia as an LLM-Native Productivity Tool and the Excel Analogy 3311 Matt asks pragmatic questions about product delivery, model selection, and enterprise customization. George re-frames Hebbia's mission around the 1985 Excel analogy, explaining how LLM-native productivity tools allow non-technical users to programmatically wrangle models.
Target Markets: Private Equity Due Diligence and Human-in-the-Loop AI 3223 Matt probes whether Hebbia acts purely as a copilot or if it actively replaces human analysts in private equity due diligence. George clarifies his product philosophy, arguing that human-in-the-loop tools amplify human productivity rather than replace jobs.
Technical Architecture: Engineering Input, Output, and Reliability 2411 Matt asks about Hebbia's underlying technical architecture and defensive moat. George provides commentary on OpenAI's function calling API released that same day, explaining how Hebbia focuses R&D on input, output, and workflow flexibility to avoid model obsolescence.
Key Platform Features: Question Lists, Document Comparison, and Semantic Alerts 4200 Matt demonstrates prep work by citing specific platform capabilities including question lists, table extraction, and semantic alerts. George elaborates on how these features serve specific workflow needs like CEO-level contract tracking.
Addressing AI Bottlenecks: Multi-Step Reasoning and Compound Errors 3521 Matt inquires about AI limitations and technical bottlenecks. George breaks down why simple prompts and naive orchestration frameworks fail, illustrating how multi-step LLM operations suffer from compounding errors across nth-order tasks.
Hebbia's Long-Term Roadmap and Vision Beyond Finance 2410 Matt asks about Hebbia's long-term roadmap before opening to audience questions. An audience member asks about building user trust in multi-step AI outputs, prompting George to explain why raw cosine similarity scores confused users and required hybrid retrieval UX.
Audience Q&A: UI/UX Innovation and Proactive AI vs Human-in-the-Loop 0430 In response to an audience member asking if AI will proactively reveal unprompted insights, George rejects the premise, asserting that unprompted outlier detection is unhelpful without human-defined intent.
Audience Q&A: Differentiating Hebbia from Enterprise Search Solutions Like Glean 1320 An audience member asks how Hebbia differentiates from enterprise search tools like Glean. George offers respect for Glean's search focus while clarifying that Hebbia is built as a general-purpose analytical productivity tool for research.

Statements from this episode (14)

Opinion
Sivulka: Microsoft Excel is the most important software ever created
“To me, it's actually the most important software to ever have been made.”
George Sivulka Jun 21, 2023 ▶ 7:33
Assertion Not checkable as stated
Hebbia serves major banks and U.S. government without fine-tuning
“We have not, so we work with many of the largest asset managers, we work with many banks, we work with the US government and we have not needed to fine tune or, you know, kind of personalize any part of the software.”
George Sivulka Jun 21, 2023 ▶ 9:07
Assertion Not checkable as stated
Sivulka: Not a single job has been replaced by Hebbia
“To date, not a single job has been replaced by Hebbia, you know, knock on wood.”
George Sivulka Jun 21, 2023 ▶ 13:27
Prediction Not checkable as stated
Sivulka predicts new jobs will be created to operate Hebbia
“I think one day there will be many more jobs that are created to run Hebbia.”
George Sivulka Jun 21, 2023 ▶ 13:36
Insight
Hebbia CEO: Users do not naturally trust large language model outputs
“People just don't by default trust the output of a large language model.”
George Sivulka Jun 21, 2023 ▶ 15:54
Assertion Not checkable as stated
Sivulka: Hebbia solved semantic comparison between complex documents
“We allow you to compare two documents semantically, which is a pretty hard problem, ah, that we're pretty excited to have solved.”
George Sivulka Jun 21, 2023 ▶ 16:47
Assertion Not checkable as stated
Sivulka: LLMs can successfully map between different taxonomies
“If it's about mapping between one taxonomy and another taxonomy, you know, you can have a large language model do that very successfully.”
George Sivulka Jun 21, 2023 ▶ 17:23
Insight
Sivulka: Multi-step AI workflows suffer from compounding accuracy errors
“If you have a 90% good system, and then another 90% good system, and then another 90% good system, and then, you know, your output is the sum, or rather the product of all those 90% good systems, you're gonna have something that is, you know, 10% good, dependi…”
George Sivulka Jun 21, 2023 ▶ 19:55
Insight
Sivulka: LLMs excel at first-order tasks but struggle with nth-order
“What we're finding is that these models are good at first order tasks. Right? You can get 90% good at a first order task. They don't quite have the context window, which will be solved, the contextualization kind of the repeated back and forth that allows them…”
George Sivulka Jun 21, 2023 ▶ 20:13
Disclosure
George Sivulka: Remaining solely in financial services means Hebbia failed
“If we stay in financial services, I will have failed.”
George Sivulka Jun 21, 2023 ▶ 21:17
Disclosure
Hebbia removed cosine similarity scores after finding they confused users
“When we used to just do, you know, something like a chroma index or whatever we would include the cosine similarity score on the thing. And people would freak out. They'd be like, what does 72% similar mean? Or, you know, what does very relevant mean? And we a…”
George Sivulka Jun 21, 2023 ▶ 24:27
Insight
Sivulka: Pairing keyword search with LLM search builds user trust
“And we noticed that it was actually really important to use systems of old Like traditional BM-F, TF-IDF kind of keyword matching algorithms to show them results alongside and make these users believe they were still seeing everything.”
George Sivulka Jun 21, 2023 ▶ 24:59
Insight
Sivulka: Human-in-the-loop guidance is essential for useful outlier detection
“Like, if I just took a data set and I said, here are all your outliers, it's actually less helpful than saying, hey, across this principal component, I want to know what's an outlier. And so kind of back to the whole human in the loop paradigm, I think it's ve…”
George Sivulka Jun 21, 2023 ▶ 28:55
Disclosure
Sivulka: Hebbia is internally procuring Glean for corporate search
“I actually think we're internally procuring Glean at Heavya”
George Sivulka Jun 21, 2023 ▶ 30:08
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.