Jun 18, 2026 · 44m · sourcery

Harvey Co-Founder Gabe Pereyra on the Token Pricing Reckoning Coming for AI · Sourcery with Molly O'Shea

Gabe Pereyra · 26m spoken Molly O'Shea · 7m spoken Niko Grupen · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Sourcery, host Molly O'Shea interviews Harvey leaders Gabe Pereyra and Niko Grupen to explore the open-sourcing of Legal Agent Bench, the operational dynamics of agentic workflows, and the impending token pricing reckoning facing enterprise AI adoption.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Molly holds 18.3% of the talking time here. How this is scored →

Molly as informed peer 4.6 Guest teaching 5.4 Guest disagreement 1.2 Molly pushing back 0.7
05100:0015:0030:000:50–5:10 · Molly as informed peer 4/10 Welcome and Setting the Stage with Gabe Pereyra Molly opens by asking Gabe to explain what the newly open-sourced Legal Agent Benchmark measures. Gabe delivers an extensive breakdown comparing legal agent evaluation to SWE-bench and coding unit tests across complex legal diligence workflows.5:10–9:21 · Molly as informed peer 6/10 Model Competition, Diversity, and Cost Saturation Molly directly challenges Gabe on why Harvey would open source benchmarks when their key research partners are also their primary competitors. Gabe reframes the dynamic, arguing law firms face ethical and platform conflict risks that necessitate multi-model neutrality.9:21–12:34 · Molly as informed peer 5/10 Data Privacy, Shared Workspaces, and Fine-Tuning Models Molly asks how Harvey trains models given strict client confidentiality and inquires about Gabe's prior experience at Google Brain and DeepMind. Gabe details Harvey's Shared Spaces architecture and contrasts Brain's bottoms-up research culture with DeepMind's top-down AGI tech tree.12:34–15:15 · Molly as informed peer 5/10 The Evolution of Gabe Pereyra's Role at Harvey Molly references Andrej Karpathy to ask how Gabe's technical role has shifted from startup inception to scale. Gabe describes building two companies in parallel: a traditional enterprise SaaS business followed by an agentic consumption model.15:15–18:10 · Molly as informed peer 5/10 The Booming AI Inference Layer and Custom Model Serving Molly prompts Gabe on the surge of activity across the AI inference layer. Gabe explains how test-time compute, reasoning benchmarks, and specialized inference providers like Base10 and Fireworks enable vertical AI applications to route away from expensive closed models.18:10–21:16 · Molly as informed peer 5/10 Sponsor Segment: Turing AI Infrastructure and Data Systems Molly frames the transition from chat copilots to agentic systems requiring massive compute and memory. Gabe details the economic reality of agent execution, highlighting single legal review runs costing between $20 and $20,000.21:16–23:48 · Molly as informed peer 6/10 Managing Token Economics, Routing, and Model Optimization Molly brings up Harvey's reported 13 trillion token consumption and questions Gabe on internal margin management. Gabe explains being one of the largest consumers of embeddings and outlines their strategy for routing and fine-tuning vertical open-source models.23:48–27:03 · Molly as informed peer 4/10 The Pricing Reckoning: Token Consumption vs. Billable Hours Gabe delivers a detailed breakdown rejecting the simple consensus that consumption pricing solves AI economics. He draws a direct analogy between the legal billable hour audit system and the impending enterprise backlash against opaque multimillion-dollar token bills.27:03–30:06 · Molly as informed peer 4/10 Market Competition and Price-Performance Equilibrium Molly acknowledges Gabe's novel perspective on token monetization, prompting him on how market dynamics will equilibrate pricing and how he consumes research. Gabe discusses model price-performance ratios and the decline of open publishing among frontier labs.30:06–35:46 · Molly as informed peer 2/10 Sponsor Showcase: VCX, Public, Merge, and Deel Following sponsor reads, Molly asks Gabe about mentors and key hiring traits. Gabe reflects on lessons learned from Winston, Barret Zoff, and Jensen Huang, emphasizing topic obsession during candidate evaluation.35:46–38:32 · Molly as informed peer 5/10 Interview with Niko Grupen: Legal Agent Bench Findings Molly introduces Nico Grupen to discuss Legal Agent Bench (LAB) results and data synthesis methodology. Nico explains how Harvey's internal applied legal research team mapped 24 practice areas using agentic synthetic data generation with lawyer-in-the-loop validation.38:32–40:51 · Molly as informed peer 4/10 Evaluating Benchmarks: Quality, Speed, and Cost Trade-offs Molly asks how practitioners should interpret benchmark results. Nico reframes benchmark evaluation from raw quality-maxing to Pareto-efficient trade-offs involving latency, unit costs, and specific legal sub-disciplines.40:51–44:27 · Molly as informed peer 5/10 Agent Harnesses and Organizational Intelligence in Legal AI Molly questions Nico on future roadmap items and agent scaffolding. Nico defines agent harnesses and outlines how legal AI is transitioning from single-agent task execution to organizational-level intelligence and multi-human collaboration.0:50–5:10 · Guest teaching 6/10 Welcome and Setting the Stage with Gabe Pereyra Molly opens by asking Gabe to explain what the newly open-sourced Legal Agent Benchmark measures. Gabe delivers an extensive breakdown comparing legal agent evaluation to SWE-bench and coding unit tests across complex legal diligence workflows.5:10–9:21 · Guest teaching 6/10 Model Competition, Diversity, and Cost Saturation Molly directly challenges Gabe on why Harvey would open source benchmarks when their key research partners are also their primary competitors. Gabe reframes the dynamic, arguing law firms face ethical and platform conflict risks that necessitate multi-model neutrality.9:21–12:34 · Guest teaching 5/10 Data Privacy, Shared Workspaces, and Fine-Tuning Models Molly asks how Harvey trains models given strict client confidentiality and inquires about Gabe's prior experience at Google Brain and DeepMind. Gabe details Harvey's Shared Spaces architecture and contrasts Brain's bottoms-up research culture with DeepMind's top-down AGI tech tree.12:34–15:15 · Guest teaching 4/10 The Evolution of Gabe Pereyra's Role at Harvey Molly references Andrej Karpathy to ask how Gabe's technical role has shifted from startup inception to scale. Gabe describes building two companies in parallel: a traditional enterprise SaaS business followed by an agentic consumption model.15:15–18:10 · Guest teaching 5/10 The Booming AI Inference Layer and Custom Model Serving Molly prompts Gabe on the surge of activity across the AI inference layer. Gabe explains how test-time compute, reasoning benchmarks, and specialized inference providers like Base10 and Fireworks enable vertical AI applications to route away from expensive closed models.18:10–21:16 · Guest teaching 6/10 Sponsor Segment: Turing AI Infrastructure and Data Systems Molly frames the transition from chat copilots to agentic systems requiring massive compute and memory. Gabe details the economic reality of agent execution, highlighting single legal review runs costing between $20 and $20,000.21:16–23:48 · Guest teaching 5/10 Managing Token Economics, Routing, and Model Optimization Molly brings up Harvey's reported 13 trillion token consumption and questions Gabe on internal margin management. Gabe explains being one of the largest consumers of embeddings and outlines their strategy for routing and fine-tuning vertical open-source models.23:48–27:03 · Guest teaching 8/10 The Pricing Reckoning: Token Consumption vs. Billable Hours Gabe delivers a detailed breakdown rejecting the simple consensus that consumption pricing solves AI economics. He draws a direct analogy between the legal billable hour audit system and the impending enterprise backlash against opaque multimillion-dollar token bills.27:03–30:06 · Guest teaching 5/10 Market Competition and Price-Performance Equilibrium Molly acknowledges Gabe's novel perspective on token monetization, prompting him on how market dynamics will equilibrate pricing and how he consumes research. Gabe discusses model price-performance ratios and the decline of open publishing among frontier labs.30:06–35:46 · Guest teaching 2/10 Sponsor Showcase: VCX, Public, Merge, and Deel Following sponsor reads, Molly asks Gabe about mentors and key hiring traits. Gabe reflects on lessons learned from Winston, Barret Zoff, and Jensen Huang, emphasizing topic obsession during candidate evaluation.35:46–38:32 · Guest teaching 6/10 Interview with Niko Grupen: Legal Agent Bench Findings Molly introduces Nico Grupen to discuss Legal Agent Bench (LAB) results and data synthesis methodology. Nico explains how Harvey's internal applied legal research team mapped 24 practice areas using agentic synthetic data generation with lawyer-in-the-loop validation.38:32–40:51 · Guest teaching 6/10 Evaluating Benchmarks: Quality, Speed, and Cost Trade-offs Molly asks how practitioners should interpret benchmark results. Nico reframes benchmark evaluation from raw quality-maxing to Pareto-efficient trade-offs involving latency, unit costs, and specific legal sub-disciplines.40:51–44:27 · Guest teaching 6/10 Agent Harnesses and Organizational Intelligence in Legal AI Molly questions Nico on future roadmap items and agent scaffolding. Nico defines agent harnesses and outlines how legal AI is transitioning from single-agent task execution to organizational-level intelligence and multi-human collaboration.0:50–5:10 · Guest disagreement 1/10 Welcome and Setting the Stage with Gabe Pereyra Molly opens by asking Gabe to explain what the newly open-sourced Legal Agent Benchmark measures. Gabe delivers an extensive breakdown comparing legal agent evaluation to SWE-bench and coding unit tests across complex legal diligence workflows.5:10–9:21 · Guest disagreement 3/10 Model Competition, Diversity, and Cost Saturation Molly directly challenges Gabe on why Harvey would open source benchmarks when their key research partners are also their primary competitors. Gabe reframes the dynamic, arguing law firms face ethical and platform conflict risks that necessitate multi-model neutrality.9:21–12:34 · Guest disagreement 1/10 Data Privacy, Shared Workspaces, and Fine-Tuning Models Molly asks how Harvey trains models given strict client confidentiality and inquires about Gabe's prior experience at Google Brain and DeepMind. Gabe details Harvey's Shared Spaces architecture and contrasts Brain's bottoms-up research culture with DeepMind's top-down AGI tech tree.12:34–15:15 · Guest disagreement 1/10 The Evolution of Gabe Pereyra's Role at Harvey Molly references Andrej Karpathy to ask how Gabe's technical role has shifted from startup inception to scale. Gabe describes building two companies in parallel: a traditional enterprise SaaS business followed by an agentic consumption model.15:15–18:10 · Guest disagreement 1/10 The Booming AI Inference Layer and Custom Model Serving Molly prompts Gabe on the surge of activity across the AI inference layer. Gabe explains how test-time compute, reasoning benchmarks, and specialized inference providers like Base10 and Fireworks enable vertical AI applications to route away from expensive closed models.18:10–21:16 · Guest disagreement 1/10 Sponsor Segment: Turing AI Infrastructure and Data Systems Molly frames the transition from chat copilots to agentic systems requiring massive compute and memory. Gabe details the economic reality of agent execution, highlighting single legal review runs costing between $20 and $20,000.21:16–23:48 · Guest disagreement 1/10 Managing Token Economics, Routing, and Model Optimization Molly brings up Harvey's reported 13 trillion token consumption and questions Gabe on internal margin management. Gabe explains being one of the largest consumers of embeddings and outlines their strategy for routing and fine-tuning vertical open-source models.23:48–27:03 · Guest disagreement 4/10 The Pricing Reckoning: Token Consumption vs. Billable Hours Gabe delivers a detailed breakdown rejecting the simple consensus that consumption pricing solves AI economics. He draws a direct analogy between the legal billable hour audit system and the impending enterprise backlash against opaque multimillion-dollar token bills.27:03–30:06 · Guest disagreement 2/10 Market Competition and Price-Performance Equilibrium Molly acknowledges Gabe's novel perspective on token monetization, prompting him on how market dynamics will equilibrate pricing and how he consumes research. Gabe discusses model price-performance ratios and the decline of open publishing among frontier labs.30:06–35:46 · Guest disagreement 0/10 Sponsor Showcase: VCX, Public, Merge, and Deel Following sponsor reads, Molly asks Gabe about mentors and key hiring traits. Gabe reflects on lessons learned from Winston, Barret Zoff, and Jensen Huang, emphasizing topic obsession during candidate evaluation.35:46–38:32 · Guest disagreement 0/10 Interview with Niko Grupen: Legal Agent Bench Findings Molly introduces Nico Grupen to discuss Legal Agent Bench (LAB) results and data synthesis methodology. Nico explains how Harvey's internal applied legal research team mapped 24 practice areas using agentic synthetic data generation with lawyer-in-the-loop validation.38:32–40:51 · Guest disagreement 1/10 Evaluating Benchmarks: Quality, Speed, and Cost Trade-offs Molly asks how practitioners should interpret benchmark results. Nico reframes benchmark evaluation from raw quality-maxing to Pareto-efficient trade-offs involving latency, unit costs, and specific legal sub-disciplines.40:51–44:27 · Guest disagreement 0/10 Agent Harnesses and Organizational Intelligence in Legal AI Molly questions Nico on future roadmap items and agent scaffolding. Nico defines agent harnesses and outlines how legal AI is transitioning from single-agent task execution to organizational-level intelligence and multi-human collaboration.0:50–5:10 · Molly pushing back 0/10 Welcome and Setting the Stage with Gabe Pereyra Molly opens by asking Gabe to explain what the newly open-sourced Legal Agent Benchmark measures. Gabe delivers an extensive breakdown comparing legal agent evaluation to SWE-bench and coding unit tests across complex legal diligence workflows.5:10–9:21 · Molly pushing back 5/10 Model Competition, Diversity, and Cost Saturation Molly directly challenges Gabe on why Harvey would open source benchmarks when their key research partners are also their primary competitors. Gabe reframes the dynamic, arguing law firms face ethical and platform conflict risks that necessitate multi-model neutrality.9:21–12:34 · Molly pushing back 1/10 Data Privacy, Shared Workspaces, and Fine-Tuning Models Molly asks how Harvey trains models given strict client confidentiality and inquires about Gabe's prior experience at Google Brain and DeepMind. Gabe details Harvey's Shared Spaces architecture and contrasts Brain's bottoms-up research culture with DeepMind's top-down AGI tech tree.12:34–15:15 · Molly pushing back 0/10 The Evolution of Gabe Pereyra's Role at Harvey Molly references Andrej Karpathy to ask how Gabe's technical role has shifted from startup inception to scale. Gabe describes building two companies in parallel: a traditional enterprise SaaS business followed by an agentic consumption model.15:15–18:10 · Molly pushing back 0/10 The Booming AI Inference Layer and Custom Model Serving Molly prompts Gabe on the surge of activity across the AI inference layer. Gabe explains how test-time compute, reasoning benchmarks, and specialized inference providers like Base10 and Fireworks enable vertical AI applications to route away from expensive closed models.18:10–21:16 · Molly pushing back 0/10 Sponsor Segment: Turing AI Infrastructure and Data Systems Molly frames the transition from chat copilots to agentic systems requiring massive compute and memory. Gabe details the economic reality of agent execution, highlighting single legal review runs costing between $20 and $20,000.21:16–23:48 · Molly pushing back 0/10 Managing Token Economics, Routing, and Model Optimization Molly brings up Harvey's reported 13 trillion token consumption and questions Gabe on internal margin management. Gabe explains being one of the largest consumers of embeddings and outlines their strategy for routing and fine-tuning vertical open-source models.23:48–27:03 · Molly pushing back 1/10 The Pricing Reckoning: Token Consumption vs. Billable Hours Gabe delivers a detailed breakdown rejecting the simple consensus that consumption pricing solves AI economics. He draws a direct analogy between the legal billable hour audit system and the impending enterprise backlash against opaque multimillion-dollar token bills.27:03–30:06 · Molly pushing back 1/10 Market Competition and Price-Performance Equilibrium Molly acknowledges Gabe's novel perspective on token monetization, prompting him on how market dynamics will equilibrate pricing and how he consumes research. Gabe discusses model price-performance ratios and the decline of open publishing among frontier labs.30:06–35:46 · Molly pushing back 0/10 Sponsor Showcase: VCX, Public, Merge, and Deel Following sponsor reads, Molly asks Gabe about mentors and key hiring traits. Gabe reflects on lessons learned from Winston, Barret Zoff, and Jensen Huang, emphasizing topic obsession during candidate evaluation.35:46–38:32 · Molly pushing back 0/10 Interview with Niko Grupen: Legal Agent Bench Findings Molly introduces Nico Grupen to discuss Legal Agent Bench (LAB) results and data synthesis methodology. Nico explains how Harvey's internal applied legal research team mapped 24 practice areas using agentic synthetic data generation with lawyer-in-the-loop validation.38:32–40:51 · Molly pushing back 1/10 Evaluating Benchmarks: Quality, Speed, and Cost Trade-offs Molly asks how practitioners should interpret benchmark results. Nico reframes benchmark evaluation from raw quality-maxing to Pareto-efficient trade-offs involving latency, unit costs, and specific legal sub-disciplines.40:51–44:27 · Molly pushing back 0/10 Agent Harnesses and Organizational Intelligence in Legal AI Molly questions Nico on future roadmap items and agent scaffolding. Nico defines agent harnesses and outlines how legal AI is transitioning from single-agent task execution to organizational-level intelligence and multi-human collaboration.

speaking balance: gold is Molly, purple is the guest (3 minute bins)

0:00 · Molly 14.5% · guest 85.5%0:00 · Molly 14.5% · guest 85.5%3:00 · Molly 8.5% · guest 91.5%3:00 · Molly 8.5% · guest 91.5%6:00 · Molly 10.1% · guest 89.9%6:00 · Molly 10.1% · guest 89.9%9:00 · Molly 9.3% · guest 90.7%9:00 · Molly 9.3% · guest 90.7%12:00 · Molly 14.6% · guest 85.4%12:00 · Molly 14.6% · guest 85.4%15:00 · Molly 32.2% · guest 67.8%15:00 · Molly 32.2% · guest 67.8%18:00 · Molly 38.6% · guest 61.4%18:00 · Molly 38.6% · guest 61.4%21:00 · Molly 16.4% · guest 83.6%21:00 · Molly 16.4% · guest 83.6%24:00 · Molly 0% · guest 100%24:00 · Molly 0% · guest 100%27:00 · Molly 11.2% · guest 88.8%27:00 · Molly 11.2% · guest 88.8%30:00 · Molly 70.5% · guest 29.5%30:00 · Molly 70.5% · guest 29.5%33:00 · Molly 6.7% · guest 93.3%33:00 · Molly 6.7% · guest 93.3%36:00 · Molly 9.6% · guest 90.4%36:00 · Molly 9.6% · guest 90.4%39:00 · Molly 7.2% · guest 92.8%39:00 · Molly 7.2% · guest 92.8%42:00 · Molly 25.3% · guest 74.7%42:00 · Molly 25.3% · guest 74.7%
Sharpest disagreement ▶ 23:51 Gabe rejects conventional VC consensus on consumption pricing

Gabe forcefully pushes back against the prevailing tech assumption that consumption pricing solves SaaS monetization, arguing that massive unpredictable token bills will cause severe customer friction.

Hardest push from Molly ▶ 6:37 Molly challenges Gabe on open-sourcing against partner-competitors

Molly presses Gabe on why Harvey is open-sourcing core evaluation tooling when frontier model providers like OpenAI and Anthropic are also their most formidable potential competitors.

Biggest teaching moment ▶ 24:20 Gabe compares token billing friction to the law firm billable hour

Gabe educates the audience on why the billable hour survived for decades and explains how enterprise token consumption will inevitably require identical granular itemization to justify runaway costs.

Molly holds their own ▶ 21:16 Molly confronts Gabe with Harvey's 13 trillion token metric

Molly leverages specific internal operational data shared by Harvey's co-founder to press Gabe on margin pressure and routing efficiencies under extreme scale.

the scores for every segment, with the reasoning behind each
ChapterTopicMolly as informed peerGuest teachingGuest disagreementMolly pushing backWhy
Welcome and Setting the Stage with Gabe Pereyra 4610 Molly opens by asking Gabe to explain what the newly open-sourced Legal Agent Benchmark measures. Gabe delivers an extensive breakdown comparing legal agent evaluation to SWE-bench and coding unit tests across complex legal diligence workflows.
Model Competition, Diversity, and Cost Saturation 6635 Molly directly challenges Gabe on why Harvey would open source benchmarks when their key research partners are also their primary competitors. Gabe reframes the dynamic, arguing law firms face ethical and platform conflict risks that necessitate multi-model neutrality.
Data Privacy, Shared Workspaces, and Fine-Tuning Models 5511 Molly asks how Harvey trains models given strict client confidentiality and inquires about Gabe's prior experience at Google Brain and DeepMind. Gabe details Harvey's Shared Spaces architecture and contrasts Brain's bottoms-up research culture with DeepMind's top-down AGI tech tree.
The Evolution of Gabe Pereyra's Role at Harvey 5410 Molly references Andrej Karpathy to ask how Gabe's technical role has shifted from startup inception to scale. Gabe describes building two companies in parallel: a traditional enterprise SaaS business followed by an agentic consumption model.
The Booming AI Inference Layer and Custom Model Serving 5510 Molly prompts Gabe on the surge of activity across the AI inference layer. Gabe explains how test-time compute, reasoning benchmarks, and specialized inference providers like Base10 and Fireworks enable vertical AI applications to route away from expensive closed models.
Sponsor Segment: Turing AI Infrastructure and Data Systems 5610 Molly frames the transition from chat copilots to agentic systems requiring massive compute and memory. Gabe details the economic reality of agent execution, highlighting single legal review runs costing between $20 and $20,000.
Managing Token Economics, Routing, and Model Optimization 6510 Molly brings up Harvey's reported 13 trillion token consumption and questions Gabe on internal margin management. Gabe explains being one of the largest consumers of embeddings and outlines their strategy for routing and fine-tuning vertical open-source models.
The Pricing Reckoning: Token Consumption vs. Billable Hours 4841 Gabe delivers a detailed breakdown rejecting the simple consensus that consumption pricing solves AI economics. He draws a direct analogy between the legal billable hour audit system and the impending enterprise backlash against opaque multimillion-dollar token bills.
Market Competition and Price-Performance Equilibrium 4521 Molly acknowledges Gabe's novel perspective on token monetization, prompting him on how market dynamics will equilibrate pricing and how he consumes research. Gabe discusses model price-performance ratios and the decline of open publishing among frontier labs.
Sponsor Showcase: VCX, Public, Merge, and Deel 2200 Following sponsor reads, Molly asks Gabe about mentors and key hiring traits. Gabe reflects on lessons learned from Winston, Barret Zoff, and Jensen Huang, emphasizing topic obsession during candidate evaluation.
Interview with Niko Grupen: Legal Agent Bench Findings 5600 Molly introduces Nico Grupen to discuss Legal Agent Bench (LAB) results and data synthesis methodology. Nico explains how Harvey's internal applied legal research team mapped 24 practice areas using agentic synthetic data generation with lawyer-in-the-loop validation.
Evaluating Benchmarks: Quality, Speed, and Cost Trade-offs 4611 Molly asks how practitioners should interpret benchmark results. Nico reframes benchmark evaluation from raw quality-maxing to Pareto-efficient trade-offs involving latency, unit costs, and specific legal sub-disciplines.
Agent Harnesses and Organizational Intelligence in Legal AI 5600 Molly questions Nico on future roadmap items and agent scaffolding. Nico defines agent harnesses and outlines how legal AI is transitioning from single-agent task execution to organizational-level intelligence and multi-human collaboration.

Statements from this episode (21)

Insight
Pereyra: Big Law work can be objectively quantified and unit-tested
“And I think there's a bit of a misconception that, oh, legal is subjective, so you can't do this. And I think the thing for especially big law is a lot of the work you can actually quantify, and this is what senior associates and partners are doing when their …”
Gabe Pereyra Jun 18, 2026 ▶ 2:35
Disclosure
Harvey is building a taxonomy of Big Law tasks to benchmark agents
“So a lot of what we're building is essentially this taxonomy of here's all the tasks that a law firm would do, and here's all the subtasks that associates would do, and how partners or senior associates would grade that work.”
Gabe Pereyra Jun 18, 2026 ▶ 4:44
Assertion Supported
Pereyra: No Single AI Model Wins Across All Legal Tasks
“What we've seen is actually every different model is good at something different. And so with the initial results we saw, Anthropics models are quite strong, but there's areas where a 5.5 is better. There's some areas where open source is better.”
Gabe Pereyra Jun 18, 2026 ▶ 5:49
Insight
Pereyra: Model Selection Is Shifting to Lowest Cost per Task
“Increasingly, it's not just which model is the best. It's which model can solve the task at the lowest price point, because as these models get better, there's kind of intelligent saturation where it's like, okay, for some simple tasks, I actually want to know…”
Gabe Pereyra Jun 18, 2026 ▶ 6:04
Insight
Pereyra: Client conflict rules mean law firms need multiple AI providers
“Most of these law firms and enterprises Can't rely on a single lab provider. So like there's a big risk if you're a law firm that if you just use Anthropic or you just use OpenAI you run into conflict risk. So imagine you're using only Anthropics models as a l…”
Gabe Pereyra Jun 18, 2026 ▶ 7:13
Disclosure
Harvey plans to open-source general AI and monetize private enterprise infrastructure
“We think of our strategy as we want to open source the general stuff and work with all of the providers to make these models as good as possible at general legal. And then we want to build infrastructure for these law firms and enterprises that help them own t…”
Gabe Pereyra Jun 18, 2026 ▶ 8:59
Assertion Not checkable as stated
Pereyra: Google Brain operated on a bottoms-up, researcher-driven model in 2016-2017
“And so when I was at Brain, it was very bottoms up. It was, I mean, both had some of the smartest AI researchers in the world, but Brain, the approach was, you know, let's get all these smart people, give them a bunch of compute, and then kind of let them do t…”
Gabe Pereyra Jun 18, 2026 ▶ 11:12
Assertion Not checkable as stated
Pereyra: DeepMind used a top-down 'tech tree' approach to reach AGI
“And then DeepMind was much more top down where Demis just had this vision of, okay, we're going to create AGI. Here's all the things that I think are required. And so they kind of had this tech tree of how they were going to do it. And then let's have all thes…”
Gabe Pereyra Jun 18, 2026 ▶ 11:36
Disclosure
Pereyra: Harvey models its vertical AI strategy on DeepMind's top-down approach
“And I think our approach is a bit more inspired by the DeepMind one, just because I think because we're in a vertical, the end goal is very clear.”
Gabe Pereyra Jun 18, 2026 ▶ 11:56
Prediction Not checkable as stated
Pereyra: AI Pricing Will Likely Return to Subscription Models From Tokens
“Buying off tokens is insanely complicated. It'll probably at some point go back to something like subscription because it's hard to like forecast these things.”
Gabe Pereyra Jun 18, 2026 ▶ 14:21
Assertion Not checkable as stated
Pereyra: Harvey is moving traffic to open-source models due to closed-model costs
“There's customers like us, where we do need to move some of our traffic to open source models because it's just becoming super expensive to serve like the largest closed frontier model.”
Gabe Pereyra Jun 18, 2026 ▶ 16:40
Assertion Not checkable as stated
Pereyra: Harvey queries cost up to $20 for drafts, $20k for reviews
“We have simple like assistant type queries where you say, you know, draft me a document that a single query can cost 20 dollars. We have like a review product where you can upload a 100,000 contracts and ask the models to review them, and some of those can cos…”
Gabe Pereyra Jun 18, 2026 ▶ 20:02
Opinion
Pereyra: Legal AI reached an adoption inflection point in late 2023/early 2024
“I think we're just starting to see in the past six months that inflection for legal where the models can now generate entire documents and they're starting to work in a way where Most lawyers who aren't using this technology just for fun, they're like, this ne…”
Gabe Pereyra Jun 18, 2026 ▶ 20:36
Assertion Not checkable as stated
Pereyra: Harvey Is Likely the Largest Embeddings Consumer for Some AI Labs
“I think for some of the labs, we are like the largest consumer of embeddings.”
Gabe Pereyra Jun 18, 2026 ▶ 22:24
Insight
Pereyra: Vertical AI Does Not Need Trillion-Parameter Models for Narrow Tasks
“A lot of these like very large frontier models are large because they're good at everything. And so I think a lot of the opportunity for these like specific verticals will be, okay, I probably don't need a trillion parameters if I just need the system to be go…”
Gabe Pereyra Jun 18, 2026 ▶ 23:21
Insight
Pereyra: The billable hour enables standardized pricing of complex work at scale
“And I think something people don't appreciate about the billable hour and why it's such a good mechanism is, It lets you price incredibly complex work at massive scale in a way that the entire industry can agree on, right?”
Gabe Pereyra Jun 18, 2026 ▶ 25:04
Prediction Not checkable as stated
Pereyra: Token billing will spawn an optimization ecosystem to counter model incentives
“I think you're going to start seeing things like this for token billing, where there's going to become this whole ecosystem of like, how do you optimize around this? Because you kind of have these weird misaligned incentives from the model providers, right? Be…”
Gabe Pereyra Jun 18, 2026 ▶ 26:11
Insight
Pereyra: Inter-firm competition keeps law firm billable hour pricing in check
“Okay, the billable hour creates some pricing misalignment, but the fact that these law firms need to compete with every other law firm means that that's kept in check. And so it's like, this actually converges to roughly the right pricing.”
Gabe Pereyra Jun 18, 2026 ▶ 27:32
Prediction Not checkable as stated
Pereyra: Competition among frontier and open-source models will check AI pricing
“And so I think the same thing will happen with the model providers, where it's like, if you look at Opus 4.7 and 5.5, Opus 4.7 is three times more expensive than 5.5, but it's, you know, 10 or 20% more performant. But like, these are going to cause pressure on…”
Gabe Pereyra Jun 18, 2026 ▶ 27:47
Opinion
Pereyra: AI research culture has completely abandoned open publishing
“That has completely changed in the sense, like, no one really publishes what they're doing anymore, and so it is harder to keep track of, like, the research breakthroughs and what's going on.”
Gabe Pereyra Jun 18, 2026 ▶ 29:15
Disclosure
Pereyra: Harvey will publish research with labs and open-source models
“We are going to start publishing some of our research that we're doing with the labs or the other providers and open sourcing more models, more of the work we're doing.”
Gabe Pereyra Jun 18, 2026 ▶ 29:43
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 160 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.