Oct 24, 2024 · 59m · mad

The Death of Big Data and Why It’s Time To Think Small | Jordan Tigani, CEO, MotherDuck

Jordan Tigani · 48m spoken Matt Turck · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, MotherDuck CEO Jordan Tigani joins Matt Turck to discuss why 'Big Data is Dead,' explaining how small data analytics powered by DuckDB offers superior latency, lower costs, and simpler architecture. Jordan details MotherDuck's hybrid execution model, open-source partnership, and key entrepreneurial lessons on transitioning from engineering to founder leadership.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 12.3% of the talking time here. How this is scored →

Matt as informed peer 3.9 Guest teaching 4.1 Guest disagreement 1.0 Matt pushing back 1.4
05100:0015:0030:0045:000:56–6:55 · Matt as informed peer 3/10 The Death of Big Data and the Rise of Small Data Matt prompts Jordan on his small data post and conference. Jordan contrasts actual user query behavior at BigQuery against industry benchmarks, with Matt adding brief contextual details on MapReduce history.6:55–15:33 · Matt as informed peer 5/10 The Small Data Manifesto and Brand Positioning Matt challenges Jordan on whether existing platforms like Snowflake or BigQuery can efficiently serve small data. Jordan explains the distributed coordination latency tax and the medallion architecture presentation tier.15:33–18:48 · Matt as informed peer 3/10 Understanding DuckDB and Its Academic Origins Matt asks for a background on DuckDB and its academic origins. Jordan explains embedded database architecture, Python integration, and CWI history.18:48–25:24 · Matt as informed peer 4/10 Connecting with DuckDB Labs and Structuring MotherDuck Matt brings up standard VC advice about commercial startups needing to own open source communities. Jordan reframes how MotherDuck maintains an aligned partnership with DuckDB Labs while keeping independent communities.25:24–31:26 · Matt as informed peer 3/10 Funding, Hybrid Execution Architecture, and Office Culture Jordan outlines MotherDuck's funding history and technical architecture. Matt asks conversational follow-up questions regarding team size and office dynamics.31:26–39:09 · Matt as informed peer 4/10 Database Mechanics, Vectorized Execution, and Local Caching Matt asks Jordan to clarify in-memory analytics and what makes DuckDB fast. Jordan delivers an technical overview of Stonebraker's paper, compiler optimizations for vectorized execution, and client caching.39:09–43:03 · Matt as informed peer 5/10 Local AI Models, On-Device Inference, and Local RAG Matt asks about current system limitations, local AI models, and whether small data hurts data ingestion vendors like Fivetran. Jordan details local RAG architectures and reframe Fivetran's core value.43:03–50:04 · Matt as informed peer 6/10 Ecosystem Integration, Open Formats, and Reducing Friction Matt cites recent industry moves like Databricks acquiring Tabular and asks if stack complexity is truly decreasing. Jordan explains how open formats like Apache Iceberg reduce data lock-in.50:22–53:58 · Matt as informed peer 3/10 Mindset and Entrepreneurial Lessons for Technical Founders Matt asks Jordan to reflect on transitioning into a first-time founder role. Jordan discusses relying on intuition and managing contradictory fundraising advice from experienced founders.53:58–58:39 · Matt as informed peer 3/10 Transitioning from Engineering to Product, Marketing, and Business Leadership Matt asks how Jordan developed commercial and marketing skills coming from engineering. Jordan shares lessons learned from seeing built technology die without product and marketing alignment.0:56–6:55 · Guest teaching 5/10 The Death of Big Data and the Rise of Small Data Matt prompts Jordan on his small data post and conference. Jordan contrasts actual user query behavior at BigQuery against industry benchmarks, with Matt adding brief contextual details on MapReduce history.6:55–15:33 · Guest teaching 5/10 The Small Data Manifesto and Brand Positioning Matt challenges Jordan on whether existing platforms like Snowflake or BigQuery can efficiently serve small data. Jordan explains the distributed coordination latency tax and the medallion architecture presentation tier.15:33–18:48 · Guest teaching 4/10 Understanding DuckDB and Its Academic Origins Matt asks for a background on DuckDB and its academic origins. Jordan explains embedded database architecture, Python integration, and CWI history.18:48–25:24 · Guest teaching 3/10 Connecting with DuckDB Labs and Structuring MotherDuck Matt brings up standard VC advice about commercial startups needing to own open source communities. Jordan reframes how MotherDuck maintains an aligned partnership with DuckDB Labs while keeping independent communities.25:24–31:26 · Guest teaching 4/10 Funding, Hybrid Execution Architecture, and Office Culture Jordan outlines MotherDuck's funding history and technical architecture. Matt asks conversational follow-up questions regarding team size and office dynamics.31:26–39:09 · Guest teaching 6/10 Database Mechanics, Vectorized Execution, and Local Caching Matt asks Jordan to clarify in-memory analytics and what makes DuckDB fast. Jordan delivers an technical overview of Stonebraker's paper, compiler optimizations for vectorized execution, and client caching.39:09–43:03 · Guest teaching 4/10 Local AI Models, On-Device Inference, and Local RAG Matt asks about current system limitations, local AI models, and whether small data hurts data ingestion vendors like Fivetran. Jordan details local RAG architectures and reframe Fivetran's core value.43:03–50:04 · Guest teaching 5/10 Ecosystem Integration, Open Formats, and Reducing Friction Matt cites recent industry moves like Databricks acquiring Tabular and asks if stack complexity is truly decreasing. Jordan explains how open formats like Apache Iceberg reduce data lock-in.50:22–53:58 · Guest teaching 3/10 Mindset and Entrepreneurial Lessons for Technical Founders Matt asks Jordan to reflect on transitioning into a first-time founder role. Jordan discusses relying on intuition and managing contradictory fundraising advice from experienced founders.53:58–58:39 · Guest teaching 2/10 Transitioning from Engineering to Product, Marketing, and Business Leadership Matt asks how Jordan developed commercial and marketing skills coming from engineering. Jordan shares lessons learned from seeing built technology die without product and marketing alignment.0:56–6:55 · Guest disagreement 1/10 The Death of Big Data and the Rise of Small Data Matt prompts Jordan on his small data post and conference. Jordan contrasts actual user query behavior at BigQuery against industry benchmarks, with Matt adding brief contextual details on MapReduce history.6:55–15:33 · Guest disagreement 2/10 The Small Data Manifesto and Brand Positioning Matt challenges Jordan on whether existing platforms like Snowflake or BigQuery can efficiently serve small data. Jordan explains the distributed coordination latency tax and the medallion architecture presentation tier.15:33–18:48 · Guest disagreement 0/10 Understanding DuckDB and Its Academic Origins Matt asks for a background on DuckDB and its academic origins. Jordan explains embedded database architecture, Python integration, and CWI history.18:48–25:24 · Guest disagreement 1/10 Connecting with DuckDB Labs and Structuring MotherDuck Matt brings up standard VC advice about commercial startups needing to own open source communities. Jordan reframes how MotherDuck maintains an aligned partnership with DuckDB Labs while keeping independent communities.25:24–31:26 · Guest disagreement 1/10 Funding, Hybrid Execution Architecture, and Office Culture Jordan outlines MotherDuck's funding history and technical architecture. Matt asks conversational follow-up questions regarding team size and office dynamics.31:26–39:09 · Guest disagreement 1/10 Database Mechanics, Vectorized Execution, and Local Caching Matt asks Jordan to clarify in-memory analytics and what makes DuckDB fast. Jordan delivers an technical overview of Stonebraker's paper, compiler optimizations for vectorized execution, and client caching.39:09–43:03 · Guest disagreement 2/10 Local AI Models, On-Device Inference, and Local RAG Matt asks about current system limitations, local AI models, and whether small data hurts data ingestion vendors like Fivetran. Jordan details local RAG architectures and reframe Fivetran's core value.43:03–50:04 · Guest disagreement 2/10 Ecosystem Integration, Open Formats, and Reducing Friction Matt cites recent industry moves like Databricks acquiring Tabular and asks if stack complexity is truly decreasing. Jordan explains how open formats like Apache Iceberg reduce data lock-in.50:22–53:58 · Guest disagreement 0/10 Mindset and Entrepreneurial Lessons for Technical Founders Matt asks Jordan to reflect on transitioning into a first-time founder role. Jordan discusses relying on intuition and managing contradictory fundraising advice from experienced founders.53:58–58:39 · Guest disagreement 0/10 Transitioning from Engineering to Product, Marketing, and Business Leadership Matt asks how Jordan developed commercial and marketing skills coming from engineering. Jordan shares lessons learned from seeing built technology die without product and marketing alignment.0:56–6:55 · Matt pushing back 1/10 The Death of Big Data and the Rise of Small Data Matt prompts Jordan on his small data post and conference. Jordan contrasts actual user query behavior at BigQuery against industry benchmarks, with Matt adding brief contextual details on MapReduce history.6:55–15:33 · Matt pushing back 4/10 The Small Data Manifesto and Brand Positioning Matt challenges Jordan on whether existing platforms like Snowflake or BigQuery can efficiently serve small data. Jordan explains the distributed coordination latency tax and the medallion architecture presentation tier.15:33–18:48 · Matt pushing back 0/10 Understanding DuckDB and Its Academic Origins Matt asks for a background on DuckDB and its academic origins. Jordan explains embedded database architecture, Python integration, and CWI history.18:48–25:24 · Matt pushing back 2/10 Connecting with DuckDB Labs and Structuring MotherDuck Matt brings up standard VC advice about commercial startups needing to own open source communities. Jordan reframes how MotherDuck maintains an aligned partnership with DuckDB Labs while keeping independent communities.25:24–31:26 · Matt pushing back 0/10 Funding, Hybrid Execution Architecture, and Office Culture Jordan outlines MotherDuck's funding history and technical architecture. Matt asks conversational follow-up questions regarding team size and office dynamics.31:26–39:09 · Matt pushing back 1/10 Database Mechanics, Vectorized Execution, and Local Caching Matt asks Jordan to clarify in-memory analytics and what makes DuckDB fast. Jordan delivers an technical overview of Stonebraker's paper, compiler optimizations for vectorized execution, and client caching.39:09–43:03 · Matt pushing back 3/10 Local AI Models, On-Device Inference, and Local RAG Matt asks about current system limitations, local AI models, and whether small data hurts data ingestion vendors like Fivetran. Jordan details local RAG architectures and reframe Fivetran's core value.43:03–50:04 · Matt pushing back 3/10 Ecosystem Integration, Open Formats, and Reducing Friction Matt cites recent industry moves like Databricks acquiring Tabular and asks if stack complexity is truly decreasing. Jordan explains how open formats like Apache Iceberg reduce data lock-in.50:22–53:58 · Matt pushing back 0/10 Mindset and Entrepreneurial Lessons for Technical Founders Matt asks Jordan to reflect on transitioning into a first-time founder role. Jordan discusses relying on intuition and managing contradictory fundraising advice from experienced founders.53:58–58:39 · Matt pushing back 0/10 Transitioning from Engineering to Product, Marketing, and Business Leadership Matt asks how Jordan developed commercial and marketing skills coming from engineering. Jordan shares lessons learned from seeing built technology die without product and marketing alignment.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 37.1% · guest 62.9%0:00 · Matt 37.1% · guest 62.9%3:00 · Matt 2% · guest 98%3:00 · Matt 2% · guest 98%6:00 · Matt 33% · guest 67%6:00 · Matt 33% · guest 67%9:00 · Matt 5% · guest 95%9:00 · Matt 5% · guest 95%12:00 · Matt 16.3% · guest 83.7%12:00 · Matt 16.3% · guest 83.7%15:00 · Matt 7.4% · guest 92.6%15:00 · Matt 7.4% · guest 92.6%18:00 · Matt 4.6% · guest 95.4%18:00 · Matt 4.6% · guest 95.4%21:00 · Matt 11% · guest 89%21:00 · Matt 11% · guest 89%24:00 · Matt 18.6% · guest 81.4%24:00 · Matt 18.6% · guest 81.4%27:00 · Matt 4.8% · guest 95.2%27:00 · Matt 4.8% · guest 95.2%30:00 · Matt 2.5% · guest 97.5%30:00 · Matt 2.5% · guest 97.5%33:00 · Matt 3.8% · guest 96.2%33:00 · Matt 3.8% · guest 96.2%36:00 · Matt 0% · guest 100%36:00 · Matt 0% · guest 100%39:00 · Matt 17.3% · guest 82.7%39:00 · Matt 17.3% · guest 82.7%42:00 · Matt 14.5% · guest 85.5%42:00 · Matt 14.5% · guest 85.5%45:00 · Matt 23.9% · guest 76.1%45:00 · Matt 23.9% · guest 76.1%48:00 · Matt 16.1% · guest 83.9%48:00 · Matt 16.1% · guest 83.9%51:00 · Matt 1.4% · guest 98.6%51:00 · Matt 1.4% · guest 98.6%54:00 · Matt 9.5% · guest 90.5%54:00 · Matt 9.5% · guest 90.5%57:00 · Matt 21.3% · guest 78.7%57:00 · Matt 21.3% · guest 78.7%
Sharpest disagreement ▶ 8:50 Jordan rejects the premise that big data platforms efficiently serve small data

Jordan forcefully counters the idea that cloud warehouses handle small workloads well, citing a 40x compute inefficiency tax and order-of-magnitude latency overheads.

Hardest push from Matt ▶ 8:39 Matt pushes on whether existing big data systems render small data engines redundant

Matt directly challenges Jordan's core positioning by asking why Snowflake, BigQuery, or Databricks cannot simply handle small data queries at appropriate pricing.

Biggest teaching moment ▶ 34:40 Jordan details vectorized execution compilation vs brittle hand-coded assembly

Jordan educates the host on database execution mechanics, explaining how DuckDB relies on compiler optimizations rather than hand-coded SIMD assembly to maintain performance across hardware.

Matt holds his own ▶ 45:15 Matt brings up Databricks acquiring Tabular and Apache Iceberg format adoption

Matt demonstrates domain expertise by citing high-profile data ecosystem M&A and asking how open table format shifts impact MotherDuck's market positioning.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Death of Big Data and the Rise of Small Data 3511 Matt prompts Jordan on his small data post and conference. Jordan contrasts actual user query behavior at BigQuery against industry benchmarks, with Matt adding brief contextual details on MapReduce history.
The Small Data Manifesto and Brand Positioning 5524 Matt challenges Jordan on whether existing platforms like Snowflake or BigQuery can efficiently serve small data. Jordan explains the distributed coordination latency tax and the medallion architecture presentation tier.
Understanding DuckDB and Its Academic Origins 3400 Matt asks for a background on DuckDB and its academic origins. Jordan explains embedded database architecture, Python integration, and CWI history.
Connecting with DuckDB Labs and Structuring MotherDuck 4312 Matt brings up standard VC advice about commercial startups needing to own open source communities. Jordan reframes how MotherDuck maintains an aligned partnership with DuckDB Labs while keeping independent communities.
Funding, Hybrid Execution Architecture, and Office Culture 3410 Jordan outlines MotherDuck's funding history and technical architecture. Matt asks conversational follow-up questions regarding team size and office dynamics.
Database Mechanics, Vectorized Execution, and Local Caching 4611 Matt asks Jordan to clarify in-memory analytics and what makes DuckDB fast. Jordan delivers an technical overview of Stonebraker's paper, compiler optimizations for vectorized execution, and client caching.
Local AI Models, On-Device Inference, and Local RAG 5423 Matt asks about current system limitations, local AI models, and whether small data hurts data ingestion vendors like Fivetran. Jordan details local RAG architectures and reframe Fivetran's core value.
Ecosystem Integration, Open Formats, and Reducing Friction 6523 Matt cites recent industry moves like Databricks acquiring Tabular and asks if stack complexity is truly decreasing. Jordan explains how open formats like Apache Iceberg reduce data lock-in.
Mindset and Entrepreneurial Lessons for Technical Founders 3300 Matt asks Jordan to reflect on transitioning into a first-time founder role. Jordan discusses relying on intuition and managing contradictory fundraising advice from experienced founders.
Transitioning from Engineering to Product, Marketing, and Business Leadership 3200 Matt asks how Jordan developed commercial and marketing skills coming from engineering. Jordan shares lessons learned from seeing built technology die without product and marketing alignment.

Statements from this episode (14)

Assertion Not checkable as stated
Google BigQuery's largest customers never ran 100-terabyte queries, says Jordan Tigani
“The query sizes they were using was a hundred terabytes, and I remembered back from my time at BigQuery you know, we had some of the, Largest customers in the world. We had Walmart, Home Depot, Equifax, HSBC, you know, like, and nobody was using anything, you …”
Jordan Tigani Oct 24, 2024 ▶ 2:17
Assertion Supported
Modern laptops are 10 to 100 times faster than MapReduce-era servers
“But you know, nowadays, like, you know, I've got a Mac M two laptop. It's two years old. It's like probably an order of magnitude to two orders of magnitude faster than the server machines were back when, you know, MapReduce came out and people started buildin…”
Jordan Tigani Oct 24, 2024 ▶ 4:02
Insight
Distributed data systems impose a massive complexity tax over single-node hardware
“There's just this huge tax that you pay to have to build a distributed system that scales out and can do, like, you know, distributed transactions and shuffling data and, you know, if you were gonna design something for kind of modern hardware, you could make …”
Jordan Tigani Oct 24, 2024 ▶ 4:43
Assertion Supported
MotherDuck executes queries in milliseconds compared to BigQuery's 400-millisecond overhead
“You know, we can do queries in sort of single digit milliseconds. And you know, in BigQuery, we were very, very happy when we got kind of the overhead down to like, 400 milliseconds.”
Jordan Tigani Oct 24, 2024 ▶ 10:03
Assertion Supported
Google BigQuery lacked time zone support for years due to engineering difficulty
“In BigQuery, we didn't implement time zones for, I think, like six or seven years, because it's just really hard to get right, and like, you know, there's all kinds of bizarro things in in, when you deal with you know, with time zones, and we're like, well, we…”
Jordan Tigani Oct 24, 2024 ▶ 19:24
Disclosure
MotherDuck gave DuckDB's creators a co-founder-sized equity stake at inception
“We also gave them a chunk of the company when we started, like essentially a co-founder share.”
Jordan Tigani Oct 24, 2024 ▶ 23:34
Disclosure
DuckDB's co-creator agreed not to assist any MotherDuck competitors
“On the other hand, he did agree not to work with anybody who's doing something similar to what we're doing.”
Jordan Tigani Oct 24, 2024 ▶ 23:56
Opinion
Tigani: DuckDB is amazing software, but not a data warehouse
“DuckDB is an amazing, amazing, amazing piece of software, but it's not a data warehouse.”
Jordan Tigani Oct 24, 2024 ▶ 26:00
Assertion Not checkable as stated
Tigani: DuckDB ported to Apple Silicon in just two hours
“When the Mac Silicon came out, it took them, like, two hours to make it work on the new Mac Silicon versus, like, having to, sort of, do all this, like, complex hand, hand coding hand coding stuff.”
Jordan Tigani Oct 24, 2024 ▶ 36:54
Assertion Not checkable as stated
Network latency makes 60 FPS data visualization impossible in traditional cloud architectures
“This is how you get the 60 frame per second kinds of you know, visual, visual visualization speeds that are literally impossible in more traditional architecture, because if the, if like, if I'm talking to a data center that is, you know, on the other side of …”
Jordan Tigani Oct 24, 2024 ▶ 38:06
Disclosure
MotherDuck is not well-suited for database working sets exceeding 10 terabytes
“So, I mean, scale is certainly one limitation, and I think, you know, if you're gonna push past working sets of, like, 10 terabytes or larger, then you know, I think then, you know, Mother Duck doesn't work well yet.”
Jordan Tigani Oct 24, 2024 ▶ 39:18
Prediction Not checkable as stated
Open data formats like Apache Iceberg will pressure incumbent database margins
“I think that's going to be hard for the incumbents, and I think it's going to be a net benefit to the, you know, kind of the people that are coming in with new with new tools and new ways of doing things, and it's going to put pressure on margins which again i…”
Jordan Tigani Oct 24, 2024 ▶ 46:20
Insight
Ex-Google engineers must realize marketing is war because nobody cares naturally
“I realized that marketing is war. You know, when you're at Google, you think, oh, well, you don't need to do marketing. Who needs to do marketing? But then you realize outside of Google is like, well, nobody cares about what you're doing. And you have to basic…”
Jordan Tigani Oct 24, 2024 ▶ 56:58
Assertion Not checkable as stated
SingleStore reached $100 million in ARR during Tigani's tenure as CPO
“I mean, there was, there were a hundred million in ARR”
Jordan Tigani Oct 24, 2024 ▶ 57:25
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.