Feb 1, 2021 · 25m · mad

Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)

Wes McKinney · 18m spoken Matt Turck · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC fireside chat, host Matt Turck interviews Wes McKinney, Founder and CEO of Ursa Computing, about the creation and impact of Pandas, the engineering rationale behind Apache Arrow, and the business models supporting open-source software development.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 21.2% of the talking time here. How this is scored →

Matt as informed peer 4.1 Guest teaching 5.0 Guest disagreement 1.4 Matt pushing back 1.6
05100:0010:0020:002:16–5:08 · Matt as informed peer 4/10 The Vital Role of Pandas in Data Preparation and Cleaning Matt sets up educational prompts about pandas, ML data preparation, and dataframe terminology. Wes explains the origins of dataframes from R and S+ and how pandas series and dataframes structure tabular data.5:08–7:58 · Matt as informed peer 3/10 Why Python Became the Dominant Data Science Language Matt asks why Python became the dominant data science language over alternatives. Wes clarifies that Python's dominance was not predetermined and details how a perfect storm of open-source tools enabled rapid prototyping.7:58–10:17 · Matt as informed peer 2/10 The Genesis and Core Mission of Apache Arrow Matt prompts Wes on the origins of Apache Arrow. Wes educates the audience on the 2015 data interoperability crisis across cloud data lakes and big data compute engines.10:17–13:21 · Matt as informed peer 6/10 Bridging Databases and Data Science Ecosystems Matt pushes back by asking why existing database connectors like ODBC and JDBC are not sufficient. Wes explains how row-oriented protocols create a severe throughput bottleneck, comparing it to drinking a thick milkshake through a small straw.13:21–17:29 · Matt as informed peer 7/10 How Apache Arrow Eliminates Data Conversion Bottlenecks Matt demonstrates strong domain knowledge by offering a detailed explanation of data movement bottlenecks between cloud data warehouses like Snowflake and machine learning frameworks. Wes agrees and expands on native Arrow export support across major data warehouse vendors.17:29–21:16 · Matt as informed peer 4/10 Commercializing Open Source: From Ursa Labs to Ursa Computing Matt asks about the transition from Ursa Labs as a non-profit consortium to Ursa Computing as a venture-backed commercial business. Wes explains the funding dynamics and the necessity of commercial backing for open-source scale.21:16–24:58 · Matt as informed peer 3/10 Audience Q&A: Apache Arrow vs. the Databricks Stack Matt relays audience Q&A regarding Databricks/Spark and the future of parallel computing. Wes offer direct criticism of Spark's architecture, noting that it fails to scale down efficiently to single-node computing compared to pandas.2:16–5:08 · Guest teaching 5/10 The Vital Role of Pandas in Data Preparation and Cleaning Matt sets up educational prompts about pandas, ML data preparation, and dataframe terminology. Wes explains the origins of dataframes from R and S+ and how pandas series and dataframes structure tabular data.5:08–7:58 · Guest teaching 5/10 Why Python Became the Dominant Data Science Language Matt asks why Python became the dominant data science language over alternatives. Wes clarifies that Python's dominance was not predetermined and details how a perfect storm of open-source tools enabled rapid prototyping.7:58–10:17 · Guest teaching 5/10 The Genesis and Core Mission of Apache Arrow Matt prompts Wes on the origins of Apache Arrow. Wes educates the audience on the 2015 data interoperability crisis across cloud data lakes and big data compute engines.10:17–13:21 · Guest teaching 7/10 Bridging Databases and Data Science Ecosystems Matt pushes back by asking why existing database connectors like ODBC and JDBC are not sufficient. Wes explains how row-oriented protocols create a severe throughput bottleneck, comparing it to drinking a thick milkshake through a small straw.13:21–17:29 · Guest teaching 4/10 How Apache Arrow Eliminates Data Conversion Bottlenecks Matt demonstrates strong domain knowledge by offering a detailed explanation of data movement bottlenecks between cloud data warehouses like Snowflake and machine learning frameworks. Wes agrees and expands on native Arrow export support across major data warehouse vendors.17:29–21:16 · Guest teaching 3/10 Commercializing Open Source: From Ursa Labs to Ursa Computing Matt asks about the transition from Ursa Labs as a non-profit consortium to Ursa Computing as a venture-backed commercial business. Wes explains the funding dynamics and the necessity of commercial backing for open-source scale.21:16–24:58 · Guest teaching 6/10 Audience Q&A: Apache Arrow vs. the Databricks Stack Matt relays audience Q&A regarding Databricks/Spark and the future of parallel computing. Wes offer direct criticism of Spark's architecture, noting that it fails to scale down efficiently to single-node computing compared to pandas.2:16–5:08 · Guest disagreement 1/10 The Vital Role of Pandas in Data Preparation and Cleaning Matt sets up educational prompts about pandas, ML data preparation, and dataframe terminology. Wes explains the origins of dataframes from R and S+ and how pandas series and dataframes structure tabular data.5:08–7:58 · Guest disagreement 1/10 Why Python Became the Dominant Data Science Language Matt asks why Python became the dominant data science language over alternatives. Wes clarifies that Python's dominance was not predetermined and details how a perfect storm of open-source tools enabled rapid prototyping.7:58–10:17 · Guest disagreement 1/10 The Genesis and Core Mission of Apache Arrow Matt prompts Wes on the origins of Apache Arrow. Wes educates the audience on the 2015 data interoperability crisis across cloud data lakes and big data compute engines.10:17–13:21 · Guest disagreement 2/10 Bridging Databases and Data Science Ecosystems Matt pushes back by asking why existing database connectors like ODBC and JDBC are not sufficient. Wes explains how row-oriented protocols create a severe throughput bottleneck, comparing it to drinking a thick milkshake through a small straw.13:21–17:29 · Guest disagreement 1/10 How Apache Arrow Eliminates Data Conversion Bottlenecks Matt demonstrates strong domain knowledge by offering a detailed explanation of data movement bottlenecks between cloud data warehouses like Snowflake and machine learning frameworks. Wes agrees and expands on native Arrow export support across major data warehouse vendors.17:29–21:16 · Guest disagreement 1/10 Commercializing Open Source: From Ursa Labs to Ursa Computing Matt asks about the transition from Ursa Labs as a non-profit consortium to Ursa Computing as a venture-backed commercial business. Wes explains the funding dynamics and the necessity of commercial backing for open-source scale.21:16–24:58 · Guest disagreement 3/10 Audience Q&A: Apache Arrow vs. the Databricks Stack Matt relays audience Q&A regarding Databricks/Spark and the future of parallel computing. Wes offer direct criticism of Spark's architecture, noting that it fails to scale down efficiently to single-node computing compared to pandas.2:16–5:08 · Matt pushing back 1/10 The Vital Role of Pandas in Data Preparation and Cleaning Matt sets up educational prompts about pandas, ML data preparation, and dataframe terminology. Wes explains the origins of dataframes from R and S+ and how pandas series and dataframes structure tabular data.5:08–7:58 · Matt pushing back 1/10 Why Python Became the Dominant Data Science Language Matt asks why Python became the dominant data science language over alternatives. Wes clarifies that Python's dominance was not predetermined and details how a perfect storm of open-source tools enabled rapid prototyping.7:58–10:17 · Matt pushing back 0/10 The Genesis and Core Mission of Apache Arrow Matt prompts Wes on the origins of Apache Arrow. Wes educates the audience on the 2015 data interoperability crisis across cloud data lakes and big data compute engines.10:17–13:21 · Matt pushing back 5/10 Bridging Databases and Data Science Ecosystems Matt pushes back by asking why existing database connectors like ODBC and JDBC are not sufficient. Wes explains how row-oriented protocols create a severe throughput bottleneck, comparing it to drinking a thick milkshake through a small straw.13:21–17:29 · Matt pushing back 2/10 How Apache Arrow Eliminates Data Conversion Bottlenecks Matt demonstrates strong domain knowledge by offering a detailed explanation of data movement bottlenecks between cloud data warehouses like Snowflake and machine learning frameworks. Wes agrees and expands on native Arrow export support across major data warehouse vendors.17:29–21:16 · Matt pushing back 1/10 Commercializing Open Source: From Ursa Labs to Ursa Computing Matt asks about the transition from Ursa Labs as a non-profit consortium to Ursa Computing as a venture-backed commercial business. Wes explains the funding dynamics and the necessity of commercial backing for open-source scale.21:16–24:58 · Matt pushing back 1/10 Audience Q&A: Apache Arrow vs. the Databricks Stack Matt relays audience Q&A regarding Databricks/Spark and the future of parallel computing. Wes offer direct criticism of Spark's architecture, noting that it fails to scale down efficiently to single-node computing compared to pandas.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 34.6% · guest 65.4%0:00 · Matt 34.6% · guest 65.4%3:00 · Matt 20.3% · guest 79.7%3:00 · Matt 20.3% · guest 79.7%6:00 · Matt 3.4% · guest 96.6%6:00 · Matt 3.4% · guest 96.6%9:00 · Matt 22.9% · guest 77.1%9:00 · Matt 22.9% · guest 77.1%12:00 · Matt 19.9% · guest 80.1%12:00 · Matt 19.9% · guest 80.1%15:00 · Matt 36.9% · guest 63.1%15:00 · Matt 36.9% · guest 63.1%18:00 · Matt 2.9% · guest 97.1%18:00 · Matt 2.9% · guest 97.1%21:00 · Matt 31.8% · guest 68.2%21:00 · Matt 31.8% · guest 68.2%24:00 · Matt 16.7% · guest 83.3%24:00 · Matt 16.7% · guest 83.3%
Sharpest disagreement ▶ 24:15 Wes's critique of Spark's single-node performance

Wes openly criticizes dominant big data engines like Spark, stating that they fail to scale down efficiently to single nodes and perform slower than pandas on single-node workloads.

Hardest push from Matt ▶ 11:42 Matt challenging the necessity of Arrow over ODBC/JDBC

Matt directly challenges the fundamental premise of needing Apache Arrow by asking why standard ODBC and JDBC connectors cannot simply extract data from databases.

Biggest teaching moment ▶ 11:53 Wes explaining row-oriented database connector bottlenecks

Wes explains the deep technical limitations of ODBC/JDBC protocols for bulk data transfer, using the memorable analogy of trying to drink a thick milkshake through a small straw.

Matt holds his own ▶ 14:28 Matt's explanation of cloud data warehouse transport bottlenecks

Matt steps beyond standard interviewer prompts to outline a sophisticated hypothesis regarding data movement costs between Snowflake and ML tools, earning validation from Wes.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The Vital Role of Pandas in Data Preparation and Cleaning 4511 Matt sets up educational prompts about pandas, ML data preparation, and dataframe terminology. Wes explains the origins of dataframes from R and S+ and how pandas series and dataframes structure tabular data.
Why Python Became the Dominant Data Science Language 3511 Matt asks why Python became the dominant data science language over alternatives. Wes clarifies that Python's dominance was not predetermined and details how a perfect storm of open-source tools enabled rapid prototyping.
The Genesis and Core Mission of Apache Arrow 2510 Matt prompts Wes on the origins of Apache Arrow. Wes educates the audience on the 2015 data interoperability crisis across cloud data lakes and big data compute engines.
Bridging Databases and Data Science Ecosystems 6725 Matt pushes back by asking why existing database connectors like ODBC and JDBC are not sufficient. Wes explains how row-oriented protocols create a severe throughput bottleneck, comparing it to drinking a thick milkshake through a small straw.
How Apache Arrow Eliminates Data Conversion Bottlenecks 7412 Matt demonstrates strong domain knowledge by offering a detailed explanation of data movement bottlenecks between cloud data warehouses like Snowflake and machine learning frameworks. Wes agrees and expands on native Arrow export support across major data warehouse vendors.
Commercializing Open Source: From Ursa Labs to Ursa Computing 4311 Matt asks about the transition from Ursa Labs as a non-profit consortium to Ursa Computing as a venture-backed commercial business. Wes explains the funding dynamics and the necessity of commercial backing for open-source scale.
Audience Q&A: Apache Arrow vs. the Databricks Stack 3631 Matt relays audience Q&A regarding Databricks/Spark and the future of parallel computing. Wes offer direct criticism of Spark's architecture, noting that it fails to scale down efficiently to single-node computing compared to pandas.

Statements from this episode (12)

Assertion Not checkable as stated
Over 90% of Python tabular data passes through Pandas
“I'd say, you know, nine, more than 90% of the data that's coming into structured data processing in, in the Python ecosystem, tabular data processing is passing through pandas at some point at some point in its lifetime.”
Wes McKinney Feb 1, 2021 ▶ 0:46
Assertion Supported
Pandas has had well over 2,000 open-source contributors
“I know there've been well over 2000 contributors at this point.”
Wes McKinney Feb 1, 2021 ▶ 1:29
Assertion Not checkable as stated
Pandas never had a significant corporate sponsor
“Pandas never really had a significant corporate sponsor who was you know putting in the majority of contributions. Like it really was a community project almost, you know from the get go.”
Wes McKinney Feb 1, 2021 ▶ 1:53
Assertion Partly supported
The term 'data frame' originated in the R programming language
“Data frame is a term that arose from originally in the R programming language which was based on the S and S plus programming languages.”
Wes McKinney Feb 1, 2021 ▶ 3:47
Disclosure
Apache Arrow was built to bridge database and data science developers
“For me, one of the primary motivators was to create a technology which could you know, proverbially tie the room together and enable that, that cross-pollination between database developers and data science developers that had just never never existed because …”
Wes McKinney Feb 1, 2021 ▶ 11:21
Assertion Supported
ODBC and JDBC were never designed for bulk data transfer
“Protocols like interfaces like ODBC and JDBC were never designed Or intended for bulk data transfer, like on the order of gigabytes, for example.”
Wes McKinney Feb 1, 2021 ▶ 12:11
Assertion Supported
Snowflake and Google BigQuery support exporting query results to Apache Arrow
“Snowflake exports supports exporting query results to arrow format. So it is big query.”
Wes McKinney Feb 1, 2021 ▶ 15:54
Prediction Not checkable as stated
McKinney predicts every data warehouse will soon support Apache Arrow
“So I think that, that, you know, in the course of the next few years you know, pretty much every database system, every data warehouse Is going to support aero based import and export in some format.”
Wes McKinney Feb 1, 2021 ▶ 16:59
Insight
Commercial entities are necessary to scale and sustain open-source projects
“It became clear to me and to many people that that to have more of a commercial engine behind Arrow and the Arrow ecosystem was important for enabling the ecosystem to continue to grow for us to be able to pour a lot more resources into the open source project…”
Wes McKinney Feb 1, 2021 ▶ 19:15
Assertion Not checkable as stated
Apache Arrow strictly complements rather than competes with Databricks
“It's neither it's neither a competitor or a replacement, so it's strictly a complimentary technology.”
Wes McKinney Feb 1, 2021 ▶ 21:40
Opinion
Apache Spark failed to effectively shrink down to single-node scale
“One of the things that you find with things like Spark is that they really failed to shrink down and, Effectively do computing at the single node scale.”
Wes McKinney Feb 1, 2021 ▶ 24:22
Assertion Supported
Running Apache Spark on a single node is slower than Pandas
“You can use spark at the single node scale as an alternative to pandas through the koalas interface, but you'll find that for many workloads, it's simply slower than pandas, which is not super impressive.”
Wes McKinney Feb 1, 2021 ▶ 24:30
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.