Wes McKinney, creator of Pandas and Apache Arrow, explains why he transitioned from open-source consortium Ursa Labs to commercial startup Ursa Computing.
“It became clear to me and to many people that that to have more of a commercial engine behind Arrow and the Arrow ecosystem was important for enabling the ecosystem to continue to grow for us to be able to pour a lot more resources into the open source project.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Wes McKinney
Opinion
Apache Spark failed to effectively shrink down to single-node scale
“One of the things that you find with things like Spark is that they really failed to shrink down and, Effectively do computing at the single node scale.”
Wes McKinneyFeb 1, 2021▶ 24:22Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
AssertionNot checkable as stated
Over 90% of Python tabular data passes through Pandas
“I'd say, you know, nine, more than 90% of the data that's coming into structured data processing in, in the Python ecosystem, tabular data processing is passing through pandas at some point at some point in its lifetime.”
Wes McKinneyFeb 1, 2021▶ 0:46Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
PredictionNot checkable as stated
McKinney predicts every data warehouse will soon support Apache Arrow
“So I think that, that, you know, in the course of the next few years you know, pretty much every database system, every data warehouse Is going to support aero based import and export in some format.”
Wes McKinneyFeb 1, 2021▶ 16:59Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
AssertionSupported
Running Apache Spark on a single node is slower than Pandas
“You can use spark at the single node scale as an alternative to pandas through the koalas interface, but you'll find that for many workloads, it's simply slower than pandas, which is not super impressive.”
Wes McKinneyFeb 1, 2021▶ 24:30Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
AssertionNot checkable as stated
Pandas never had a significant corporate sponsor
“Pandas never really had a significant corporate sponsor who was you know putting in the majority of contributions. Like it really was a community project almost, you know from the get go.”
Wes McKinneyFeb 1, 2021▶ 1:53Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
AssertionSupported
ODBC and JDBC were never designed for bulk data transfer
“Protocols like interfaces like ODBC and JDBC were never designed
Or intended for bulk data transfer, like on the order of gigabytes, for example.”
Wes McKinneyFeb 1, 2021▶ 12:11Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.