Matei Zaharia, co-founder and CTO of Databricks and creator of Apache Spark, describes Databricks' commercial strategy of offering Spark as a managed cloud service while keeping open-source Spark fully featured.
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support on it yourself, and there isn't any tension for us between you know do we put something in Spark, or does it become some kind of premium feature?”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Matei Zaharia
Insight
Zaharia: Data scientists prefer advanced work over answering routine user queries
“The interesting thing is nobody wants only the insiders to work with data, basically. Everyone wants to be able to access it directly. Actually, there was a great keynote about this at the Spark Summit by Gloria Lau, where she said that also the insiders thems…”
Matei ZahariaJan 2, 2019▶ 4:55a16z Podcast | A Conversation With the Inventor of Spark
AssertionNot publicly verifiable
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Matei ZahariaJan 2, 2019▶ 9:15a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Mentoring open-source contributors is slower initially but scales community
“At the beginning, you know, if you're someone working on it every day and someone comes in and wants help to, you know, to get some idea in, it's always faster for you to do it yourself than to help this other person. But you have to do that. You have to Help …”
Matei ZahariaJan 2, 2019▶ 12:09a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Testing infrastructure is essential to maintain development speed in open source
“The third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this kind of infrastructure, much the same as y…”
Matei ZahariaJan 2, 2019▶ 13:10a16z Podcast | A Conversation With the Inventor of Spark
AssertionNot checkable as stated
Zaharia: Apache Spark is easier to use than prior big data systems
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or…”
Matei ZahariaJan 2, 2019▶ 0:29a16z Podcast | A Conversation With the Inventor of Spark
AssertionSupported
Zaharia: MapReduce was created by Google for nightly web indexing
“MapReduce initially came out of Google, where it was used for web indexing, and the whole point was, I will run this giant job every night, and in the morning, it's built a new index of the web.”
Matei ZahariaJan 2, 2019▶ 3:49a16z Podcast | A Conversation With the Inventor of Spark
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 1,000 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.