Matei Zaharia, creator of Apache Spark and CTO of Databricks, explains how Lester Mackey's recommendation workload for the Netflix Prize was one of the initial target applications for Spark at UC Berkeley.
“So it's actually one of the applications that I first tried to support in Spark was you know, the recommendation algorithm he was working on.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Matei Zaharia
Insight
Zaharia: Data scientists prefer advanced work over answering routine user queries
“The interesting thing is nobody wants only the insiders to work with data, basically. Everyone wants to be able to access it directly. Actually, there was a great keynote about this at the Spark Summit by Gloria Lau, where she said that also the insiders thems…”
Matei ZahariaJan 2, 2019▶ 4:55a16z Podcast | A Conversation With the Inventor of Spark
AssertionNot publicly verifiable
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Matei ZahariaJan 2, 2019▶ 9:15a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Mentoring open-source contributors is slower initially but scales community
“At the beginning, you know, if you're someone working on it every day and someone comes in and wants help to, you know, to get some idea in, it's always faster for you to do it yourself than to help this other person. But you have to do that. You have to Help …”
Matei ZahariaJan 2, 2019▶ 12:09a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Testing infrastructure is essential to maintain development speed in open source
“The third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this kind of infrastructure, much the same as y…”
Matei ZahariaJan 2, 2019▶ 13:10a16z Podcast | A Conversation With the Inventor of Spark
AssertionContradicted
Zaharia: Databricks includes all engine improvements in open-source Spark
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support o…”
Matei ZahariaJan 2, 2019▶ 18:31a16z Podcast | A Conversation With the Inventor of Spark
AssertionNot checkable as stated
Zaharia: Apache Spark is easier to use than prior big data systems
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or…”
Matei ZahariaJan 2, 2019▶ 0:29a16z Podcast | A Conversation With the Inventor of Spark
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 1,000 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.