Assertion certainty 4/5 debate potential 2/5

Zaharia: Apache Spark is easier to use than prior big data systems

Matei Zaharia · a16z Podcast | A Conversation With the Inventor of Spark · Jan 2, 2019 · at 0:29

Matei Zaharia is the creator of Apache Spark and co-founder/CTO of Databricks. He explains the unique design features of Apache Spark compared to earlier data processing frameworks.

0:00 / 0:23exact quote · 23.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or graph computation or stream processing. And second, it's designed to be very easy to use, much easier to use than previous systems for working with large data.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Matei Zaharia

Insight
Zaharia: Data scientists prefer advanced work over answering routine user queries
“The interesting thing is nobody wants only the insiders to work with data, basically. Everyone wants to be able to access it directly. Actually, there was a great keynote about this at the Spark Summit by Gloria Lau, where she said that also the insiders thems…”
Matei Zaharia Jan 2, 2019 ▶ 4:55 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Not publicly verifiable
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Matei Zaharia Jan 2, 2019 ▶ 9:15 a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Mentoring open-source contributors is slower initially but scales community
“At the beginning, you know, if you're someone working on it every day and someone comes in and wants help to, you know, to get some idea in, it's always faster for you to do it yourself than to help this other person. But you have to do that. You have to Help …”
Matei Zaharia Jan 2, 2019 ▶ 12:09 a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Testing infrastructure is essential to maintain development speed in open source
“The third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this kind of infrastructure, much the same as y…”
Matei Zaharia Jan 2, 2019 ▶ 13:10 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Contradicted
Zaharia: Databricks includes all engine improvements in open-source Spark
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support o…”
Matei Zaharia Jan 2, 2019 ▶ 18:31 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Supported
Zaharia: MapReduce was created by Google for nightly web indexing
“MapReduce initially came out of Google, where it was used for web indexing, and the whole point was, I will run this giant job every night, and in the morning, it's built a new index of the web.”
Matei Zaharia Jan 2, 2019 ▶ 3:49 a16z Podcast | A Conversation With the Inventor of Spark
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.