Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 1/5

Legacy Hadoop tools like Hive, Pig, and Mahout now run on Spark

Matei Zaharia · a16z Podcast | A Conversation With the Inventor of Spark · Jan 2, 2019 · at 13:59

Matei Zaharia, creator of Apache Spark and Databricks CTO, describes how legacy Hadoop ecosystem projects adapted to integrate directly with Spark for better performance.

0:00 / 0:19exact quote · 19.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So in particular you know, one of the things we saw is many of the projects that were built on top of Hadoop, such as Hive, which is a SQL processing at scale and Pig and Mahout for machine learning are starting to run on top of Spark as well, so that users of those can get, you know, the speed ups and the benefits from using Spark.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Matei Zaharia

Insight
Zaharia: Data scientists prefer advanced work over answering routine user queries
“The interesting thing is nobody wants only the insiders to work with data, basically. Everyone wants to be able to access it directly. Actually, there was a great keynote about this at the Spark Summit by Gloria Lau, where she said that also the insiders thems…”
Matei Zaharia Jan 2, 2019 ▶ 4:55 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Not publicly verifiable
Zaharia: Apache Spark is the most active open-source data processing project
“It's actually the most active open source project in data processing in general as far as we can tell.”
Matei Zaharia Jan 2, 2019 ▶ 9:15 a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Mentoring open-source contributors is slower initially but scales community
“At the beginning, you know, if you're someone working on it every day and someone comes in and wants help to, you know, to get some idea in, it's always faster for you to do it yourself than to help this other person. But you have to do that. You have to Help …”
Matei Zaharia Jan 2, 2019 ▶ 12:09 a16z Podcast | A Conversation With the Inventor of Spark
Insight
Zaharia: Testing infrastructure is essential to maintain development speed in open source
“The third thing I need that's really important to keep a project moving quickly is just really great infrastructure for testing, checking the quality, making sure that it continues to be good. And by investing in this kind of infrastructure, much the same as y…”
Matei Zaharia Jan 2, 2019 ▶ 13:10 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Contradicted
Zaharia: Databricks includes all engine improvements in open-source Spark
“It's the same Spark that anyone else gets in the open source. All the libraries, all the improvements we put into the engine, you can just download them and run them yourselves. Or if you want you know, you can talk to a vendor that provides support. Support o…”
Matei Zaharia Jan 2, 2019 ▶ 18:31 a16z Podcast | A Conversation With the Inventor of Spark
Assertion Not checkable as stated
Zaharia: Apache Spark is easier to use than prior big data systems
“So Spark is software for processing large volumes of data on a cluster, and the things that make it unique are, first of all, it has a very powerful programming model that lets you do many kinds of advanced analytics and processing, such as machine learning or…”
Matei Zaharia Jan 2, 2019 ▶ 0:29 a16z Podcast | A Conversation With the Inventor of Spark
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.