Apache Spark, every mention
26 scenes (2021), the whole family · ← back to Apache Spark
every year 2021 anyone Matt Turck 59Haoyuan Li 13Ali Ghodsi 11Stefan Groschupf 10Praveen Murugesan 10Julien Le Dem 9Tobi Knaup 6Matt Housley 6Christopher Nguyen 6Prat Moghe 5
Verbatim, from the transcripts: the passages where Apache Spark comes up
Top 10 Trends in AI, Machine Learning and Data for 2022
- ▶ 7:43 Matt Turck And then the spark came along and that was like another whole thing that took several years.
Fireside Chat: Zhamak Dehghani (Founder, Data Mesh) with Matt Turck (Partner, FirstMark)
- ▶ 19:01 Zhamak Dehghani So a lot of people still use a spark or beam or, you know, whatever.
Fireside Chat: Abe Gong (Founder & CEO, Superconductive) with Matt Turck (Partner, FirstMark)
- ▶ 11:50 Abe Gong Even in other places, like within, um, typed data frames in Spark or, you know, your choice of data warehouses, being able to do sets, ranges, regular expressions, uh, distributions, uh, correlations, uh, things like that start to be…
- ▶ 17:14 Abe Gong Uh, we also do Spark data frames, um, and then SQL, uh, through SQL alchemy, uh, which also implies a whole slew of different SQL dialects.
Fireside Chat: Nick Schrock (Founder & CEO, Elementl) with Matt Turck (Partner, FirstMark)
- ▶ 4:38 Nick Schrock You know, Hadoop and then spark and now the cloud data warehouse.
- ▶ 26:29 Nick Schrock And then actually, you know, I've been really impressed with the development of Spark over the last few years.
Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark)
- ▶ 0:18 Matt Turck So Amplabs, Spark, and Databricks, how did it all start?
- ▶ 6:30 Matt Turck So to close on, on, on that, um, you know, chapter of the early years, um, how did you go from this academic, uh, very popular open source project, uh, which was Spark to 2 times in the scene
- ▶ 11:17 Matt Turck Uh, yeah, I had the, um, uh, pleasure and honor of, like, hosting your co-founder and CEO at the time, Stoica in 2015, and the conversation, I rewatched it before this, and the conversation at the time was all about, you know, the, the,… 4 times in the scene
- ▶ 22:58 Ali Ghodsi So in the past, when someone wanted to do SQL or warehousing on Databricks, we would offer them Spark. 4 times in the scene
- ▶ 26:03 Ali Ghodsi When we had spark and the founders were discussing, what should the name of the company be? 4 times in the scene
- ▶ 28:48 Ali Ghodsi Of course, some of these projects, when they get older, like spark, they move into the maintenance side. 2 times in the scene
- ▶ 33:25 Ali Ghodsi And if we were just doing spark, like we were on this show in, that would have been probably 10, five percent because, you know, over time, these technologies become mature and, you know, the excitement around them, uh, wanes.
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 5:58 Dave Burgess And so we query Intelli, uh, to Presto, to Hive, uh, to Spark SQL, uh, to MySQL, but you can do it with other engines too.
- ▶ 7:56 Dave Burgess So data engineering is, uh, we, we have many, many tools and we can, we can maybe cover that a bit later, but the, for the analytics itself, uh, we focus on, uh, using Hadoop and spark. 2 times in the scene
- ▶ 31:08 Dave Burgess They usually either Spark or Hive or Presto jobs or Spark SQL and, uh, just process the data in every step and, and persist the data back to S three along the way.
Fireside Chat: Bindu Reddy (Founder & CEO, Abacus.AI) with Matt Turck (Partner, FirstMark)
- ▶ 12:02 Bindu Reddy Uh, Kubernetes, um, you know, Spark, um, Redis.
- ▶ 28:26 Bindu Reddy The other thing I think from a data science perspective, the thing which has really, really born the test of time has been Spark. 2 times in the scene
Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark)
- ▶ 3:26 Savin Goyal Uh, so like the storage layer for our data warehouse, and we use Spark, Presto, Snowflake, uh, as our query engines. 2 times in the scene
Introducing Kedro
- ▶ 11:59 Kedro Product Manager Um, the data catalog supports multiple integrations of pandas, spark, desk,
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 7:35 Julien Le Dem You know, so I talked to Wes, uh, obviously, uh, we like spark contributors, um, GBT and all the, the air flow and all the very, um, popular frameworks to schedule and process data. 6 times in the scene
- ▶ 22:42 Julien Le Dem I think right now, today, Spark is one of the big projects that people are using. 3 times in the scene
Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark)
- ▶ 10:31 Matt Turck And then on the other side, you have, uh, the world, uh, of the machine learning and data analysis tools, which is like Spark and NumPy and, and so on and so forth.
- ▶ 21:50 Wes McKinney Spark, uh, Spark supports Arrow as a, as an interchange format, and it's used heavily in the interface with Python and R, for example. 2 times in the scene
- ▶ 22:42 Matt Turck After HPC, um, slash Hadoop, slash Spark, slash Ray, what's the long-term future of parallel compute for data intensive workflows? 3 times in the scene
Fireside Chat: Alok Gupta (Head of Data Science & ML, DoorDash) with Matt Turck (Partner, FirstMark)
- ▶ 7:24 Alok Gupta We use Python and Databricks to, uh, and Spark to pull data in, build models. 2 times in the scene