Apache Spark, every mention

108 scenes, the whole family · ← back to Apache Spark

tap a year for its mentions
003086015201420152016201720182019202020212022202320242025episodesmentions
0815201420152016201720182019202020212022202320242025episodes it came up in
0047.5815201420152016201720182019202020212022202320242025episodesmentions per episode

every year anyone Matt Turck 59Haoyuan Li 13Ali Ghodsi 11Stefan Groschupf 10Praveen Murugesan 10Julien Le Dem 9Tobi Knaup 6Matt Housley 6Christopher Nguyen 6Prat Moghe 5

Verbatim, from the transcripts: the passages where Apache Spark comes up

loading…

Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy Jan 23, 2025 · 5 mentions

  • ▶ 20:58 Matt Turck Spark or Kafka or like any of those frameworks, what would you recommend next once I have my language, I have my SQL, uh, what do I do next? 3 times in the scene
  • ▶ 27:12 Ben Rogojan (Seattle Data Guy) Like there, there are cases where maybe they don't know what a data warehouse is, or, you know, when you say data pipeline, maybe they, they don't understand, or when you reference Spark, they're like, I, you know, why do I care what Spark
  • ▶ 40:31 Ben Rogojan (Seattle Data Guy) And then from there, you could just say like, I'm going to use Spark.

AI at ZoomInfo: Superpowering GTM teams | Ali Dasdan, CTO, ZoomInfo Sep 19, 2024 · 1 mention

  • ▶ 16:26 Ali Dasdan All these connections are there that we are extracting out of that, and that is done through, you know, different technologies, you know, either spark code or data flow that GCP has, and all kinds of basic processing.

Vector databases and the $8 trillion open source market | Bob van Luijt, CEO of Weaviate Feb 15, 2024 · 1 mention

  • ▶ 18:54 Bob van Luijt The, on a serious note, why this becomes so interesting, at some point we saw, like, it started for Weaviate in an uptick that we saw, so community contributions to, um, our Spark connector.

Build Fast APIs Faster Over Data at Scale | Tinybird Founder & CEO Jorge Gomez Sancha Mar 9, 2023 · 1 mention

Data Visibility & Control | BigID Co-Founder & CEO Dimitri Sirota Jan 30, 2023 · 1 mention

  • ▶ 8:57 Dimitri Sirota Now, they may differ, so for instance, for Hadoop, we could do, like, MapReduce, uh, we could do, um, direct, uh, connectivity to, um, uh, uh, Hive or Spark we could use, but they all use native protocols to scan the underlying system.

Fundamentals of Data Engineering | Joe Reis and Matt Housley Oct 24, 2022 · 6 mentions

  • ▶ 2:40 Matt Housley Yeah, exactly, and, and I'll kind of skip a bullet point and then go back, but like this one right here, what we kept hearing a lot is that, you know, data engineering is Spark, or data engineering is Kafka.
  • ▶ 14:38 Matt Housley Or, or we get, like, things like, well, it's really all about Spark. 2 times in the scene
  • ▶ 30:50 Matt Housley Um, I, I think when we, I, I think, and correct me if I misunderstood the question, but I think when we talk about not defining, setting definitions around technology, we mean specifically not saying that data engineering is about Spark,… 3 times in the scene

A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy Sep 12, 2022 · 1 mention

  • ▶ 2:56 Gleb Mezhanskiy For example, your Airflow orchestrator scheduler is broken, or your cluster, like Spark cluster is underwater and backlogged, or your vendor that you use to buy data ships to something which is, ah, of low quality.

The Next Layer of the Modern Data Stack | dbt's Tristan Handy Jun 28, 2022 · 1 mention

  • ▶ 7:34 Tristan Handy Um, and that's, it's, again, I, sometimes, like, people get defensive, the data engineers in the audience, this is not a diatribe against data engineers, it's just that there are actually two orders of magnitude more human beings on the…

Top 10 Trends in AI, Machine Learning and Data for 2022 Oct 27, 2021 · 1 mention

  • ▶ 7:43 Matt Turck And then the spark came along and that was like another whole thing that took several years.

Fireside Chat: Zhamak Dehghani (Founder, Data Mesh) with Matt Turck (Partner, FirstMark) Oct 27, 2021 · 1 mention

Fireside Chat: Abe Gong (Founder & CEO, Superconductive) with Matt Turck (Partner, FirstMark) Jun 21, 2021 · 2 mentions

  • ▶ 11:50 Abe Gong Even in other places, like within, um, typed data frames in Spark or, you know, your choice of data warehouses, being able to do sets, ranges, regular expressions, uh, distributions, uh, correlations, uh, things like that start to be…
  • ▶ 17:14 Abe Gong Uh, we also do Spark data frames, um, and then SQL, uh, through SQL alchemy, uh, which also implies a whole slew of different SQL dialects.

Fireside Chat: Nick Schrock (Founder & CEO, Elementl) with Matt Turck (Partner, FirstMark) Jun 21, 2021 · 2 mentions

  • ▶ 4:38 Nick Schrock You know, Hadoop and then spark and now the cloud data warehouse.
  • ▶ 26:29 Nick Schrock And then actually, you know, I've been really impressed with the development of Spark over the last few years.

Fireside Chat: Ali Ghodsi (Founder & CEO, Databricks) with Matt Turck (Partner, FirstMark) May 24, 2021 · 18 mentions

  • ▶ 0:18 Matt Turck So Amplabs, Spark, and Databricks, how did it all start?
  • ▶ 6:30 Matt Turck So to close on, on, on that, um, you know, chapter of the early years, um, how did you go from this academic, uh, very popular open source project, uh, which was Spark to 2 times in the scene
  • ▶ 11:17 Matt Turck Uh, yeah, I had the, um, uh, pleasure and honor of, like, hosting your co-founder and CEO at the time, Stoica in 2015, and the conversation, I rewatched it before this, and the conversation at the time was all about, you know, the, the,… 4 times in the scene
  • ▶ 22:58 Ali Ghodsi So in the past, when someone wanted to do SQL or warehousing on Databricks, we would offer them Spark. 4 times in the scene
  • ▶ 26:03 Ali Ghodsi When we had spark and the founders were discussing, what should the name of the company be? 4 times in the scene
  • ▶ 28:48 Ali Ghodsi Of course, some of these projects, when they get older, like spark, they move into the maintenance side. 2 times in the scene
  • ▶ 33:25 Ali Ghodsi And if we were just doing spark, like we were on this show in, that would have been probably 10, five percent because, you know, over time, these technologies become mature and, you know, the excitement around them, uh, wanes.

Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark) Apr 5, 2021 · 4 mentions

  • ▶ 5:58 Dave Burgess And so we query Intelli, uh, to Presto, to Hive, uh, to Spark SQL, uh, to MySQL, but you can do it with other engines too.
  • ▶ 7:56 Dave Burgess So data engineering is, uh, we, we have many, many tools and we can, we can maybe cover that a bit later, but the, for the analytics itself, uh, we focus on, uh, using Hadoop and spark. 2 times in the scene
  • ▶ 31:08 Dave Burgess They usually either Spark or Hive or Presto jobs or Spark SQL and, uh, just process the data in every step and, and persist the data back to S three along the way.

Fireside Chat: Bindu Reddy (Founder & CEO, Abacus.AI) with Matt Turck (Partner, FirstMark) Apr 5, 2021 · 3 mentions

  • ▶ 12:02 Bindu Reddy Uh, Kubernetes, um, you know, Spark, um, Redis.
  • ▶ 28:26 Bindu Reddy The other thing I think from a data science perspective, the thing which has really, really born the test of time has been Spark. 2 times in the scene

Fireside Chat: Savin Goyal (ML Infra team (Metaflow), Netflix) with Matt Turck (Partner, FirstMark) Feb 17, 2021 · 2 mentions

  • ▶ 3:26 Savin Goyal Uh, so like the storage layer for our data warehouse, and we use Spark, Presto, Snowflake, uh, as our query engines. 2 times in the scene

Introducing Kedro Feb 17, 2021 · 1 mention

Data Observability and Pipelines: OpenLineage and Marquez Feb 1, 2021 · 9 mentions

  • ▶ 7:35 Julien Le Dem You know, so I talked to Wes, uh, obviously, uh, we like spark contributors, um, GBT and all the, the air flow and all the very, um, popular frameworks to schedule and process data. 6 times in the scene
  • ▶ 22:42 Julien Le Dem I think right now, today, Spark is one of the big projects that people are using. 3 times in the scene

Fireside Chat: Wes McKinney (Founder & CEO, Ursa Computing) with Matt Turck (Partner, FirstMark) Feb 1, 2021 · 6 mentions

  • ▶ 10:31 Matt Turck And then on the other side, you have, uh, the world, uh, of the machine learning and data analysis tools, which is like Spark and NumPy and, and so on and so forth.
  • ▶ 21:50 Wes McKinney Spark, uh, Spark supports Arrow as a, as an interchange format, and it's used heavily in the interface with Python and R, for example. 2 times in the scene
  • ▶ 22:42 Matt Turck After HPC, um, slash Hadoop, slash Spark, slash Ray, what's the long-term future of parallel compute for data intensive workflows? 3 times in the scene

Fireside Chat: Alok Gupta (Head of Data Science & ML, DoorDash) with Matt Turck (Partner, FirstMark) Feb 1, 2021 · 2 mentions

  • ▶ 7:24 Alok Gupta We use Python and Databricks to, uh, and Spark to pull data in, build models. 2 times in the scene

Apache Druid & An Introduction to Data Rivers // FJ Yang, Imply (FirstMark's Data Driven NYC) Jun 12, 2019 · 2 mentions

  • ▶ 1:12 FJ Yang And, ah, how many people have worked with what kind of different types of big data technologies like Hadoop, Spark, and Kafka, and others?
  • ▶ 7:30 FJ Yang You now have machine learning and AI engines like Apache Spark that basically follow this architecture model.

Fireside Chat: Solmaz Shahalizadeh, VP of Data Science & Engineering at Shopify (Data Driven NYC) Jun 12, 2019 · 4 mentions

  • ▶ 8:37 Solmaz Shahalizadeh So, um, we basically moved from, uh, using Vertica, uh, to building a ETL tool in-house using, uh, Spark, uh, and, uh, Python. 4 times in the scene

Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC) Mar 19, 2019 · 5 mentions

  • ▶ 9:35 Justin Borgman That means, you know, Hadoop's various projects can read this data, Spark can read this data, and of course, Presto can read this data as well, um, but because it's an open file format,
  • ▶ 17:54 unnamed speaker I was wondering, given that Spark is also separated from storage, um, how does Starburst stay dependent with companies like Databricks? 4 times in the scene

Fireside Chat: Mike Tuchen, CEO of Talend (TLND) (FirstMark's Data Driven NYC) Oct 17, 2018 · 7 mentions

  • ▶ 8:07 Matt Turck Uh, created, but you have, so the databases, this is what I'm talking about, the databases, um, on one side as a source, and then you have the processing layer, uh, that used to be Teradata Oracle IBM, which is now Spark, Snowflake,… 7 times in the scene

The Launch of Dataiku 5 // Florian Douetteau, Dataiku (FirstMark's Data Driven NYC) Sep 17, 2018 · 1 mention

  • ▶ 20:35 Florian Douetteau So to give you some example, last year we added support for a Spark ML for machine learning, and we added support for TensorFlow with this release.

3 Heretical Ideas on the Future of Data // Ajay Kulkarni, TimescaleDB (FirstMark's Data Driven NYC) Sep 17, 2018 · 1 mention

Fireside Chat: Chris Dixon, General Partner at Andreessen Horowitz (FirstMark's Data Driven) Jun 8, 2018 · 1 mention

  • ▶ 52:57 Chris Dixon a company called Databricks, which is, uh, built around Spark, which is essentially, you know,

Fireside Chat with Bob Muglia, CEO at Snowflake (FirstMark's Data Driven) Apr 9, 2018 · 2 mentions

  • ▶ 17:19 Matt Turck So the Hadoop and Spark ecosystems are friend or foe?
  • ▶ 25:37 Bob Muglia So we use, we, you know, we internally, uh, uh, we use R and Spark and things like that internally to do machine analytics.

Where Should Machines Go to Learn? // Auren Hoffman, SafeGraph (FirstMark's Data Driven) Nov 20, 2017 · 2 mentions

  • ▶ 3:41 Auren Hoffman You've got some sort of, like, Spark infrastructure streaming the data in with Kafka.
  • ▶ 11:43 Auren Hoffman So, these are companies like Palantir, they're the BI tools, even things like, you know, Hadoop or Spark, um, you know, basically any, most of these companies are basically, let me take your own data and help you make better decisions with…

Three Loops of Analytics Efficiency // Sean Kandel, Trifacta (FirstMark's Data Driven) Jul 13, 2017 · 1 mention

  • ▶ 6:41 Sean Kandel Uh, so probably still today the most common is using kind of hand coding tools, um, so programming languages, Python, Spark,

Big Data as a Service // Prat Moghe, Cazena (FirstMark's Data Driven) May 24, 2017 · 5 mentions

  • ▶ 4:52 Prat Moghe It's Hadoop, it's Spark, it's Python, it's R, and it's like every three months there's a new open source project around it.
  • ▶ 10:17 Prat Moghe You're running Spark.
  • ▶ 13:22 Prat Moghe Or, I'm running a data engineering job on Spark, and I got certain SLA, and I have certain price points, and go run it for me.
  • ▶ 13:52 Prat Moghe It could, they could be SQL, they could be Spark, um, and we picked certain, uh, tools here. 2 times in the scene

Project Jupyter // Jason Grout & Sylvain Corlay, Bloomberg (FirstMark's Data Driven) Feb 3, 2017 · 5 mentions

  • ▶ 18:18 unnamed speaker So as Spark starts to kind of work its tentacles into every corner of the big data ecosystem, I'm seeing a lot of interest in projects like the Apache Zeppelin project as compared to, to Jupiter. 5 times in the scene

Making Big Data Accessible Using the Cloud // Ashish Thusoo, Qubole [FirstMark's Data Driven] Nov 9, 2016 · 1 mention

  • ▶ 3:07 Ashish Thusoo So, um, you know, big data, essentially the emergence of these new systems, we hear about systems like Hadoop, Spark, Hive, and so on and so forth.

Lessons Learned from Advanced Data Science Orgs // Domino Data Lab [FirstMark's Data Driven] Nov 9, 2016 · 5 mentions

  • ▶ 15:56 Nick (Domino Data Lab) Probably the most, well, I don't know whether it's surprising or not, um, I think there's still more hype around Spark than actual value extraction from it. 5 times in the scene

Making On-Demand Delivery Profitable // Jeremy Stanley, Instacart (Data Driven NYC / FirstMark) Sep 30, 2016 · 1 mention

  • ▶ 17:37 Jeremy Stanley Um, we use Spark, uh, when we have to, uh, try to stay away from having to use those systems as long as we can.

The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark) Sep 30, 2016 · 10 mentions

  • ▶ 4:41 Praveen Murugesan And, ah, on top of HDFS, we basically have, like, Spark, and, ah, Presto, and Hive.
  • ▶ 7:18 Praveen Murugesan So, a few things, like I talked about, strict schema management, so we actually built, like, a central schema repository which is used for schema management, and then, ah, we unlocked, like, a whole bunch of new tools with, like, data on… 2 times in the scene
  • ▶ 8:19 Praveen Murugesan So, basically, think about it as, like, if you are a first-time Spark developer, you, you, your, your goal is not to learn Spark, like, in-depth and, like, play around with, like, a hundred different knobs. 5 times in the scene
  • ▶ 15:24 Praveen Murugesan And, ah, so given that it's in a Hive UDF, you also can use it on Spark. 2 times in the scene

Why Marketing is All About Data // Nitay Joffe, ActionIQ [FirstMark's Data Driven] Jun 16, 2016 · 2 mentions

  • ▶ 8:48 Nitay Joffe And so in the upper left, you see a suite of BI systems, Hadoop, Spark, and so on and so forth. 2 times in the scene

The Path to A.I. Augmented Human Intelligence // Christopher Nguyen, Arimo [FirstMark's Data Driven] Jun 16, 2016 · 6 mentions

  • ▶ 5:02 Christopher Nguyen I did not pre-aggregate this, and this is, we use Spark for in-memory processing, um, but the idea is that we've made it so simple to use that I, as a business user, I have not done any, ah, you know, any kind of SQL.
  • ▶ 9:03 Christopher Nguyen If you, ah, if I move too fast through here, just Google Spark Deep Learning Reference Architecture. 4 times in the scene
  • ▶ 12:03 Christopher Nguyen One is Spark Only.

Big Data in Insurance // Louis DiModugno, Chief Data Officer at AXA US [FirstMark's Data Driven] May 23, 2016 · 1 mention

The Journey to Information for Everyone // Prakash Nanduri, Paxata (Hosted by FirstMark) Apr 13, 2016 · 1 mention

  • ▶ 13:33 Prakash Nanduri The technology around distributed computing and what's going on with in-memory scale-out, Spark, and Luxia, and all the wonderful solutions that, ah, enabling technologies that are coming up are really important.

A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark) Apr 13, 2016 · 13 mentions

  • ▶ 0:51 Haoyuan Li Funding commuter of Apache Spark as well.
  • ▶ 2:08 Haoyuan Li The same lab produced Apache Mesos as well as Apache Spark, and we open sourced it second year April, which was around three years ago, on the Apache two-point-one license, and the latest release is version 1.1, which was actually March… 2 times in the scene
  • ▶ 5:05 Haoyuan Li Say for example, Sim Lab will produce Apache Spark.
  • ▶ 5:52 Haoyuan Li You have Spark, you have MapReduce, you have Flink, you have HBase, Presto.
  • ▶ 11:22 Haoyuan Li Basically, they run Spark and Spark SQL on top of Aluxio. 7 times in the scene
  • ▶ 17:01 Haoyuan Li people, like, inclined to believe in the project created by this lab, so when I started this project, it's already, like, four years later than other projects from the lab, like, ah, particularly, like, Mesos and Spark.

The Solitude of the Data Team Manager // Florian Douetteau, Dataiku (Hosted by FirstMark) Apr 13, 2016 · 3 mentions

  • ▶ 3:56 Florian Douetteau You might want to do everything in Spark, but you would find out that maybe the geo-analytics team in the company is really into Postgres and PostGIS, and want to do everything in Postgres and PostGIS.
  • ▶ 20:21 Florian Douetteau Um, because the, the, many companies have the vision of having, like, data, data analysts, largely speaking, moving from a SQL, and possibly SAS world, to a kind of Hadoop-ish, Spark-ish world, on the long term. 2 times in the scene

Combining Machine Learning With Expert Human Judgement // Eric Colson, Stitch Fix Mar 18, 2016 · 1 mention

The Benefits of Fast Business Intelligence // Amir Orad, Sisense (Hosted by FirstMark Capital) Jan 25, 2016 · 2 mentions

  • ▶ 17:23 Matt Turck When, when, when, uh, presumably you go talk to IT buyers or even marketing buyers, uh, this jungle of companies out there, uh, and people have heard all the buzz terms and all the Hadoop and the Spark and 2 times in the scene

B2B Big Data Challenges, Nick Mehta, Gainsight (Data Driven NYC / FirstMark Capital) Dec 17, 2015 · 1 mention

  • ▶ 21:00 Nick Mehta Like, by the way, we're going to do, um, Spark, and we'll probably, maybe after the sales pitch, I'll do MapR and Datamir and everything else I do, right?

A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark) Dec 17, 2015 · 2 mentions

  • ▶ 20:58 M.C. Srivas We introduced JSON to Hadoop and Spark and in MapReduce and in Hive and everywhere. 2 times in the scene

The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer Dec 17, 2015 · 10 mentions

  • ▶ 10:12 Stefan Groschupf Um, who thinks Spark is disruptive? 6 times in the scene
  • ▶ 11:51 Stefan Groschupf Oh yeah, we need spark, we need real time, we need this, this, this, but the, the, the thought, the think process, the thought process, I really believe needs to have happen upside down, bottom up, because it's, think about how difficult… 3 times in the scene
  • ▶ 15:36 Stefan Groschupf And finally, Yarn is very batch-centric, where again, now we kind of make it work with Spark and, and, and other things now, but, uh, Mesosphere really has a very flexible scheduling mechanism.

Black Boxes and Unicorns - DataRobot CEO Jeremy Achin Nov 23, 2015 · 1 mention

  • ▶ 18:30 Jeremy Achin Um, so we'll, we'll compete algorithms from R, from Python, from H-to-O, Spark, um, and just kind of compete them all against each other and see what works best for yours, right?

Liz Crawford, Birchbox // Data Science & Analytics at Birchbox (Hosted by FirstMark Capital) Oct 21, 2015 · 2 mentions

  • ▶ 4:49 Liz Crawford So we don't expect our data scientists to be able to produce our new spark infrastructure, for example.
  • ▶ 13:02 Liz Crawford We didn't start out with all of them on a, on Spark, for example.

Ramana Rao, Livefyre // Real-Time Social Engagement (Hosted by FirstMark Capital) Oct 21, 2015 · 2 mentions

Joseph Essas, OpenTable // Mining Diner Talk (Hosted by FirstMark Capital) Jun 19, 2015 · 2 mentions

  • ▶ 2:32 Joseph Essas Uh, all of our events flowing through Kafka, they've been populated into Cassandra, which then we run Spark instances that kind of model on top of the data. 2 times in the scene

Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital) Apr 2, 2015 · 37 mentions

  • ▶ 0:19 Matt Turck So, uh, so as a quick intro, you are one of the co-creators of Apache Spark, uh, which is a unified framework for building, uh, data pipelines, and also one of the most active, um, open source projects in, in big data. 2 times in the scene
  • ▶ 5:45 Ion Stoica Um, the next project was Spark, and Spark was, ah, you know, we targeted first some workloads which are not covered by Hadoop, and from all this experience I mentioned earlier, we look at iterative, iterative computations to support…
  • ▶ 6:15 Matt Turck Uh, so, uh, you know, to the sort of uninitiated, uh, the, you know, Hadoop was the big thing, uh, for the last two to three years, and then Spark sort of appeared on the scene. 4 times in the scene
  • ▶ 7:24 Matt Turck Yes, so perhaps more specifically, can you go through the, the, the key advantages of Spark over MapReduce? 6 times in the scene
  • ▶ 9:37 Ion Stoica So there's a key here, you know, one way to look at Spark is that think about 3 times in the scene
  • ▶ 12:07 Matt Turck And to take just one of the, one of the things you mentioned, um, so maybe contrast Spark streaming with Storm, which also does micro-batching. 4 times in the scene
  • ▶ 15:30 Ion Stoica And it has, ah, a few things to make it easy, ah, in addition to Spark.
  • ▶ 16:31 Matt Turck What is the, uh, relationship between Databricks and the Spark community specifically? 6 times in the scene
  • ▶ 19:06 unnamed speaker I just wanted to get your thought on that and whether you guys were considering that for Spark. 2 times in the scene
  • ▶ 20:44 Matt Turck You've officially been appointed Spark Questionnaire. 8 times in the scene
page 1 of 2 · 100 scenes per page · newest episode first next →
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.