Apache Spark, every mention
5 scenes (2022), the whole family · ← back to Apache Spark
every year 2022 anyone Matt Turck 59Haoyuan Li 13Ali Ghodsi 11Stefan Groschupf 10Praveen Murugesan 10Julien Le Dem 9Tobi Knaup 6Matt Housley 6Christopher Nguyen 6Prat Moghe 5
Verbatim, from the transcripts: the passages where Apache Spark comes up
Fundamentals of Data Engineering | Joe Reis and Matt Housley
- ▶ 2:40 Matt Housley Yeah, exactly, and, and I'll kind of skip a bullet point and then go back, but like this one right here, what we kept hearing a lot is that, you know, data engineering is Spark, or data engineering is Kafka.
- ▶ 14:38 Matt Housley Or, or we get, like, things like, well, it's really all about Spark. 2 times in the scene
- ▶ 30:50 Matt Housley Um, I, I think when we, I, I think, and correct me if I misunderstood the question, but I think when we talk about not defining, setting definitions around technology, we mean specifically not saying that data engineering is about Spark,… 3 times in the scene
A Novel Approach to Data Quality for the Modern Data Stack | Datafold’s Gleb Mezhanskiy
- ▶ 2:56 Gleb Mezhanskiy For example, your Airflow orchestrator scheduler is broken, or your cluster, like Spark cluster is underwater and backlogged, or your vendor that you use to buy data ships to something which is, ah, of low quality.
The Next Layer of the Modern Data Stack | dbt's Tristan Handy
- ▶ 7:34 Tristan Handy Um, and that's, it's, again, I, sometimes, like, people get defensive, the data engineers in the audience, this is not a diatribe against data engineers, it's just that there are actually two orders of magnitude more human beings on the…