Spark SQL, every mention
8 scenes across 2 shows · ← back to Spark SQL
tap a year for its mentions
the MAD Podcast 11
Latent Space 2
every year every show
the MAD Podcast 11
Latent Space 2
Verbatim, from the transcripts: passages where Spark SQL comes up on the MAD Podcast, Latent Space
The Agent Cloud: Databricks’ Bet on the Future of AI — Matei Zaharia and Reynold Xin
- ▶ 53:10 Reynold Xin Instead of worrying about PostgreSQL and maybe Spark SQL, why not just one? 2 times in the scene
Fireside Chat: Dave Burgess (Head of Data Engineering, Pinterest) w/ Matt Turck (Partner, FirstMark)
- ▶ 8:16 Dave Burgess Uh, we have a query platform, which is Presto and Spark SQL. 2 times in the scene
- ▶ 31:08 Dave Burgess They usually either Spark or Hive or Presto jobs or Spark SQL and, uh, just process the data in every step and, and persist the data back to S three along the way.
Data Observability and Pipelines: OpenLineage and Marquez
- ▶ 25:28 Julien Le Dem Spark, uh, built Spark SQL on top of Parquet, uh, Drill, uh, same thing used Parquet initially, um, you had Hive, Impala,
Optionality in Data Architecture // Justin Borgman, Starburst Data (FirstMark's Data Driven NYC)
- ▶ 18:41 Justin Borgman That's our target, you know, market, and yes, there is Spark SQL, which is a sub-project, and that's the five percent overlap, but I think anyone who's trying to do high concurrency will usually realize that
A Virtual Distributed Storage System // Haoyuan Li, Alluxio (Hosted by FirstMark)
- ▶ 11:22 Haoyuan Li Basically, they run Spark and Spark SQL on top of Aluxio.
A Fireside Chat with MapR CTO M.C. Srivas (Data Driven NYC / FirstMark)
- ▶ 24:40 M.C. Srivas So, and in Spark SQL.
Ion Stoica, Databricks // Creating Apache Spark // Data Driven NYC (FirstMark Capital)
- ▶ 20:58 unnamed speaker But, you know, there's been a lot of investment in optimizing Spark SQL, and, and I think a lot of the use cases that people in practice actually use it for are SQL-like workloads. 4 times in the scene