Sep 30, 2016 · 20m · mad

The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)

Praveen Murugesan · 17m spoken Matt Turck · 53s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC presentation, Praveen Murugesan of Uber details the rapid growth and architectural evolution of Uber's big data infrastructure from a basic warehouse to a massive, centralized Hadoop data lake. He highlights internal developer frameworks like Spark UDK and Magellan that enable high-scale geospatial analytics and seamless data access across the global enterprise.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.7% of the talking time here. How this is scored →

Matt as informed peer 0.0 Guest teaching 3.6 Guest disagreement 0.0 Matt pushing back 0.0
05100:0010:0020:000:54–3:18 · Matt as informed peer 0/10 Uber's Core Mission and Global Scale Praveen delivers a presentation monologue outlining Uber's scale across 75 countries and detailing the distinct data requirements of city operations, data scientists, and engineers. The host does not participate in this presentation segment.3:18–8:11 · Matt as informed peer 0/10 Evolution of Uber Data Architecture (2014 vs Present) Praveen explains the architectural shift from legacy Vertica and S3 setups to a Hadoop-based data lake, detailing the transition from ETL to EL and JSON to Parquet/Avro schema management. The host remains silent throughout the monologue.8:11–11:53 · Matt as informed peer 0/10 Deep Dive: Spark UDK (Uber Developer Kit) Praveen details the design of the Uber Developer Kit (UDK) for Spark, explaining how SCBuilder and Spark Log abstract SRE constraints to prevent cluster trashing. The host provides no input during this section.11:53–16:14 · Matt as informed peer 0/10 Geospatial Processing and Spatial Join Optimization Praveen presents Uber's spatial join optimization, breaking down how Cartesian join complexity is bypassed using generated Hive UDF indexes. The host is inactive during this presentation segment.16:14–17:01 · Matt as informed peer 0/10 Key Takeaways and Talk Conclusion Praveen concludes his presentation with key architectural takeaways regarding EL systems, internal frameworks, and open-source adoption. The host does not interject during this concluding slide.0:54–3:18 · Guest teaching 2/10 Uber's Core Mission and Global Scale Praveen delivers a presentation monologue outlining Uber's scale across 75 countries and detailing the distinct data requirements of city operations, data scientists, and engineers. The host does not participate in this presentation segment.3:18–8:11 · Guest teaching 4/10 Evolution of Uber Data Architecture (2014 vs Present) Praveen explains the architectural shift from legacy Vertica and S3 setups to a Hadoop-based data lake, detailing the transition from ETL to EL and JSON to Parquet/Avro schema management. The host remains silent throughout the monologue.8:11–11:53 · Guest teaching 4/10 Deep Dive: Spark UDK (Uber Developer Kit) Praveen details the design of the Uber Developer Kit (UDK) for Spark, explaining how SCBuilder and Spark Log abstract SRE constraints to prevent cluster trashing. The host provides no input during this section.11:53–16:14 · Guest teaching 5/10 Geospatial Processing and Spatial Join Optimization Praveen presents Uber's spatial join optimization, breaking down how Cartesian join complexity is bypassed using generated Hive UDF indexes. The host is inactive during this presentation segment.16:14–17:01 · Guest teaching 3/10 Key Takeaways and Talk Conclusion Praveen concludes his presentation with key architectural takeaways regarding EL systems, internal frameworks, and open-source adoption. The host does not interject during this concluding slide.0:54–3:18 · Guest disagreement 0/10 Uber's Core Mission and Global Scale Praveen delivers a presentation monologue outlining Uber's scale across 75 countries and detailing the distinct data requirements of city operations, data scientists, and engineers. The host does not participate in this presentation segment.3:18–8:11 · Guest disagreement 0/10 Evolution of Uber Data Architecture (2014 vs Present) Praveen explains the architectural shift from legacy Vertica and S3 setups to a Hadoop-based data lake, detailing the transition from ETL to EL and JSON to Parquet/Avro schema management. The host remains silent throughout the monologue.8:11–11:53 · Guest disagreement 0/10 Deep Dive: Spark UDK (Uber Developer Kit) Praveen details the design of the Uber Developer Kit (UDK) for Spark, explaining how SCBuilder and Spark Log abstract SRE constraints to prevent cluster trashing. The host provides no input during this section.11:53–16:14 · Guest disagreement 0/10 Geospatial Processing and Spatial Join Optimization Praveen presents Uber's spatial join optimization, breaking down how Cartesian join complexity is bypassed using generated Hive UDF indexes. The host is inactive during this presentation segment.16:14–17:01 · Guest disagreement 0/10 Key Takeaways and Talk Conclusion Praveen concludes his presentation with key architectural takeaways regarding EL systems, internal frameworks, and open-source adoption. The host does not interject during this concluding slide.0:54–3:18 · Matt pushing back 0/10 Uber's Core Mission and Global Scale Praveen delivers a presentation monologue outlining Uber's scale across 75 countries and detailing the distinct data requirements of city operations, data scientists, and engineers. The host does not participate in this presentation segment.3:18–8:11 · Matt pushing back 0/10 Evolution of Uber Data Architecture (2014 vs Present) Praveen explains the architectural shift from legacy Vertica and S3 setups to a Hadoop-based data lake, detailing the transition from ETL to EL and JSON to Parquet/Avro schema management. The host remains silent throughout the monologue.8:11–11:53 · Matt pushing back 0/10 Deep Dive: Spark UDK (Uber Developer Kit) Praveen details the design of the Uber Developer Kit (UDK) for Spark, explaining how SCBuilder and Spark Log abstract SRE constraints to prevent cluster trashing. The host provides no input during this section.11:53–16:14 · Matt pushing back 0/10 Geospatial Processing and Spatial Join Optimization Praveen presents Uber's spatial join optimization, breaking down how Cartesian join complexity is bypassed using generated Hive UDF indexes. The host is inactive during this presentation segment.16:14–17:01 · Matt pushing back 0/10 Key Takeaways and Talk Conclusion Praveen concludes his presentation with key architectural takeaways regarding EL systems, internal frameworks, and open-source adoption. The host does not interject during this concluding slide.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 20.4% · guest 79.6%15:00 · Matt 20.4% · guest 79.6%18:00 · Matt 17.5% · guest 82.5%18:00 · Matt 17.5% · guest 82.5%
Sharpest disagreement ▶ 17:38 Rejection of naive data lake implementation framing

Praveen politely reframes the host's premise about data lakes, clarifying that Uber did not blindly build a data dump but strategically addressed critical production source query costs.

Hardest push from Matt ▶ 17:02 Host challenges data lake efficacy and ordering

Matt Turck points out that data lakes often become unused dumping grounds in Fortune 1000 companies and questions whether Uber built the lake before analytics.

Biggest teaching moment ▶ 12:45 Technical breakdown of spatial join optimization

Praveen educates the audience on the mathematical expense of spatial boundary joins and how Uber engineered automated index UDFs to solve Cartesian explosion.

Matt holds his own ▶ 17:02 Host demonstrates enterprise data architecture context

Matt Turck demonstrates industry expertise by referencing common enterprise data lake failures before asking a targeted architectural question.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Uber's Core Mission and Global Scale 0200 Praveen delivers a presentation monologue outlining Uber's scale across 75 countries and detailing the distinct data requirements of city operations, data scientists, and engineers. The host does not participate in this presentation segment.
Evolution of Uber Data Architecture (2014 vs Present) 0400 Praveen explains the architectural shift from legacy Vertica and S3 setups to a Hadoop-based data lake, detailing the transition from ETL to EL and JSON to Parquet/Avro schema management. The host remains silent throughout the monologue.
Deep Dive: Spark UDK (Uber Developer Kit) 0400 Praveen details the design of the Uber Developer Kit (UDK) for Spark, explaining how SCBuilder and Spark Log abstract SRE constraints to prevent cluster trashing. The host provides no input during this section.
Geospatial Processing and Spatial Join Optimization 0500 Praveen presents Uber's spatial join optimization, breaking down how Cartesian join complexity is bypassed using generated Hive UDF indexes. The host is inactive during this presentation segment.
Key Takeaways and Talk Conclusion 0300 Praveen concludes his presentation with key architectural takeaways regarding EL systems, internal frameworks, and open-source adoption. The host does not interject during this concluding slide.

Statements from this episode (9)

Assertion Not checkable as stated
Murugesan: Uber practically had no data infrastructure when he joined in 2014
“We practically did not have an infrastructure is what the honest reality is.”
Praveen Murugesan Sep 30, 2016 ▶ 0:20
Assertion Not checkable as stated
Murugesan: City operations personnel constitute most of Uber's data consumers
“Most of Uber's data consumers are actually these operations people.”
Praveen Murugesan Sep 30, 2016 ▶ 2:09
Assertion Not checkable as stated
Uber consolidated all log and business data into an HDFS data lake
“What we really created was, like, using HDFS, like, a data lake, where we basically copied the whole data sets from, like whatever we get from, like, analytical logs or, like, all our business data sources, too, into HDFS.”
Praveen Murugesan Sep 30, 2016 ▶ 4:28
Assertion Not checkable as stated
Streamific powers all data ingestion and aggregation at Uber
“There's a system called Streamific, which powers all of the data ingestion aggregation at this point.”
Praveen Murugesan Sep 30, 2016 ▶ 5:37
Assertion Not checkable as stated
Uber transitioned from ETL into Vertica to EL into Hadoop
“We went from an ETL model, where we scraped from, like, the original source, transformed the data and loaded to Vertica, to, like, just an EL model, where we just, like, just copy the data as soon as possible into, like, Hadoop, and all the transformation can …”
Praveen Murugesan Sep 30, 2016 ▶ 6:19
Insight
Schemaless JSON fails as companies scale beyond 50 employees
“It works well, like, if you're, like, a ten-person company, or, like, even to a fifty-person company, when you're, like, scaling to, like, thousands of people, like, you actually need a proper negotiation in between.”
Praveen Murugesan Sep 30, 2016 ▶ 6:56
Assertion Not checkable as stated
Uber engineers frequently crashed Kafka clusters with unthrottled Spark executor writes
“Kafka was, in general, like, a nice way where people used to pipe the results of, like, their Spark jobs. But often cases, what they do is, like, they hit Kafka hard and bring Kafka down because they're trying to, like, actually send data from, like, hundred e…”
Praveen Murugesan Sep 30, 2016 ▶ 10:21
Assertion Not checkable as stated
Uber deployed trip similarity algorithms to fight incentive fraud in China
“Uber used to run this incentive program in China, and, ah, people are trying to game it, so they used to always have, like, similar trips simulated From their, ah, various different devices. So what we did was actually, like, try to find an algorithm where we …”
Praveen Murugesan Sep 30, 2016 ▶ 12:31
Disclosure
Uber uses cost accounting chargebacks to track data storage costs by unit
“We have something called cost accounting chargebacks is what we call it. So we actually try to, we have a lineage model where we try to figure out who's storing the data. And, ah, we actually kind of can go get a dollar amount of what we are storing.”
Praveen Murugesan Sep 30, 2016 ▶ 18:52
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.