Sep 30, 2016 · 20m · mad
The Uber Big Data Story // Praveen Murugesan, Uber (Data Driven NYC / FirstMark)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Data Driven NYC presentation, Praveen Murugesan of Uber details the rapid growth and architectural evolution of Uber's big data infrastructure from a basic warehouse to a massive, centralized Hadoop data lake. He highlights internal developer frameworks like Spark UDK and Magellan that enable high-scale geospatial analytics and seamless data access across the global enterprise.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.7% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Praveen politely reframes the host's premise about data lakes, clarifying that Uber did not blindly build a data dump but strategically addressed critical production source query costs.
Hardest push from Matt ▶ 17:02 Host challenges data lake efficacy and orderingMatt Turck points out that data lakes often become unused dumping grounds in Fortune 1000 companies and questions whether Uber built the lake before analytics.
Biggest teaching moment ▶ 12:45 Technical breakdown of spatial join optimizationPraveen educates the audience on the mathematical expense of spatial boundary joins and how Uber engineered automated index UDFs to solve Cartesian explosion.
Matt holds his own ▶ 17:02 Host demonstrates enterprise data architecture contextMatt Turck demonstrates industry expertise by referencing common enterprise data lake failures before asking a targeted architectural question.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Uber's Core Mission and Global Scale | 0 | 2 | 0 | 0 | Praveen delivers a presentation monologue outlining Uber's scale across 75 countries and detailing the distinct data requirements of city operations, data scientists, and engineers. The host does not participate in this presentation segment. | |
| Evolution of Uber Data Architecture (2014 vs Present) | 0 | 4 | 0 | 0 | Praveen explains the architectural shift from legacy Vertica and S3 setups to a Hadoop-based data lake, detailing the transition from ETL to EL and JSON to Parquet/Avro schema management. The host remains silent throughout the monologue. | |
| Deep Dive: Spark UDK (Uber Developer Kit) | 0 | 4 | 0 | 0 | Praveen details the design of the Uber Developer Kit (UDK) for Spark, explaining how SCBuilder and Spark Log abstract SRE constraints to prevent cluster trashing. The host provides no input during this section. | |
| Geospatial Processing and Spatial Join Optimization | 0 | 5 | 0 | 0 | Praveen presents Uber's spatial join optimization, breaking down how Cartesian join complexity is bypassed using generated Hive UDF indexes. The host is inactive during this presentation segment. | |
| Key Takeaways and Talk Conclusion | 0 | 3 | 0 | 0 | Praveen concludes his presentation with key architectural takeaways regarding EL systems, internal frameworks, and open-source adoption. The host does not interject during this concluding slide. |