Feb 25, 2019 · 22m · mad

Dynamic Range Sharding with Spanner // Daniel Chia, Google Spanner (FirstMark's Data Driven NYC)

Daniel Chia · 17m spoken Matt Turck · 28s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this FirstMark Data Driven NYC presentation, Google Spanner engineer Daniel Chia explains the internal architecture and dynamic sharding mechanisms behind Google's globally distributed relational database. He details how Spanner overcomes traditional static hashing limitations through dynamic range sharding, adaptive load splitting, and hardware-backed time synchronization to achieve seamless horizontal scale and high availability.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 2.2% of the talking time here. How this is scored →

Matt as informed peer 0.3 Guest teaching 0.7 Guest disagreement 0.2 Matt pushing back 0.3
05100:0010:0020:001:17–3:28 · Matt as informed peer 0/10 What is Google Spanner? This is a solo presentation segment by guest Daniel Chia explaining Google Spanner's core features like TrueTime and external consistency. Because the host does not speak, all host-related scores are set to zero.3:28–6:23 · Matt as informed peer 0/10 The Challenge of Data Scale Daniel Chia presents a monologue on the challenges of database scaling and the engineering trade-offs between shard balance and query locality. The host is inactive during this segment.6:23–8:31 · Matt as informed peer 0/10 Algorithmic Sharding on Full Primary Key Daniel Chia breaks down algorithmic sharding schemes and their failure modes during presentation. Host scores remain zero as the host does not speak.8:31–11:53 · Matt as informed peer 0/10 Spanner's Approach: Dynamic Range Sharding Guest continues his solo presentation detailing dynamic range sharding and the location service inside Spanner. Host does not participate.11:53–15:20 · Matt as informed peer 0/10 Dynamic Clean-Up: Adaptive Merging Daniel Chia concludes his slide talk by outlining the coprocessor framework and edge cases such as unsplittable workloads. Host scores are zero due to zero host involvement.15:20–22:49 · Matt as informed peer 2/10 Audience Q&A and Conclusion Matt Turck opens the Q&A session with introductory questions on Spanner's internal usage and product history. Daniel Chia politely corrects Matt's assumption that Spanner was open sourced, clarifying that it was made available as Cloud Spanner.1:17–3:28 · Guest teaching 0/10 What is Google Spanner? This is a solo presentation segment by guest Daniel Chia explaining Google Spanner's core features like TrueTime and external consistency. Because the host does not speak, all host-related scores are set to zero.3:28–6:23 · Guest teaching 0/10 The Challenge of Data Scale Daniel Chia presents a monologue on the challenges of database scaling and the engineering trade-offs between shard balance and query locality. The host is inactive during this segment.6:23–8:31 · Guest teaching 0/10 Algorithmic Sharding on Full Primary Key Daniel Chia breaks down algorithmic sharding schemes and their failure modes during presentation. Host scores remain zero as the host does not speak.8:31–11:53 · Guest teaching 0/10 Spanner's Approach: Dynamic Range Sharding Guest continues his solo presentation detailing dynamic range sharding and the location service inside Spanner. Host does not participate.11:53–15:20 · Guest teaching 0/10 Dynamic Clean-Up: Adaptive Merging Daniel Chia concludes his slide talk by outlining the coprocessor framework and edge cases such as unsplittable workloads. Host scores are zero due to zero host involvement.15:20–22:49 · Guest teaching 4/10 Audience Q&A and Conclusion Matt Turck opens the Q&A session with introductory questions on Spanner's internal usage and product history. Daniel Chia politely corrects Matt's assumption that Spanner was open sourced, clarifying that it was made available as Cloud Spanner.1:17–3:28 · Guest disagreement 0/10 What is Google Spanner? This is a solo presentation segment by guest Daniel Chia explaining Google Spanner's core features like TrueTime and external consistency. Because the host does not speak, all host-related scores are set to zero.3:28–6:23 · Guest disagreement 0/10 The Challenge of Data Scale Daniel Chia presents a monologue on the challenges of database scaling and the engineering trade-offs between shard balance and query locality. The host is inactive during this segment.6:23–8:31 · Guest disagreement 0/10 Algorithmic Sharding on Full Primary Key Daniel Chia breaks down algorithmic sharding schemes and their failure modes during presentation. Host scores remain zero as the host does not speak.8:31–11:53 · Guest disagreement 0/10 Spanner's Approach: Dynamic Range Sharding Guest continues his solo presentation detailing dynamic range sharding and the location service inside Spanner. Host does not participate.11:53–15:20 · Guest disagreement 0/10 Dynamic Clean-Up: Adaptive Merging Daniel Chia concludes his slide talk by outlining the coprocessor framework and edge cases such as unsplittable workloads. Host scores are zero due to zero host involvement.15:20–22:49 · Guest disagreement 1/10 Audience Q&A and Conclusion Matt Turck opens the Q&A session with introductory questions on Spanner's internal usage and product history. Daniel Chia politely corrects Matt's assumption that Spanner was open sourced, clarifying that it was made available as Cloud Spanner.1:17–3:28 · Matt pushing back 0/10 What is Google Spanner? This is a solo presentation segment by guest Daniel Chia explaining Google Spanner's core features like TrueTime and external consistency. Because the host does not speak, all host-related scores are set to zero.3:28–6:23 · Matt pushing back 0/10 The Challenge of Data Scale Daniel Chia presents a monologue on the challenges of database scaling and the engineering trade-offs between shard balance and query locality. The host is inactive during this segment.6:23–8:31 · Matt pushing back 0/10 Algorithmic Sharding on Full Primary Key Daniel Chia breaks down algorithmic sharding schemes and their failure modes during presentation. Host scores remain zero as the host does not speak.8:31–11:53 · Matt pushing back 0/10 Spanner's Approach: Dynamic Range Sharding Guest continues his solo presentation detailing dynamic range sharding and the location service inside Spanner. Host does not participate.11:53–15:20 · Matt pushing back 0/10 Dynamic Clean-Up: Adaptive Merging Daniel Chia concludes his slide talk by outlining the coprocessor framework and edge cases such as unsplittable workloads. Host scores are zero due to zero host involvement.15:20–22:49 · Matt pushing back 2/10 Audience Q&A and Conclusion Matt Turck opens the Q&A session with introductory questions on Spanner's internal usage and product history. Daniel Chia politely corrects Matt's assumption that Spanner was open sourced, clarifying that it was made available as Cloud Spanner.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 16.6% · guest 83.4%15:00 · Matt 16.6% · guest 83.4%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 2.2% · guest 97.8%21:00 · Matt 2.2% · guest 97.8%
Sharpest disagreement ▶ 22:16 Gentle rejection of competitor comparison

Daniel Chia politely reframes an audience question regarding MapRDB and Apache Drill, rejecting the assertion that they are direct competitors by emphasizing specific workload requirements.

Hardest push from Matt ▶ 16:51 Matt Turck probing Spanner's open source trajectory

Matt Turck presses Daniel Chia on Spanner's product evolution, asking if the system was open sourced before becoming a commercial cloud service.

Biggest teaching moment ▶ 16:56 Correcting the open-source misconception

Daniel Chia directly clarifies to host Matt Turck that Spanner was never open sourced, distinguishing internal deployment from Cloud Spanner.

Matt holds his own ▶ 16:03 Matt Turck demonstrating technical context

Matt Turck displays background knowledge regarding Spanner's history at Google, framing the Q&A around its long operational history.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
What is Google Spanner? 0000 This is a solo presentation segment by guest Daniel Chia explaining Google Spanner's core features like TrueTime and external consistency. Because the host does not speak, all host-related scores are set to zero.
The Challenge of Data Scale 0000 Daniel Chia presents a monologue on the challenges of database scaling and the engineering trade-offs between shard balance and query locality. The host is inactive during this segment.
Algorithmic Sharding on Full Primary Key 0000 Daniel Chia breaks down algorithmic sharding schemes and their failure modes during presentation. Host scores remain zero as the host does not speak.
Spanner's Approach: Dynamic Range Sharding 0000 Guest continues his solo presentation detailing dynamic range sharding and the location service inside Spanner. Host does not participate.
Dynamic Clean-Up: Adaptive Merging 0000 Daniel Chia concludes his slide talk by outlining the coprocessor framework and edge cases such as unsplittable workloads. Host scores are zero due to zero host involvement.
Audience Q&A and Conclusion 2412 Matt Turck opens the Q&A session with introductory questions on Spanner's internal usage and product history. Daniel Chia politely corrects Matt's assumption that Spanner was open sourced, clarifying that it was made available as Cloud Spanner.

Statements from this episode (15)

Assertion Supported
Google Spanner provides ACID transactions and SQL semantics
“Well, Spanner is Google's relational database. It has asset semantics and SQL, sorry, asset transactions and SQL semantics.”
Daniel Chia Feb 25, 2019 ▶ 1:29
Assertion Supported
Google Spanner handles petabytes of data and tens of millions of QPS
“The mission critical database for Google, it supports hundreds of Google applications, including Google AdWords and Google Play, and we manage, you know, petabytes of data and have tens of millions of QPS every day.”
Daniel Chia Feb 25, 2019 ▶ 1:39
Assertion Supported
Google Spanner synchronously replicates data across data centers using Paxos
“What's special about Spanner is that it's able to synchronously replicate data between different data centers using algorithms such as Paxos and two-phase commit.”
Daniel Chia Feb 25, 2019 ▶ 1:55
Assertion Partly supported
Chia: Google Spanner uses TrueTime technology to assign global timestamps
“We use this technology called TrueTime, which allows us to get an estimate of what is the exact accurate time anywhere in the world to assign a unique timestamp to every transaction that Spanner processes.”
Daniel Chia Feb 25, 2019 ▶ 2:17
Assertion Supported
Chia: Spanner provides external consistency for concurrent transactions
“This allows us to offer developers a consistency model called external consistency, where even though Spanner might be processing hundreds of transactions concurrently they all appear to have committed In serial order according to their timestamp.”
Daniel Chia Feb 25, 2019 ▶ 2:31
Insight
Chia: Shard granularity is exceptionally difficult to alter post-implementation
“Once you shard by a certain level of granularity, it's very hard to redesign your system to try and think about it a different way.”
Daniel Chia Feb 25, 2019 ▶ 4:59
Assertion Contradicted
Chia: Google Spanner eliminates operational differences between and within shards
“And in Spanner, it's a little bit nicer in that every shard behaves the same, and because we have technologies like, ah, two-phase commit, We are, you don't really have any difference between operating within a shard or across shards, but it's not always true …”
Daniel Chia Feb 25, 2019 ▶ 5:06
Insight
Algorithmic sharding on full primary keys degrades database query locality
“With every shard being on a side lead on a different row, sorry, every row being on a different shard, you end up needing to check all shards to answer this query, which kind of defies this prior, this criteria of locality that I talked about earlier.”
Daniel Chia Feb 25, 2019 ▶ 7:19
Insight
Daniel Chia: Sharding database tables purely by entity keys risks giant hot shards
“The problem here is that, you know, one day you're going to get some singer that's so popular and so prolific and has so many songs that That one singer shard is gonna be huge, and then you are kind of in trouble, right? You either need to do a lot of special …”
Daniel Chia Feb 25, 2019 ▶ 8:06
Insight
Chia: Location services allow dynamic database sharding over hardcoding
“The really nice thing about having a location service is that you can adaptively change your shards rather than having to hard code it upfront.”
Daniel Chia Feb 25, 2019 ▶ 10:08
Assertion Supported
Chia: Google Spanner automatically merges shards when read traffic normalizes
“When the reload goes back to normal, We can then eventually remove these shard points that we added earlier, and this is a natural way of the system cleaning up itself so that we then shrink back the number of shards to reduce both unnecessary shards and also …”
Daniel Chia Feb 25, 2019 ▶ 11:53
Assertion Supported
Chia: All Spanner RPCs address specific data rather than machine names
“All RPCs in Spanner are addressed to specific pieces of data rather than the machine name.”
Daniel Chia Feb 25, 2019 ▶ 13:07
Insight
Daniel Chia: Dynamic sharding systems struggle to quickly adapt to sudden workload shifts
“One of the biggest problems you can face with a dynamic sharding system is that it takes a while for us to adapt to your workload.”
Daniel Chia Feb 25, 2019 ▶ 13:46
Insight
Daniel Chia: Row-based sharding can outperform dynamic range sharding for uniformly random workloads
“In fact, this is a case where one of the earlier methods of sharding just by row would potentially work better, because if your incoming new workload is just randomly Distributed across rows perfectly. It will distribute across all the shards very well.”
Daniel Chia Feb 25, 2019 ▶ 14:23
Assertion Supported
Google Spanner provides full SQL capabilities and transaction guarantees across shards
“Between shards, You still have full transaction guarantees and full SQL capabilities. In a lot of SQL sharding schemes, you, within charts, you have full SQL capabilities, but across charts, there's no such thing.”
Daniel Chia Feb 25, 2019 ▶ 21:00
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.