Aug 11, 2022 · 20m · mad

Building Real-Time Data Pipelines | Estuary's Johnny Graettinger

Johnny Graettinger · 17m spoken Matt Turck · 8s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC talk, Estuary CTO and co-founder Johnny Graettinger presents Estuary Flow, a real-time data ops platform designed to seamlessly capture, transform, and materialize database changes into downstream applications without degrading production system performance.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 0.8% of the talking time here. How this is scored →

Matt as informed peer 0.8 Guest teaching 4.0 Guest disagreement 0.3 Matt pushing back 0.5
05100:0010:0020:000:08–2:42 · Matt as informed peer 0/10 The E-Commerce Logistics Problem Johnny Graettinger sets up an e-commerce logistics problem and introduces Estuary Flow's real-time data integration architecture. Because this segment is entirely a guest monologue and presentation, the host metrics remain at zero.2:42–5:29 · Matt as informed peer 0/10 Estuary Flow Architecture: Captures, Derivations, and Materializations Johnny demonstrates capturing Postgres database changes and materializing them live into Google Sheets, executing a real-time SQL update. As the host does not speak during this demo section, host scores are strictly zero.5:29–9:41 · Matt as informed peer 0/10 Live Demo: Real-Time Spreadsheet Replication and SQL Execution Johnny explains how derivations incrementally roll up 22 million database rows into Google Sheets without overloading performance, driving live spreadsheet chart updates. The host remains silent throughout this presentation segment.9:41–20:06 · Matt as informed peer 3/10 Collection Reusability and Presentation Wrap-Up Matt Turck enters the discussion to question Johnny about destination endpoints and target customer personas across engineering and business teams. Audience members and Matt receive detailed explanations on RocksDB stateful joins, shard auto-scaling, and usage pricing.0:08–2:42 · Guest teaching 3/10 The E-Commerce Logistics Problem Johnny Graettinger sets up an e-commerce logistics problem and introduces Estuary Flow's real-time data integration architecture. Because this segment is entirely a guest monologue and presentation, the host metrics remain at zero.2:42–5:29 · Guest teaching 4/10 Estuary Flow Architecture: Captures, Derivations, and Materializations Johnny demonstrates capturing Postgres database changes and materializing them live into Google Sheets, executing a real-time SQL update. As the host does not speak during this demo section, host scores are strictly zero.5:29–9:41 · Guest teaching 4/10 Live Demo: Real-Time Spreadsheet Replication and SQL Execution Johnny explains how derivations incrementally roll up 22 million database rows into Google Sheets without overloading performance, driving live spreadsheet chart updates. The host remains silent throughout this presentation segment.9:41–20:06 · Guest teaching 5/10 Collection Reusability and Presentation Wrap-Up Matt Turck enters the discussion to question Johnny about destination endpoints and target customer personas across engineering and business teams. Audience members and Matt receive detailed explanations on RocksDB stateful joins, shard auto-scaling, and usage pricing.0:08–2:42 · Guest disagreement 0/10 The E-Commerce Logistics Problem Johnny Graettinger sets up an e-commerce logistics problem and introduces Estuary Flow's real-time data integration architecture. Because this segment is entirely a guest monologue and presentation, the host metrics remain at zero.2:42–5:29 · Guest disagreement 0/10 Estuary Flow Architecture: Captures, Derivations, and Materializations Johnny demonstrates capturing Postgres database changes and materializing them live into Google Sheets, executing a real-time SQL update. As the host does not speak during this demo section, host scores are strictly zero.5:29–9:41 · Guest disagreement 0/10 Live Demo: Real-Time Spreadsheet Replication and SQL Execution Johnny explains how derivations incrementally roll up 22 million database rows into Google Sheets without overloading performance, driving live spreadsheet chart updates. The host remains silent throughout this presentation segment.9:41–20:06 · Guest disagreement 1/10 Collection Reusability and Presentation Wrap-Up Matt Turck enters the discussion to question Johnny about destination endpoints and target customer personas across engineering and business teams. Audience members and Matt receive detailed explanations on RocksDB stateful joins, shard auto-scaling, and usage pricing.0:08–2:42 · Matt pushing back 0/10 The E-Commerce Logistics Problem Johnny Graettinger sets up an e-commerce logistics problem and introduces Estuary Flow's real-time data integration architecture. Because this segment is entirely a guest monologue and presentation, the host metrics remain at zero.2:42–5:29 · Matt pushing back 0/10 Estuary Flow Architecture: Captures, Derivations, and Materializations Johnny demonstrates capturing Postgres database changes and materializing them live into Google Sheets, executing a real-time SQL update. As the host does not speak during this demo section, host scores are strictly zero.5:29–9:41 · Matt pushing back 0/10 Live Demo: Real-Time Spreadsheet Replication and SQL Execution Johnny explains how derivations incrementally roll up 22 million database rows into Google Sheets without overloading performance, driving live spreadsheet chart updates. The host remains silent throughout this presentation segment.9:41–20:06 · Matt pushing back 2/10 Collection Reusability and Presentation Wrap-Up Matt Turck enters the discussion to question Johnny about destination endpoints and target customer personas across engineering and business teams. Audience members and Matt receive detailed explanations on RocksDB stateful joins, shard auto-scaling, and usage pricing.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 5% · guest 95%12:00 · Matt 5% · guest 95%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%
Sharpest disagreement ▶ 15:30 Reframing the go-to-market question

Johnny gently reframes Matt's question about target buyer personas by comparing Estuary Flow to a broad-use database rather than a narrow point solution.

Hardest push from Matt ▶ 14:30 Challenging customer persona alignment

Matt Turck presses Johnny on who Estuary will actually sell to, noting the tension between engineering personnel and spreadsheet users.

Biggest teaching moment ▶ 17:42 Explaining horizontally scalable RocksDB shards

Johnny educates the audience on how Flow handles stateful joins across massive datasets without running out of memory by utilizing RocksDB and recursively splitting shards online.

Matt holds his own ▶ 14:30 Probing target buyer personas

Matt Turck displays informed product insight by identifying that Estuary bridges two vastly different user groups—backend infrastructure engineers and spreadsheet users.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
The E-Commerce Logistics Problem 0300 Johnny Graettinger sets up an e-commerce logistics problem and introduces Estuary Flow's real-time data integration architecture. Because this segment is entirely a guest monologue and presentation, the host metrics remain at zero.
Estuary Flow Architecture: Captures, Derivations, and Materializations 0400 Johnny demonstrates capturing Postgres database changes and materializing them live into Google Sheets, executing a real-time SQL update. As the host does not speak during this demo section, host scores are strictly zero.
Live Demo: Real-Time Spreadsheet Replication and SQL Execution 0400 Johnny explains how derivations incrementally roll up 22 million database rows into Google Sheets without overloading performance, driving live spreadsheet chart updates. The host remains silent throughout this presentation segment.
Collection Reusability and Presentation Wrap-Up 3512 Matt Turck enters the discussion to question Johnny about destination endpoints and target customer personas across engineering and business teams. Audience members and Matt receive detailed explanations on RocksDB stateful joins, shard auto-scaling, and usage pricing.

Statements from this episode (6)

Disclosure
Graettinger: Estuary Flow is not a database and stores no data copies
“But a key thing to understand about Flow is that it is not a database. It is not storing copies of your data. It is really concerned just with coordinating your existing data systems and how data is moving between them.”
Johnny Graettinger Aug 11, 2022 ▶ 2:22
Assertion Supported
Graettinger: Estuary Flow processes data incrementally by tracking changes
“Flo's really concerned with just the movement of data through this topology through the data systems that you have, and it's doing it in a very incremental way. So it's looking just at what's changing within any given collection and propagating the effects.”
Johnny Graettinger Aug 11, 2022 ▶ 3:29
Assertion Supported
Estuary's materialization connectors are implemented as Docker container plugins
“Connectors for materializations are essentially Docker container plugins.”
Johnny Graettinger Aug 11, 2022 ▶ 4:24
Assertion Supported
Graettinger: Google Sheets rate limits updates faster than once per second
“That's actually a limitation of Google Sheets. If you try and make updates to Google Sheets any faster than that, Google starts to rate limit you.”
Johnny Graettinger Aug 11, 2022 ▶ 8:19
Assertion Not checkable as stated
Graettinger: Estuary Flow collections enable reuse without impacting source databases
“So these collections are, can, you know, can be repeatedly used for different downstream kind of data processing tasks without ever impacting or touching that database again.”
Johnny Graettinger Aug 11, 2022 ▶ 10:52
Disclosure
Graettinger: Estuary charges customers based on compute resource usage
“We're charging on usage, essentially. So we're charging on the actual, like, compute resource that's required in order to get these workflows up and running.”
Johnny Graettinger Aug 11, 2022 ▶ 19:04
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.