Jun 12, 2019 · 23m · mad

What's Next for Open-Source Time Series Data? // Paul Dix, Influx Data (FirstMark's Data Driven NYC)

Paul Dix · 19m spoken Matt Turck · 18s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC keynote, InfluxData CTO and Founder Paul Dix explores the unique database requirements of time series data and introduces InfluxDB 2.0 alongside Flux, a novel functional scripting language designed to unify data collection, querying, and stream processing.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1.4% of the talking time here. How this is scored →

Matt as informed peer 0.0 Guest teaching 0.3 Guest disagreement 0.3 Matt pushing back 0.0
05100:0010:0020:000:48–3:06 · Matt as informed peer 0/10 Defining Time Series Data Across Key Use Cases Paul Dix provides an introductory presentation defining time series data and distinguishing between regular metrics and irregular event streams. Because this is a solo keynote segment, all host-side dynamic scores are zero.3:06–6:06 · Matt as informed peer 0/10 Unique Database Workload Characteristics of Time Series Paul explains database workload characteristics unique to time series, such as high write throughput, large range scans, and retention eviction. Host interaction scores remain zero for this monologue segment.6:06–9:30 · Matt as informed peer 0/10 The Origin and Evolution of the TICK Stack Paul traces the origin of InfluxDB and the emergence of Telegraf, Kapacitor, and Chronograf to form the TICK stack. This is a solo technical talk with no host participation.9:30–12:23 · Matt as informed peer 0/10 InfluxDB Line Protocol Schema and Nanosecond Precision Paul details the Line Protocol schema, support for varied field types, and nanosecond timestamp precision. Host interaction scores are zero during this monologue presentation.12:23–14:25 · Matt as informed peer 0/10 The Unified Architecture of InfluxDB 2.0 Paul describes consolidating the separate TICK stack components into a unified InfluxDB 2.0 platform. Host scores remain zero as this is part of the main presentation.14:25–16:52 · Matt as informed peer 0/10 Client Libraries and Modular Visualization Tools Paul covers official client libraries, UI JavaScript components, and the Flux engine pluggable parser architecture. Host scores are zero due to the solo monologue format.16:52–18:57 · Matt as informed peer 0/10 Flux Code Demonstration and Serverless Execution Platform Paul demonstrates Flux syntax, functional pipe operators, and its capability as a serverless execution engine inside the database. Host-side scores are zero for this monologue.18:57–23:24 · Matt as informed peer 0/10 InfluxDB 2.0 Distribution Options and Keynote Conclusion Host Matt Turck opens the floor to audience questions regarding indexing and language design rationale. Paul playfully dismisses Lisp and alternative options while answering the audience, maintaining full conversational authority.0:48–3:06 · Guest teaching 0/10 Defining Time Series Data Across Key Use Cases Paul Dix provides an introductory presentation defining time series data and distinguishing between regular metrics and irregular event streams. Because this is a solo keynote segment, all host-side dynamic scores are zero.3:06–6:06 · Guest teaching 0/10 Unique Database Workload Characteristics of Time Series Paul explains database workload characteristics unique to time series, such as high write throughput, large range scans, and retention eviction. Host interaction scores remain zero for this monologue segment.6:06–9:30 · Guest teaching 0/10 The Origin and Evolution of the TICK Stack Paul traces the origin of InfluxDB and the emergence of Telegraf, Kapacitor, and Chronograf to form the TICK stack. This is a solo technical talk with no host participation.9:30–12:23 · Guest teaching 0/10 InfluxDB Line Protocol Schema and Nanosecond Precision Paul details the Line Protocol schema, support for varied field types, and nanosecond timestamp precision. Host interaction scores are zero during this monologue presentation.12:23–14:25 · Guest teaching 0/10 The Unified Architecture of InfluxDB 2.0 Paul describes consolidating the separate TICK stack components into a unified InfluxDB 2.0 platform. Host scores remain zero as this is part of the main presentation.14:25–16:52 · Guest teaching 0/10 Client Libraries and Modular Visualization Tools Paul covers official client libraries, UI JavaScript components, and the Flux engine pluggable parser architecture. Host scores are zero due to the solo monologue format.16:52–18:57 · Guest teaching 0/10 Flux Code Demonstration and Serverless Execution Platform Paul demonstrates Flux syntax, functional pipe operators, and its capability as a serverless execution engine inside the database. Host-side scores are zero for this monologue.18:57–23:24 · Guest teaching 2/10 InfluxDB 2.0 Distribution Options and Keynote Conclusion Host Matt Turck opens the floor to audience questions regarding indexing and language design rationale. Paul playfully dismisses Lisp and alternative options while answering the audience, maintaining full conversational authority.0:48–3:06 · Guest disagreement 0/10 Defining Time Series Data Across Key Use Cases Paul Dix provides an introductory presentation defining time series data and distinguishing between regular metrics and irregular event streams. Because this is a solo keynote segment, all host-side dynamic scores are zero.3:06–6:06 · Guest disagreement 0/10 Unique Database Workload Characteristics of Time Series Paul explains database workload characteristics unique to time series, such as high write throughput, large range scans, and retention eviction. Host interaction scores remain zero for this monologue segment.6:06–9:30 · Guest disagreement 0/10 The Origin and Evolution of the TICK Stack Paul traces the origin of InfluxDB and the emergence of Telegraf, Kapacitor, and Chronograf to form the TICK stack. This is a solo technical talk with no host participation.9:30–12:23 · Guest disagreement 0/10 InfluxDB Line Protocol Schema and Nanosecond Precision Paul details the Line Protocol schema, support for varied field types, and nanosecond timestamp precision. Host interaction scores are zero during this monologue presentation.12:23–14:25 · Guest disagreement 0/10 The Unified Architecture of InfluxDB 2.0 Paul describes consolidating the separate TICK stack components into a unified InfluxDB 2.0 platform. Host scores remain zero as this is part of the main presentation.14:25–16:52 · Guest disagreement 0/10 Client Libraries and Modular Visualization Tools Paul covers official client libraries, UI JavaScript components, and the Flux engine pluggable parser architecture. Host scores are zero due to the solo monologue format.16:52–18:57 · Guest disagreement 0/10 Flux Code Demonstration and Serverless Execution Platform Paul demonstrates Flux syntax, functional pipe operators, and its capability as a serverless execution engine inside the database. Host-side scores are zero for this monologue.18:57–23:24 · Guest disagreement 2/10 InfluxDB 2.0 Distribution Options and Keynote Conclusion Host Matt Turck opens the floor to audience questions regarding indexing and language design rationale. Paul playfully dismisses Lisp and alternative options while answering the audience, maintaining full conversational authority.0:48–3:06 · Matt pushing back 0/10 Defining Time Series Data Across Key Use Cases Paul Dix provides an introductory presentation defining time series data and distinguishing between regular metrics and irregular event streams. Because this is a solo keynote segment, all host-side dynamic scores are zero.3:06–6:06 · Matt pushing back 0/10 Unique Database Workload Characteristics of Time Series Paul explains database workload characteristics unique to time series, such as high write throughput, large range scans, and retention eviction. Host interaction scores remain zero for this monologue segment.6:06–9:30 · Matt pushing back 0/10 The Origin and Evolution of the TICK Stack Paul traces the origin of InfluxDB and the emergence of Telegraf, Kapacitor, and Chronograf to form the TICK stack. This is a solo technical talk with no host participation.9:30–12:23 · Matt pushing back 0/10 InfluxDB Line Protocol Schema and Nanosecond Precision Paul details the Line Protocol schema, support for varied field types, and nanosecond timestamp precision. Host interaction scores are zero during this monologue presentation.12:23–14:25 · Matt pushing back 0/10 The Unified Architecture of InfluxDB 2.0 Paul describes consolidating the separate TICK stack components into a unified InfluxDB 2.0 platform. Host scores remain zero as this is part of the main presentation.14:25–16:52 · Matt pushing back 0/10 Client Libraries and Modular Visualization Tools Paul covers official client libraries, UI JavaScript components, and the Flux engine pluggable parser architecture. Host scores are zero due to the solo monologue format.16:52–18:57 · Matt pushing back 0/10 Flux Code Demonstration and Serverless Execution Platform Paul demonstrates Flux syntax, functional pipe operators, and its capability as a serverless execution engine inside the database. Host-side scores are zero for this monologue.18:57–23:24 · Matt pushing back 0/10 InfluxDB 2.0 Distribution Options and Keynote Conclusion Host Matt Turck opens the floor to audience questions regarding indexing and language design rationale. Paul playfully dismisses Lisp and alternative options while answering the audience, maintaining full conversational authority.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 7.4% · guest 92.6%18:00 · Matt 7.4% · guest 92.6%21:00 · Matt 6% · guest 94%21:00 · Matt 6% · guest 94%
Sharpest disagreement ▶ 21:50 Dismissing Lisp and JavaScript Subsets

Paul forcefully and humorously dismisses adopting existing languages like Lisp or JavaScript for Flux, arguing that if Paul Graham and Rich Hickey could not make Lisp popular, nobody will.

Hardest push from Matt ▶ 19:38 Host Caps Talk Duration for Q&A

Host Matt Turck steps in to end the formal presentation due to time limits and redirects the remaining time to quick audience questions.

Biggest teaching moment ▶ 20:10 Explaining Dual Inverted Index Architecture

Paul educates an audience member on how InfluxDB combines a columnar data store for time series values with an inverted index mapping metadata tag key-value pairs.

Matt holds his own ▶ 19:38 Host Contextualizes Company Progress

Host Matt Turck highlights the massive progress made by InfluxData since Paul last presented at the Data Driven NYC event.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining Time Series Data Across Key Use Cases 0000 Paul Dix provides an introductory presentation defining time series data and distinguishing between regular metrics and irregular event streams. Because this is a solo keynote segment, all host-side dynamic scores are zero.
Unique Database Workload Characteristics of Time Series 0000 Paul explains database workload characteristics unique to time series, such as high write throughput, large range scans, and retention eviction. Host interaction scores remain zero for this monologue segment.
The Origin and Evolution of the TICK Stack 0000 Paul traces the origin of InfluxDB and the emergence of Telegraf, Kapacitor, and Chronograf to form the TICK stack. This is a solo technical talk with no host participation.
InfluxDB Line Protocol Schema and Nanosecond Precision 0000 Paul details the Line Protocol schema, support for varied field types, and nanosecond timestamp precision. Host interaction scores are zero during this monologue presentation.
The Unified Architecture of InfluxDB 2.0 0000 Paul describes consolidating the separate TICK stack components into a unified InfluxDB 2.0 platform. Host scores remain zero as this is part of the main presentation.
Client Libraries and Modular Visualization Tools 0000 Paul covers official client libraries, UI JavaScript components, and the Flux engine pluggable parser architecture. Host scores are zero due to the solo monologue format.
Flux Code Demonstration and Serverless Execution Platform 0000 Paul demonstrates Flux syntax, functional pipe operators, and its capability as a serverless execution engine inside the database. Host-side scores are zero for this monologue.
InfluxDB 2.0 Distribution Options and Keynote Conclusion 0220 Host Matt Turck opens the floor to audience questions regarding indexing and language design rationale. Paul playfully dismisses Lisp and alternative options while answering the audience, maintaining full conversational authority.

Statements from this episode (10)

Insight
Paul Dix: Regular time series data summarizes irregular event data
“The thing that's interesting about irregular time series data is that you can actually induce a regular time series from irregular time series data. Basically, a regular time series is just a summary of an irregular Series.”
Paul Dix Jun 12, 2019 ▶ 2:37
Insight
Why traditional databases fail at time series workloads
“And actually no Regular database is designed with this kind of workload in mind. Databases, for the most part, assume that you want to keep your data around for a very long time. So, these kind of unique aspects of time series make it kind of a degenerate case…”
Paul Dix Jun 12, 2019 ▶ 5:47
Assertion Not checkable as stated
Paul Dix: Telegraf is likely on tens to hundreds of millions of servers
“This project actually is our most popular open source project by far. We don't have any sort of tracking on it right now, but I would estimate that it's probably deployed on tens of millions of servers across the world, if not hundreds of millions at this poin…”
Paul Dix Jun 12, 2019 ▶ 7:44
Assertion Not checkable as stated
An InfluxData customer deployed Telegraf to 45,000 servers in a single day
“Our, we've had customers, like, deploy it to 45,000 servers in a single day, so.”
Paul Dix Jun 12, 2019 ▶ 8:01
Assertion Not checkable as stated
HFT firms use InfluxDB and atomic clocks for sub-300ns drift
“There are high frequency trading firms that use it to track latencies in their network infrastructure, and they actually have, like, atomic clocks deployed in their data centers, so they guarantee less than 300 nanoseconds of clock drift Worldwide.”
Paul Dix Jun 12, 2019 ▶ 10:42
Disclosure
InfluxDB 2.0 unifies the TICK stack into a single database
“So, my idea was within FluxDB two dot O, we could collapse these things into one cohesive whole, and have one language that ties all of it together.”
Paul Dix Jun 12, 2019 ▶ 12:23
Disclosure
Flux combines a query optimizer, VM, and Turing-complete scripting language
“Flux is basically a new language that we're creating for two dot O. It's a combination of a bunch of things. It's basically a query planner, it's a query optimizer, but it's also a Turing complete scripting language, which includes a virtual machine, And a que…”
Paul Dix Jun 12, 2019 ▶ 15:14
Disclosure
InfluxData is building PromQL support directly into Flux
“And we also are in the middle of building in support for PromQL, the Prometheus query language.”
Paul Dix Jun 12, 2019 ▶ 15:46
Assertion Supported
Flux turns InfluxDB into a serverless execution platform for time series
“It essentially turns the database into a serverless execution platform for time series data. The idea is you can define any sort of custom logic that you want, inject it into the database, and it will periodically run that logic over the data that you're writi…”
Paul Dix Jun 12, 2019 ▶ 17:56
Prediction Not checkable as stated
Paul Dix: Nobody will ever make the Lisp programming language popular
“Lisp, we're not going to do because Paul Graham and Rich Hickey couldn't make Lisp popular, and neither will we. Nobody will.”
Paul Dix Jun 12, 2019 ▶ 22:16
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.