Apr 2, 2015 · 22m · mad

Paul Dix, InfluxDB // Open-Source Time Series Database // Data Driven NYC (FirstMark Capital)

Paul Dix · 16m spoken Matt Turck · 31s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At a Data Driven NYC event, InfluxDB creator and CEO Paul Dix presents the open-source time series database, detailing its architecture, query capabilities, and real-world applications before engaging in a fireside chat and Q&A on the company's origin and technical engineering choices.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3% of the talking time here. How this is scored →

Matt as informed peer 0.4 Guest teaching 4.8 Guest disagreement 0.6 Matt pushing back 0.2
05100:0010:0020:000:22–7:35 · Matt as informed peer 0/10 Use Cases for Time Series Data Paul Dix presents a solo slide deck outlining why traditional SQL and Cassandra databases fail to scale for time series data. The host does not speak during this segment, requiring zero host scores. Dix thoroughly educates the room on wide-row Cassandra pitfalls and data retention challenges.7:35–10:01 · Matt as informed peer 0/10 InfluxDB Architecture, Data Model, and Write API Dix walks through InfluxDB's core data model, schema-on-the-fly approach, and HTTP write API in a standard monologue. The host remains silent throughout the technical explanation. The segment dynamic is purely instructional with no confrontation.10:01–13:26 · Matt as informed peer 0/10 InfluxDB Query Language and Tag Discovery Dix demonstrates InfluxDB's SQL-like query language and internal inverted index for tag discovery. The host is non-existent in this segment as it is part of the ongoing technical presentation. The content continues to be educational regarding database architecture.13:26–16:33 · Matt as informed peer 1/10 Retention Policies, Continuous Queries, and Grafana Integration Dix finishes his deck right on time, and host Matt Turck steps in to compliment his timing and pivot to company history and background. Turck offers light host commentary without challenging any technical points. Dix collaboratively shares the founding story and YC journey.16:33–22:27 · Matt as informed peer 1/10 Audience Q&A and Event Conclusion Audience members ask about record compression, financial tech use cases, and comparisons to KDB. Dix gently challenges an audience premise regarding KDB's open-source status, pointing out it is proprietary. Host Matt Turck acts strictly as a Q&A facilitator.0:22–7:35 · Guest teaching 6/10 Use Cases for Time Series Data Paul Dix presents a solo slide deck outlining why traditional SQL and Cassandra databases fail to scale for time series data. The host does not speak during this segment, requiring zero host scores. Dix thoroughly educates the room on wide-row Cassandra pitfalls and data retention challenges.7:35–10:01 · Guest teaching 5/10 InfluxDB Architecture, Data Model, and Write API Dix walks through InfluxDB's core data model, schema-on-the-fly approach, and HTTP write API in a standard monologue. The host remains silent throughout the technical explanation. The segment dynamic is purely instructional with no confrontation.10:01–13:26 · Guest teaching 5/10 InfluxDB Query Language and Tag Discovery Dix demonstrates InfluxDB's SQL-like query language and internal inverted index for tag discovery. The host is non-existent in this segment as it is part of the ongoing technical presentation. The content continues to be educational regarding database architecture.13:26–16:33 · Guest teaching 3/10 Retention Policies, Continuous Queries, and Grafana Integration Dix finishes his deck right on time, and host Matt Turck steps in to compliment his timing and pivot to company history and background. Turck offers light host commentary without challenging any technical points. Dix collaboratively shares the founding story and YC journey.16:33–22:27 · Guest teaching 5/10 Audience Q&A and Event Conclusion Audience members ask about record compression, financial tech use cases, and comparisons to KDB. Dix gently challenges an audience premise regarding KDB's open-source status, pointing out it is proprietary. Host Matt Turck acts strictly as a Q&A facilitator.0:22–7:35 · Guest disagreement 1/10 Use Cases for Time Series Data Paul Dix presents a solo slide deck outlining why traditional SQL and Cassandra databases fail to scale for time series data. The host does not speak during this segment, requiring zero host scores. Dix thoroughly educates the room on wide-row Cassandra pitfalls and data retention challenges.7:35–10:01 · Guest disagreement 0/10 InfluxDB Architecture, Data Model, and Write API Dix walks through InfluxDB's core data model, schema-on-the-fly approach, and HTTP write API in a standard monologue. The host remains silent throughout the technical explanation. The segment dynamic is purely instructional with no confrontation.10:01–13:26 · Guest disagreement 0/10 InfluxDB Query Language and Tag Discovery Dix demonstrates InfluxDB's SQL-like query language and internal inverted index for tag discovery. The host is non-existent in this segment as it is part of the ongoing technical presentation. The content continues to be educational regarding database architecture.13:26–16:33 · Guest disagreement 0/10 Retention Policies, Continuous Queries, and Grafana Integration Dix finishes his deck right on time, and host Matt Turck steps in to compliment his timing and pivot to company history and background. Turck offers light host commentary without challenging any technical points. Dix collaboratively shares the founding story and YC journey.16:33–22:27 · Guest disagreement 2/10 Audience Q&A and Event Conclusion Audience members ask about record compression, financial tech use cases, and comparisons to KDB. Dix gently challenges an audience premise regarding KDB's open-source status, pointing out it is proprietary. Host Matt Turck acts strictly as a Q&A facilitator.0:22–7:35 · Matt pushing back 0/10 Use Cases for Time Series Data Paul Dix presents a solo slide deck outlining why traditional SQL and Cassandra databases fail to scale for time series data. The host does not speak during this segment, requiring zero host scores. Dix thoroughly educates the room on wide-row Cassandra pitfalls and data retention challenges.7:35–10:01 · Matt pushing back 0/10 InfluxDB Architecture, Data Model, and Write API Dix walks through InfluxDB's core data model, schema-on-the-fly approach, and HTTP write API in a standard monologue. The host remains silent throughout the technical explanation. The segment dynamic is purely instructional with no confrontation.10:01–13:26 · Matt pushing back 0/10 InfluxDB Query Language and Tag Discovery Dix demonstrates InfluxDB's SQL-like query language and internal inverted index for tag discovery. The host is non-existent in this segment as it is part of the ongoing technical presentation. The content continues to be educational regarding database architecture.13:26–16:33 · Matt pushing back 1/10 Retention Policies, Continuous Queries, and Grafana Integration Dix finishes his deck right on time, and host Matt Turck steps in to compliment his timing and pivot to company history and background. Turck offers light host commentary without challenging any technical points. Dix collaboratively shares the founding story and YC journey.16:33–22:27 · Matt pushing back 0/10 Audience Q&A and Event Conclusion Audience members ask about record compression, financial tech use cases, and comparisons to KDB. Dix gently challenges an audience premise regarding KDB's open-source status, pointing out it is proprietary. Host Matt Turck acts strictly as a Q&A facilitator.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 6.9% · guest 93.1%12:00 · Matt 6.9% · guest 93.1%15:00 · Matt 9.9% · guest 90.1%15:00 · Matt 9.9% · guest 90.1%18:00 · Matt 3.8% · guest 96.2%18:00 · Matt 3.8% · guest 96.2%21:00 · Matt 5.3% · guest 94.7%21:00 · Matt 5.3% · guest 94.7%
Sharpest disagreement ▶ 21:07 Premise challenge on KDB database licensing

When an audience member claims KDB is an open source database, Dix immediately interrupts and corrects the premise, pointing out KDB is not open source.

Hardest push from Matt ▶ 14:47 Host transitions from pitch to origin story

Matt Turck steps in right at the end of the presentation to steer the conversation away from technical slides toward Dix's personal background and startup origin.

Biggest teaching moment ▶ 3:30 Deep dive into distributed database bottlenecks

Dix details how distributed databases like Cassandra force developers to write application-level routing logic and create hot spots across cluster replication groups.

Matt holds his own ▶ 14:47 Host asserts control of session structure

Host Matt Turck praises Dix's precise timing and takes charge of steering the interview phase before opening the floor to audience questions.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Use Cases for Time Series Data 0610 Paul Dix presents a solo slide deck outlining why traditional SQL and Cassandra databases fail to scale for time series data. The host does not speak during this segment, requiring zero host scores. Dix thoroughly educates the room on wide-row Cassandra pitfalls and data retention challenges.
InfluxDB Architecture, Data Model, and Write API 0500 Dix walks through InfluxDB's core data model, schema-on-the-fly approach, and HTTP write API in a standard monologue. The host remains silent throughout the technical explanation. The segment dynamic is purely instructional with no confrontation.
InfluxDB Query Language and Tag Discovery 0500 Dix demonstrates InfluxDB's SQL-like query language and internal inverted index for tag discovery. The host is non-existent in this segment as it is part of the ongoing technical presentation. The content continues to be educational regarding database architecture.
Retention Policies, Continuous Queries, and Grafana Integration 1301 Dix finishes his deck right on time, and host Matt Turck steps in to compliment his timing and pivot to company history and background. Turck offers light host commentary without challenging any technical points. Dix collaboratively shares the founding story and YC journey.
Audience Q&A and Event Conclusion 1520 Audience members ask about record compression, financial tech use cases, and comparisons to KDB. Dix gently challenges an audience premise regarding KDB's open-source status, pointing out it is proprietary. Host Matt Turck acts strictly as a Q&A facilitator.

Statements from this episode (11)

Disclosure
Dix: InfluxDB is targeting IoT consumer and industrial sensor data
“And then the last one that we're really targeting is sensor data. So this is IOT, both consumer and industrial, right? You're thinking power generation, oil and gas wells, and then on the consumer side, you know, fitness trackers to smart home stuff, all that …”
Paul Dix Apr 2, 2015 ▶ 1:34
Insight
DevOps metrics and IoT sensor data spaces look surprisingly similar
“And those, the sensor data and the DevOps data spaces, they look surprisingly similar. Because when you think about DevOps data, and you think about the metrics that you're collecting, the sensors are just software that you have on your servers. Whereas in sen…”
Paul Dix Apr 2, 2015 ▶ 1:52
Disclosure
Dix built time series databases on Cassandra twice before InfluxDB
“I've actually built a quote unquote time series database on top of Cassandra on two separate occasions. One for a fintech company and another for a metrics SAS developer monitoring thing.”
Paul Dix Apr 2, 2015 ▶ 3:38
Insight
Dix: Ship code to where data lives, not data to code
“This is the key thing that we learned from Hadoop and Google's MapReduce framework, which is You want to ship the code to where the data lives, not the other way around.”
Paul Dix Apr 2, 2015 ▶ 5:16
Insight
Building analytics apps in 2015 resembles web development in 1998
“Building an application with an analytics component today is like building a web application in 1998. You spend months and millions of dollars building infrastructure before you get to the actual thing you want to build that's driving user value, right?”
Paul Dix Apr 2, 2015 ▶ 7:10
Assertion Supported
Dix: InfluxDB operates with zero external software dependencies
“It's an open source time series database with no external dependencies.”
Paul Dix Apr 2, 2015 ▶ 7:38
Assertion Supported
Dix: InfluxDB pairs a time series engine with an in-memory index
“So the InfluxDB is actually, it's kind of like two databases in one. So the time, there's the time series database, and that's useful for storing both regular and irregular time series. Regular is collected on fixed intervals, like once every 10 seconds. Irreg…”
Paul Dix Apr 2, 2015 ▶ 11:47
Assertion Not checkable as stated
Dix: Most InfluxDB users use Grafana for data visualization
“And we find that most of the people using InfluxDB use this to visualize their data and create dashboards for all the data that's going into Influx.”
Paul Dix Apr 2, 2015 ▶ 14:29
Assertion Not checkable as stated
No open-source projects focused on time-series databases in 2013
“So we looked at the open source time series space and found that nobody was really focused on it, and it seemed like A need that was emerging.”
Paul Dix Apr 2, 2015 ▶ 16:03
Assertion Supported
Paul Dix: InfluxDB compresses measurement names and tags into eight-byte IDs
“When you send in a measurement name and a tag set, we compress all of that down into a single ID, a single eight-byte ID.”
Paul Dix Apr 2, 2015 ▶ 17:55
Assertion Supported
Go's garbage collection makes InfluxDB unsuitable for sub-millisecond trading
“This isn't, this is written in Go, so, which is a garbage collected language, so worst case response time can be worse than what you would want in that kind of setup.”
Paul Dix Apr 2, 2015 ▶ 20:20
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.