Jun 21, 2021 · 26m · mad

Fireside Chat: Abe Gong (Founder & CEO, Superconductive) with Matt Turck (Partner, FirstMark)

Abe Gong · 18m spoken Matt Turck · 5m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Data Driven NYC, Abe Gong, Founder and CEO of Superconductive, joins host Matt Turck to discuss the open-source data quality framework Great Expectations, exploring how explicit validation rules, community-driven development, and commercial cloud features solve modern data pipeline challenges.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 21.7% of the talking time here. How this is scored →

Matt as informed peer 2.9 Guest teaching 6.0 Guest disagreement 0.8 Matt pushing back 1.3
05100:0010:0020:000:11–4:05 · Matt as informed peer 2/10 Defining Data Quality and Business Impact Matt introduces Abe and sets up the overarching question about data quality, playfully imposing a 10-second limit. Abe breaks down the definition between precise technical invariance for engineers and poorly articulated business concerns.4:05–7:07 · Matt as informed peer 1/10 Root Causes of Data Errors and Healthcare Case Study Matt asks a high-level question about why data errors occur. Abe educates the audience using a concrete healthcare integration case study where upstream gender column codes changed without warning, silently corrupting downstream models.7:07–9:15 · Matt as informed peer 3/10 Rules-Based Expectations vs. Anomaly Detection Matt asks about market approaches to data quality. Abe takes a clear stance against black-box anomaly detection, arguing that explicit rule-based expectations make tacit knowledge visible without introducing uninterpretable systems.9:15–13:04 · Matt as informed peer 5/10 Understanding Expectations and Framework Mechanics Matt demonstrates understanding of developer workflows by asking specific scenario-based questions about alerts, rules, documentation, and extensibility. Abe details how expectations go beyond basic schema into statistical distributions.13:04–15:45 · Matt as informed peer 4/10 Validation, Automated Documentation, and Data Profiling Matt asks Abe to clarify validation, automated documentation, and profiling, then cites specific enterprise and startup customers from Superconductive's roster. Abe explains how automated living docs bridge the gap between engineers and non-technical stakeholders.15:45–18:04 · Matt as informed peer 2/10 Multi-Engine Support and Data Warehouse Migrations Matt listens as Abe details support across Dask, Pandas, Spark, and SQL dialects. Abe highlights how assertions enable smooth data warehouse migrations like moving to Snowflake by ensuring invariants hold across platforms.18:04–21:05 · Matt as informed peer 3/10 Open Source Community Governance and Roadmap Matt prompts Abe to plug and explain their upcoming open source community roadmap event. Abe explains how their user contribution model differs from lower-level infrastructure projects like Docker.21:05–23:08 · Matt as informed peer 3/10 Superconductive Enterprise Strategy and Cloud SaaS Matt guides the conversation toward commercial strategy and pushes for a launch timeline for Great Expectations Cloud. Abe explains the commercial SaaS layer while keeping specific release dates tentative.23:08–25:56 · Matt as informed peer 3/10 Rapid-Fire: Favorite Data Tools and Recommended Resources Matt leads rapid-fire questions on favorite tools and learning resources, asking Abe to spell tool names for the audience. Abe recommends Hasura, SQLFluff, Locally Optimistic, and Amplify.0:11–4:05 · Guest teaching 6/10 Defining Data Quality and Business Impact Matt introduces Abe and sets up the overarching question about data quality, playfully imposing a 10-second limit. Abe breaks down the definition between precise technical invariance for engineers and poorly articulated business concerns.4:05–7:07 · Guest teaching 7/10 Root Causes of Data Errors and Healthcare Case Study Matt asks a high-level question about why data errors occur. Abe educates the audience using a concrete healthcare integration case study where upstream gender column codes changed without warning, silently corrupting downstream models.7:07–9:15 · Guest teaching 7/10 Rules-Based Expectations vs. Anomaly Detection Matt asks about market approaches to data quality. Abe takes a clear stance against black-box anomaly detection, arguing that explicit rule-based expectations make tacit knowledge visible without introducing uninterpretable systems.9:15–13:04 · Guest teaching 5/10 Understanding Expectations and Framework Mechanics Matt demonstrates understanding of developer workflows by asking specific scenario-based questions about alerts, rules, documentation, and extensibility. Abe details how expectations go beyond basic schema into statistical distributions.13:04–15:45 · Guest teaching 6/10 Validation, Automated Documentation, and Data Profiling Matt asks Abe to clarify validation, automated documentation, and profiling, then cites specific enterprise and startup customers from Superconductive's roster. Abe explains how automated living docs bridge the gap between engineers and non-technical stakeholders.15:45–18:04 · Guest teaching 7/10 Multi-Engine Support and Data Warehouse Migrations Matt listens as Abe details support across Dask, Pandas, Spark, and SQL dialects. Abe highlights how assertions enable smooth data warehouse migrations like moving to Snowflake by ensuring invariants hold across platforms.18:04–21:05 · Guest teaching 6/10 Open Source Community Governance and Roadmap Matt prompts Abe to plug and explain their upcoming open source community roadmap event. Abe explains how their user contribution model differs from lower-level infrastructure projects like Docker.21:05–23:08 · Guest teaching 5/10 Superconductive Enterprise Strategy and Cloud SaaS Matt guides the conversation toward commercial strategy and pushes for a launch timeline for Great Expectations Cloud. Abe explains the commercial SaaS layer while keeping specific release dates tentative.23:08–25:56 · Guest teaching 5/10 Rapid-Fire: Favorite Data Tools and Recommended Resources Matt leads rapid-fire questions on favorite tools and learning resources, asking Abe to spell tool names for the audience. Abe recommends Hasura, SQLFluff, Locally Optimistic, and Amplify.0:11–4:05 · Guest disagreement 1/10 Defining Data Quality and Business Impact Matt introduces Abe and sets up the overarching question about data quality, playfully imposing a 10-second limit. Abe breaks down the definition between precise technical invariance for engineers and poorly articulated business concerns.4:05–7:07 · Guest disagreement 0/10 Root Causes of Data Errors and Healthcare Case Study Matt asks a high-level question about why data errors occur. Abe educates the audience using a concrete healthcare integration case study where upstream gender column codes changed without warning, silently corrupting downstream models.7:07–9:15 · Guest disagreement 2/10 Rules-Based Expectations vs. Anomaly Detection Matt asks about market approaches to data quality. Abe takes a clear stance against black-box anomaly detection, arguing that explicit rule-based expectations make tacit knowledge visible without introducing uninterpretable systems.9:15–13:04 · Guest disagreement 1/10 Understanding Expectations and Framework Mechanics Matt demonstrates understanding of developer workflows by asking specific scenario-based questions about alerts, rules, documentation, and extensibility. Abe details how expectations go beyond basic schema into statistical distributions.13:04–15:45 · Guest disagreement 1/10 Validation, Automated Documentation, and Data Profiling Matt asks Abe to clarify validation, automated documentation, and profiling, then cites specific enterprise and startup customers from Superconductive's roster. Abe explains how automated living docs bridge the gap between engineers and non-technical stakeholders.15:45–18:04 · Guest disagreement 0/10 Multi-Engine Support and Data Warehouse Migrations Matt listens as Abe details support across Dask, Pandas, Spark, and SQL dialects. Abe highlights how assertions enable smooth data warehouse migrations like moving to Snowflake by ensuring invariants hold across platforms.18:04–21:05 · Guest disagreement 1/10 Open Source Community Governance and Roadmap Matt prompts Abe to plug and explain their upcoming open source community roadmap event. Abe explains how their user contribution model differs from lower-level infrastructure projects like Docker.21:05–23:08 · Guest disagreement 1/10 Superconductive Enterprise Strategy and Cloud SaaS Matt guides the conversation toward commercial strategy and pushes for a launch timeline for Great Expectations Cloud. Abe explains the commercial SaaS layer while keeping specific release dates tentative.23:08–25:56 · Guest disagreement 0/10 Rapid-Fire: Favorite Data Tools and Recommended Resources Matt leads rapid-fire questions on favorite tools and learning resources, asking Abe to spell tool names for the audience. Abe recommends Hasura, SQLFluff, Locally Optimistic, and Amplify.0:11–4:05 · Matt pushing back 2/10 Defining Data Quality and Business Impact Matt introduces Abe and sets up the overarching question about data quality, playfully imposing a 10-second limit. Abe breaks down the definition between precise technical invariance for engineers and poorly articulated business concerns.4:05–7:07 · Matt pushing back 1/10 Root Causes of Data Errors and Healthcare Case Study Matt asks a high-level question about why data errors occur. Abe educates the audience using a concrete healthcare integration case study where upstream gender column codes changed without warning, silently corrupting downstream models.7:07–9:15 · Matt pushing back 2/10 Rules-Based Expectations vs. Anomaly Detection Matt asks about market approaches to data quality. Abe takes a clear stance against black-box anomaly detection, arguing that explicit rule-based expectations make tacit knowledge visible without introducing uninterpretable systems.9:15–13:04 · Matt pushing back 2/10 Understanding Expectations and Framework Mechanics Matt demonstrates understanding of developer workflows by asking specific scenario-based questions about alerts, rules, documentation, and extensibility. Abe details how expectations go beyond basic schema into statistical distributions.13:04–15:45 · Matt pushing back 1/10 Validation, Automated Documentation, and Data Profiling Matt asks Abe to clarify validation, automated documentation, and profiling, then cites specific enterprise and startup customers from Superconductive's roster. Abe explains how automated living docs bridge the gap between engineers and non-technical stakeholders.15:45–18:04 · Matt pushing back 0/10 Multi-Engine Support and Data Warehouse Migrations Matt listens as Abe details support across Dask, Pandas, Spark, and SQL dialects. Abe highlights how assertions enable smooth data warehouse migrations like moving to Snowflake by ensuring invariants hold across platforms.18:04–21:05 · Matt pushing back 1/10 Open Source Community Governance and Roadmap Matt prompts Abe to plug and explain their upcoming open source community roadmap event. Abe explains how their user contribution model differs from lower-level infrastructure projects like Docker.21:05–23:08 · Matt pushing back 2/10 Superconductive Enterprise Strategy and Cloud SaaS Matt guides the conversation toward commercial strategy and pushes for a launch timeline for Great Expectations Cloud. Abe explains the commercial SaaS layer while keeping specific release dates tentative.23:08–25:56 · Matt pushing back 1/10 Rapid-Fire: Favorite Data Tools and Recommended Resources Matt leads rapid-fire questions on favorite tools and learning resources, asking Abe to spell tool names for the audience. Abe recommends Hasura, SQLFluff, Locally Optimistic, and Amplify.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 24.5% · guest 75.5%0:00 · Matt 24.5% · guest 75.5%3:00 · Matt 8.2% · guest 91.8%3:00 · Matt 8.2% · guest 91.8%6:00 · Matt 10% · guest 90%6:00 · Matt 10% · guest 90%9:00 · Matt 36% · guest 64%9:00 · Matt 36% · guest 64%12:00 · Matt 15.4% · guest 84.6%12:00 · Matt 15.4% · guest 84.6%15:00 · Matt 20.5% · guest 79.5%15:00 · Matt 20.5% · guest 79.5%18:00 · Matt 21.4% · guest 78.6%18:00 · Matt 21.4% · guest 78.6%21:00 · Matt 27.2% · guest 72.8%21:00 · Matt 27.2% · guest 72.8%24:00 · Matt 33% · guest 67%24:00 · Matt 33% · guest 67%
Sharpest disagreement ▶ 7:55 Dismissing anomaly detection approaches

Abe explicitly rejects the popular industry trend of automated anomaly detection, labeling black-box systems as prone to introducing hidden tacit knowledge that shoots teams in the foot.

Hardest push from Matt ▶ 22:34 Pressing on commercial launch timeline

Matt refuses to accept vague descriptions of the Cloud SaaS offering and directly presses Abe on when the product will actually launch.

Biggest teaching moment ▶ 5:50 Healthcare pipeline data corruption case study

Abe delivers an insightful explanation using a real-world healthcare insurance example where a gender column changing from 1 and 2 to 1, 2, 4, and 9 caused silent algorithmic failure.

Matt holds his own ▶ 15:22 Referencing customer profile and stack diversity

Matt displays deep industry familiarity by citing specific Superconductive clients across high-growth startups and traditional Fortune 1000 enterprises to challenge stack compatibility assumptions.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Defining Data Quality and Business Impact 2612 Matt introduces Abe and sets up the overarching question about data quality, playfully imposing a 10-second limit. Abe breaks down the definition between precise technical invariance for engineers and poorly articulated business concerns.
Root Causes of Data Errors and Healthcare Case Study 1701 Matt asks a high-level question about why data errors occur. Abe educates the audience using a concrete healthcare integration case study where upstream gender column codes changed without warning, silently corrupting downstream models.
Rules-Based Expectations vs. Anomaly Detection 3722 Matt asks about market approaches to data quality. Abe takes a clear stance against black-box anomaly detection, arguing that explicit rule-based expectations make tacit knowledge visible without introducing uninterpretable systems.
Understanding Expectations and Framework Mechanics 5512 Matt demonstrates understanding of developer workflows by asking specific scenario-based questions about alerts, rules, documentation, and extensibility. Abe details how expectations go beyond basic schema into statistical distributions.
Validation, Automated Documentation, and Data Profiling 4611 Matt asks Abe to clarify validation, automated documentation, and profiling, then cites specific enterprise and startup customers from Superconductive's roster. Abe explains how automated living docs bridge the gap between engineers and non-technical stakeholders.
Multi-Engine Support and Data Warehouse Migrations 2700 Matt listens as Abe details support across Dask, Pandas, Spark, and SQL dialects. Abe highlights how assertions enable smooth data warehouse migrations like moving to Snowflake by ensuring invariants hold across platforms.
Open Source Community Governance and Roadmap 3611 Matt prompts Abe to plug and explain their upcoming open source community roadmap event. Abe explains how their user contribution model differs from lower-level infrastructure projects like Docker.
Superconductive Enterprise Strategy and Cloud SaaS 3512 Matt guides the conversation toward commercial strategy and pushes for a launch timeline for Great Expectations Cloud. Abe explains the commercial SaaS layer while keeping specific release dates tentative.
Rapid-Fire: Favorite Data Tools and Recommended Resources 3501 Matt leads rapid-fire questions on favorite tools and learning resources, asking Abe to spell tool names for the audience. Abe recommends Hasura, SQLFluff, Locally Optimistic, and Amplify.

Statements from this episode (12)

Assertion Not checkable as stated
Abe Gong: Data quality has become a C-level issue for many companies
“Data quality has kind of bubbled up to the point where it's a C-level issue for a lot of companies”
Abe Gong Jun 21, 2021 ▶ 1:00
Insight
Abe Gong: Data teams must detect broken dashboards before stakeholders notice
“You can't always completely prevent the dashboard from breaking because it might be because of upstream data that you don't control, but at the very least you want to know about it before your stakeholders know about it.”
Abe Gong Jun 21, 2021 ▶ 3:39
Disclosure
Superconductive takes a fine-grained rules-based approach to data quality
“We've taken a very rules-based approach in the sense of declaring exactly what you expect of data you know, very fine grained detail.”
Abe Gong Jun 21, 2021 ▶ 7:38
Opinion
Gong is skeptical of anomaly detection approaches for data quality
“Frankly, I'm pretty skeptical of the anomaly detection approach.”
Abe Gong Jun 21, 2021 ▶ 8:17
Insight
Abe Gong: Software tests are valuable because they make tacit knowledge explicit.
“You need tests because they make tacit knowledge explicit. They make it very clear what the system is intended to do.”
Abe Gong Jun 21, 2021 ▶ 8:31
Opinion
Abe Gong: No rival project has deeply invested in automated data documentation
“Data documentation is a thing that we consider a really important part of the project and frankly a technology that I don't think anybody else has invested deeply in building.”
Abe Gong Jun 21, 2021 ▶ 13:54
Insight
Abe Gong: Great Expectations guarantees documentation reflects live data state
“Unlike the, like the classic outdated data wiki that most teams have lived with in the past with great expectations, as long as you're running your tests, you know, that your documentation actually reflects the current state of the data, which is kind of a gua…”
Abe Gong Jun 21, 2021 ▶ 15:04
Assertion Contradicted
Abe Gong: Great Expectations assertions execute consistently across Pandas, Spark, and SQL
“The kind of guarantee is any given expectation will execute in any of those environments and it'll execute consistently.”
Abe Gong Jun 21, 2021 ▶ 17:25
Assertion Supported
Superconductive pivoted from healthcare analytics to open source around 2019
“We actually started in healthcare data analytics and deliberately pivoted the company to open source about two years ago now, when great expectations started to just really take on a life of its own.”
Abe Gong Jun 21, 2021 ▶ 21:27
Disclosure
Abe Gong: Superconductive will permanently keep open-source Great Expectations open
“You know, have committed, will always commit to keeping everything that's in open source, great expectations open.”
Abe Gong Jun 21, 2021 ▶ 22:16
Opinion
Abe Gong calls Hasura 'just add water GraphQL' for Postgres databases.
“It's just add water GraphQL for Postgres. So if you've got a Postgres database, then you can point Hasura at it and extremely quickly spin up a GraphQL API for it. And subscriptions, queries, mutations, like you get all of those things. It's all code gen. So t…”
Abe Gong Jun 21, 2021 ▶ 23:49
Opinion
Gong: The data ecosystem critically needs a SQL linter like SQLFluff
“The other one that I, I'd point to that I haven't used deeply myself, but I just think is a thing that the world really needs is SQL fluff. SQL FL U F F. And it's a linter for SQL. And I'm just kind of shocked that we have got to 20 21 without there being a, l…”
Abe Gong Jun 21, 2021 ▶ 24:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.