Jan 23, 2025 · 1h 2m · mad

Understanding Data Engineering in 2025 | Ben Rogojan, Seattle Data Guy

Ben Rogojan (Seattle Data Guy) · 44m spoken Matt Turck · 13m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Ben Rogojan ('Seattle Data Guy') to explore the core responsibilities, technical skills, and practical workflows of data engineers. They analyze current market trends, tool consolidation, open table standards like Apache Iceberg, and strategic industry predictions for AI and data infrastructure in 2025.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 22.6% of the talking time here. How this is scored →

Matt as informed peer 4.2 Guest teaching 5.8 Guest disagreement 0.4 Matt pushing back 0.6
05100:0015:0030:0045:001:00:001:10–4:41 · Matt as informed peer 3/10 Welcome and Ben Rogojan's Journey into Data Engineering Matt opens with a warm welcome and sets up 2025 trends before asking Ben about his career trajectory. Ben details his background transitioning from healthcare analytics to startups and Big Tech at Facebook.4:41–9:23 · Matt as informed peer 3/10 Building a Personal Brand and Content Creation Insights Matt asks about building a brand and content creation, as well as core definitions of data engineering. Ben explains the importance of consistency and defines data engineering as making disparate data usable for humans and machines.9:23–13:01 · Matt as informed peer 3/10 Distinguishing Data Engineers, Data Scientists, and Data Analysts Matt prompts Ben to differentiate between data analysts, data scientists, and data engineers. Ben breaks down the distinct scopes of work, noting how smaller companies often combine these roles.13:01–15:26 · Matt as informed peer 4/10 Data Engineering Reality: Big Tech vs. Non-Silicon Valley Companies Matt asks how data engineering differs between Big Tech and traditional non-Silicon Valley companies. Ben explains that Big Tech has standardized internal infrastructure, whereas non-tech firms deal with fragmented tools and VP-driven software choices.15:26–17:54 · Matt as informed peer 3/10 Career Pathways and Key Technical Skills for Data Engineers Matt asks about technical entry points and required skills for aspiring data engineers. Ben outlines core language requirements like Python and command line proficiency.17:54–20:51 · Matt as informed peer 4/10 SQL Mastery and Foundations of Data Modeling Matt asks about SQL mastery and requests a definition of data modeling. Ben details the contrast between transactional OLTP models and analytical OLAP dimensional modeling.20:51–23:22 · Matt as informed peer 4/10 Essential Data Engineering Frameworks, Cloud Platforms, and Orchestration Matt asks about secondary tech stack requirements like Spark, Kafka, and orchestration tools. Ben outlines cloud platform preferences and recommends tools like Airflow for orchestration.23:22–28:46 · Matt as informed peer 5/10 Developing Soft Skills, Business Acumen, and Stakeholder Collaboration Matt explores soft skills and summarizes that technical teams must learn business speak while business leaders learn tech. Ben agrees and emphasizes how proactive data analysis creates strategic value.28:46–32:12 · Matt as informed peer 4/10 AI Automation in Data Engineering: Practical Realities and System Risks Matt brings up AI code automation and asks if engineers will be replaced. Ben highlights that writing code faster isn't the main goal, sharing a story of Databricks Genie generating flawed queries that required human expertise to fix.32:12–36:53 · Matt as informed peer 5/10 Ingestion Challenges, Automated Connectors, and Pipeline Economics Matt asks why the tooling ecosystem remains so fragmented, referencing his MAD landscape map. Ben playfully calls out venture capitalists for overfunding redundant tools, prompting Matt's humorous acceptance of responsibility.36:53–42:05 · Matt as informed peer 7/10 Vendor Battles, Platform Consolidation, and the Apache Iceberg Standard Matt demonstrates high domain knowledge when contextualizing platform consolidation and vendor wars between Databricks and Snowflake over open table standards like Apache Iceberg.42:18–46:02 · Matt as informed peer 5/10 SQL Server and On-Premises Cloud Migrations Matt references Ben's writing on real-world on-premise SQL Server migrations and fractional data teams. Ben details the limitations of older cloud services like Redshift and why companies seek practical cloud transitions.46:02–48:43 · Matt as informed peer 4/10 Navigating Data Architecture for Early-Stage Companies Matt asks how early-stage startups should structure data infrastructure and hiring. Ben recommends leveraging fractional consultants for initial setup before hiring full-time embedded analysts.48:43–51:34 · Matt as informed peer 4/10 Evaluating Data Warehousing and Orchestration Tools Matt prompts Ben to evaluate specific data platforms and tools. Ben outlines his preference for Snowflake and BigQuery over Azure or legacy architectures due to operational simplicity.51:34–53:55 · Matt as informed peer 4/10 2025 Prediction 2: SQL Isn't Going Anywhere Ben predicts SQL will remain dominant despite text-to-SQL AI promises. Matt asks about natural language query interfaces, and Ben explains why complex queries still demand human engineering.53:55–58:31 · Matt as informed peer 6/10 2025 Prediction 3: AI Moves From Press Releases to Production When Ben notes that self-service analytics remains an unfulfilled holy grail, Matt pushes back forcefully, questioning why decades of industry experience and big data hype haven't produced a standard playbook.58:31–1:02:01 · Matt as informed peer 3/10 2025 Prediction 5: Vertical-Specific Data Solutions Ben shares his final prediction regarding vertical-specific data solutions like healthcare data standardization. Matt closes out the interview and thanks the guest.1:10–4:41 · Guest teaching 5/10 Welcome and Ben Rogojan's Journey into Data Engineering Matt opens with a warm welcome and sets up 2025 trends before asking Ben about his career trajectory. Ben details his background transitioning from healthcare analytics to startups and Big Tech at Facebook.4:41–9:23 · Guest teaching 6/10 Building a Personal Brand and Content Creation Insights Matt asks about building a brand and content creation, as well as core definitions of data engineering. Ben explains the importance of consistency and defines data engineering as making disparate data usable for humans and machines.9:23–13:01 · Guest teaching 6/10 Distinguishing Data Engineers, Data Scientists, and Data Analysts Matt prompts Ben to differentiate between data analysts, data scientists, and data engineers. Ben breaks down the distinct scopes of work, noting how smaller companies often combine these roles.13:01–15:26 · Guest teaching 7/10 Data Engineering Reality: Big Tech vs. Non-Silicon Valley Companies Matt asks how data engineering differs between Big Tech and traditional non-Silicon Valley companies. Ben explains that Big Tech has standardized internal infrastructure, whereas non-tech firms deal with fragmented tools and VP-driven software choices.15:26–17:54 · Guest teaching 6/10 Career Pathways and Key Technical Skills for Data Engineers Matt asks about technical entry points and required skills for aspiring data engineers. Ben outlines core language requirements like Python and command line proficiency.17:54–20:51 · Guest teaching 7/10 SQL Mastery and Foundations of Data Modeling Matt asks about SQL mastery and requests a definition of data modeling. Ben details the contrast between transactional OLTP models and analytical OLAP dimensional modeling.20:51–23:22 · Guest teaching 6/10 Essential Data Engineering Frameworks, Cloud Platforms, and Orchestration Matt asks about secondary tech stack requirements like Spark, Kafka, and orchestration tools. Ben outlines cloud platform preferences and recommends tools like Airflow for orchestration.23:22–28:46 · Guest teaching 5/10 Developing Soft Skills, Business Acumen, and Stakeholder Collaboration Matt explores soft skills and summarizes that technical teams must learn business speak while business leaders learn tech. Ben agrees and emphasizes how proactive data analysis creates strategic value.28:46–32:12 · Guest teaching 6/10 AI Automation in Data Engineering: Practical Realities and System Risks Matt brings up AI code automation and asks if engineers will be replaced. Ben highlights that writing code faster isn't the main goal, sharing a story of Databricks Genie generating flawed queries that required human expertise to fix.32:12–36:53 · Guest teaching 5/10 Ingestion Challenges, Automated Connectors, and Pipeline Economics Matt asks why the tooling ecosystem remains so fragmented, referencing his MAD landscape map. Ben playfully calls out venture capitalists for overfunding redundant tools, prompting Matt's humorous acceptance of responsibility.36:53–42:05 · Guest teaching 5/10 Vendor Battles, Platform Consolidation, and the Apache Iceberg Standard Matt demonstrates high domain knowledge when contextualizing platform consolidation and vendor wars between Databricks and Snowflake over open table standards like Apache Iceberg.42:18–46:02 · Guest teaching 6/10 SQL Server and On-Premises Cloud Migrations Matt references Ben's writing on real-world on-premise SQL Server migrations and fractional data teams. Ben details the limitations of older cloud services like Redshift and why companies seek practical cloud transitions.46:02–48:43 · Guest teaching 6/10 Navigating Data Architecture for Early-Stage Companies Matt asks how early-stage startups should structure data infrastructure and hiring. Ben recommends leveraging fractional consultants for initial setup before hiring full-time embedded analysts.48:43–51:34 · Guest teaching 6/10 Evaluating Data Warehousing and Orchestration Tools Matt prompts Ben to evaluate specific data platforms and tools. Ben outlines his preference for Snowflake and BigQuery over Azure or legacy architectures due to operational simplicity.51:34–53:55 · Guest teaching 6/10 2025 Prediction 2: SQL Isn't Going Anywhere Ben predicts SQL will remain dominant despite text-to-SQL AI promises. Matt asks about natural language query interfaces, and Ben explains why complex queries still demand human engineering.53:55–58:31 · Guest teaching 5/10 2025 Prediction 3: AI Moves From Press Releases to Production When Ben notes that self-service analytics remains an unfulfilled holy grail, Matt pushes back forcefully, questioning why decades of industry experience and big data hype haven't produced a standard playbook.58:31–1:02:01 · Guest teaching 5/10 2025 Prediction 5: Vertical-Specific Data Solutions Ben shares his final prediction regarding vertical-specific data solutions like healthcare data standardization. Matt closes out the interview and thanks the guest.1:10–4:41 · Guest disagreement 0/10 Welcome and Ben Rogojan's Journey into Data Engineering Matt opens with a warm welcome and sets up 2025 trends before asking Ben about his career trajectory. Ben details his background transitioning from healthcare analytics to startups and Big Tech at Facebook.4:41–9:23 · Guest disagreement 0/10 Building a Personal Brand and Content Creation Insights Matt asks about building a brand and content creation, as well as core definitions of data engineering. Ben explains the importance of consistency and defines data engineering as making disparate data usable for humans and machines.9:23–13:01 · Guest disagreement 0/10 Distinguishing Data Engineers, Data Scientists, and Data Analysts Matt prompts Ben to differentiate between data analysts, data scientists, and data engineers. Ben breaks down the distinct scopes of work, noting how smaller companies often combine these roles.13:01–15:26 · Guest disagreement 1/10 Data Engineering Reality: Big Tech vs. Non-Silicon Valley Companies Matt asks how data engineering differs between Big Tech and traditional non-Silicon Valley companies. Ben explains that Big Tech has standardized internal infrastructure, whereas non-tech firms deal with fragmented tools and VP-driven software choices.15:26–17:54 · Guest disagreement 0/10 Career Pathways and Key Technical Skills for Data Engineers Matt asks about technical entry points and required skills for aspiring data engineers. Ben outlines core language requirements like Python and command line proficiency.17:54–20:51 · Guest disagreement 0/10 SQL Mastery and Foundations of Data Modeling Matt asks about SQL mastery and requests a definition of data modeling. Ben details the contrast between transactional OLTP models and analytical OLAP dimensional modeling.20:51–23:22 · Guest disagreement 0/10 Essential Data Engineering Frameworks, Cloud Platforms, and Orchestration Matt asks about secondary tech stack requirements like Spark, Kafka, and orchestration tools. Ben outlines cloud platform preferences and recommends tools like Airflow for orchestration.23:22–28:46 · Guest disagreement 0/10 Developing Soft Skills, Business Acumen, and Stakeholder Collaboration Matt explores soft skills and summarizes that technical teams must learn business speak while business leaders learn tech. Ben agrees and emphasizes how proactive data analysis creates strategic value.28:46–32:12 · Guest disagreement 1/10 AI Automation in Data Engineering: Practical Realities and System Risks Matt brings up AI code automation and asks if engineers will be replaced. Ben highlights that writing code faster isn't the main goal, sharing a story of Databricks Genie generating flawed queries that required human expertise to fix.32:12–36:53 · Guest disagreement 4/10 Ingestion Challenges, Automated Connectors, and Pipeline Economics Matt asks why the tooling ecosystem remains so fragmented, referencing his MAD landscape map. Ben playfully calls out venture capitalists for overfunding redundant tools, prompting Matt's humorous acceptance of responsibility.36:53–42:05 · Guest disagreement 0/10 Vendor Battles, Platform Consolidation, and the Apache Iceberg Standard Matt demonstrates high domain knowledge when contextualizing platform consolidation and vendor wars between Databricks and Snowflake over open table standards like Apache Iceberg.42:18–46:02 · Guest disagreement 0/10 SQL Server and On-Premises Cloud Migrations Matt references Ben's writing on real-world on-premise SQL Server migrations and fractional data teams. Ben details the limitations of older cloud services like Redshift and why companies seek practical cloud transitions.46:02–48:43 · Guest disagreement 0/10 Navigating Data Architecture for Early-Stage Companies Matt asks how early-stage startups should structure data infrastructure and hiring. Ben recommends leveraging fractional consultants for initial setup before hiring full-time embedded analysts.48:43–51:34 · Guest disagreement 0/10 Evaluating Data Warehousing and Orchestration Tools Matt prompts Ben to evaluate specific data platforms and tools. Ben outlines his preference for Snowflake and BigQuery over Azure or legacy architectures due to operational simplicity.51:34–53:55 · Guest disagreement 0/10 2025 Prediction 2: SQL Isn't Going Anywhere Ben predicts SQL will remain dominant despite text-to-SQL AI promises. Matt asks about natural language query interfaces, and Ben explains why complex queries still demand human engineering.53:55–58:31 · Guest disagreement 1/10 2025 Prediction 3: AI Moves From Press Releases to Production When Ben notes that self-service analytics remains an unfulfilled holy grail, Matt pushes back forcefully, questioning why decades of industry experience and big data hype haven't produced a standard playbook.58:31–1:02:01 · Guest disagreement 0/10 2025 Prediction 5: Vertical-Specific Data Solutions Ben shares his final prediction regarding vertical-specific data solutions like healthcare data standardization. Matt closes out the interview and thanks the guest.1:10–4:41 · Matt pushing back 0/10 Welcome and Ben Rogojan's Journey into Data Engineering Matt opens with a warm welcome and sets up 2025 trends before asking Ben about his career trajectory. Ben details his background transitioning from healthcare analytics to startups and Big Tech at Facebook.4:41–9:23 · Matt pushing back 0/10 Building a Personal Brand and Content Creation Insights Matt asks about building a brand and content creation, as well as core definitions of data engineering. Ben explains the importance of consistency and defines data engineering as making disparate data usable for humans and machines.9:23–13:01 · Matt pushing back 0/10 Distinguishing Data Engineers, Data Scientists, and Data Analysts Matt prompts Ben to differentiate between data analysts, data scientists, and data engineers. Ben breaks down the distinct scopes of work, noting how smaller companies often combine these roles.13:01–15:26 · Matt pushing back 0/10 Data Engineering Reality: Big Tech vs. Non-Silicon Valley Companies Matt asks how data engineering differs between Big Tech and traditional non-Silicon Valley companies. Ben explains that Big Tech has standardized internal infrastructure, whereas non-tech firms deal with fragmented tools and VP-driven software choices.15:26–17:54 · Matt pushing back 0/10 Career Pathways and Key Technical Skills for Data Engineers Matt asks about technical entry points and required skills for aspiring data engineers. Ben outlines core language requirements like Python and command line proficiency.17:54–20:51 · Matt pushing back 0/10 SQL Mastery and Foundations of Data Modeling Matt asks about SQL mastery and requests a definition of data modeling. Ben details the contrast between transactional OLTP models and analytical OLAP dimensional modeling.20:51–23:22 · Matt pushing back 0/10 Essential Data Engineering Frameworks, Cloud Platforms, and Orchestration Matt asks about secondary tech stack requirements like Spark, Kafka, and orchestration tools. Ben outlines cloud platform preferences and recommends tools like Airflow for orchestration.23:22–28:46 · Matt pushing back 1/10 Developing Soft Skills, Business Acumen, and Stakeholder Collaboration Matt explores soft skills and summarizes that technical teams must learn business speak while business leaders learn tech. Ben agrees and emphasizes how proactive data analysis creates strategic value.28:46–32:12 · Matt pushing back 0/10 AI Automation in Data Engineering: Practical Realities and System Risks Matt brings up AI code automation and asks if engineers will be replaced. Ben highlights that writing code faster isn't the main goal, sharing a story of Databricks Genie generating flawed queries that required human expertise to fix.32:12–36:53 · Matt pushing back 1/10 Ingestion Challenges, Automated Connectors, and Pipeline Economics Matt asks why the tooling ecosystem remains so fragmented, referencing his MAD landscape map. Ben playfully calls out venture capitalists for overfunding redundant tools, prompting Matt's humorous acceptance of responsibility.36:53–42:05 · Matt pushing back 2/10 Vendor Battles, Platform Consolidation, and the Apache Iceberg Standard Matt demonstrates high domain knowledge when contextualizing platform consolidation and vendor wars between Databricks and Snowflake over open table standards like Apache Iceberg.42:18–46:02 · Matt pushing back 0/10 SQL Server and On-Premises Cloud Migrations Matt references Ben's writing on real-world on-premise SQL Server migrations and fractional data teams. Ben details the limitations of older cloud services like Redshift and why companies seek practical cloud transitions.46:02–48:43 · Matt pushing back 0/10 Navigating Data Architecture for Early-Stage Companies Matt asks how early-stage startups should structure data infrastructure and hiring. Ben recommends leveraging fractional consultants for initial setup before hiring full-time embedded analysts.48:43–51:34 · Matt pushing back 0/10 Evaluating Data Warehousing and Orchestration Tools Matt prompts Ben to evaluate specific data platforms and tools. Ben outlines his preference for Snowflake and BigQuery over Azure or legacy architectures due to operational simplicity.51:34–53:55 · Matt pushing back 0/10 2025 Prediction 2: SQL Isn't Going Anywhere Ben predicts SQL will remain dominant despite text-to-SQL AI promises. Matt asks about natural language query interfaces, and Ben explains why complex queries still demand human engineering.53:55–58:31 · Matt pushing back 6/10 2025 Prediction 3: AI Moves From Press Releases to Production When Ben notes that self-service analytics remains an unfulfilled holy grail, Matt pushes back forcefully, questioning why decades of industry experience and big data hype haven't produced a standard playbook.58:31–1:02:01 · Matt pushing back 0/10 2025 Prediction 5: Vertical-Specific Data Solutions Ben shares his final prediction regarding vertical-specific data solutions like healthcare data standardization. Matt closes out the interview and thanks the guest.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 60.5% · guest 39.5%0:00 · Matt 60.5% · guest 39.5%3:00 · Matt 17% · guest 83%3:00 · Matt 17% · guest 83%6:00 · Matt 11.9% · guest 88.1%6:00 · Matt 11.9% · guest 88.1%9:00 · Matt 26% · guest 74%9:00 · Matt 26% · guest 74%12:00 · Matt 11.5% · guest 88.5%12:00 · Matt 11.5% · guest 88.5%15:00 · Matt 21.2% · guest 78.8%15:00 · Matt 21.2% · guest 78.8%18:00 · Matt 8.7% · guest 91.3%18:00 · Matt 8.7% · guest 91.3%21:00 · Matt 13.2% · guest 86.8%21:00 · Matt 13.2% · guest 86.8%24:00 · Matt 21.2% · guest 78.8%24:00 · Matt 21.2% · guest 78.8%27:00 · Matt 36.8% · guest 63.2%27:00 · Matt 36.8% · guest 63.2%30:00 · Matt 13.4% · guest 86.6%30:00 · Matt 13.4% · guest 86.6%33:00 · Matt 21.4% · guest 78.6%33:00 · Matt 21.4% · guest 78.6%36:00 · Matt 33.8% · guest 66.2%36:00 · Matt 33.8% · guest 66.2%39:00 · Matt 24.1% · guest 75.9%39:00 · Matt 24.1% · guest 75.9%42:00 · Matt 40.3% · guest 59.7%42:00 · Matt 40.3% · guest 59.7%45:00 · Matt 15.5% · guest 84.5%45:00 · Matt 15.5% · guest 84.5%48:00 · Matt 14.9% · guest 85.1%48:00 · Matt 14.9% · guest 85.1%51:00 · Matt 17.4% · guest 82.6%51:00 · Matt 17.4% · guest 82.6%54:00 · Matt 24.4% · guest 75.6%54:00 · Matt 24.4% · guest 75.6%57:00 · Matt 2.9% · guest 97.1%57:00 · Matt 2.9% · guest 97.1%1:00:00 · Matt 43% · guest 57%1:00:00 · Matt 43% · guest 57%
Sharpest disagreement ▶ 35:26 VC Blame Game

Ben playfully targets host Matt Turck as a venture capitalist, stating 'VC funding was pretty high, so you can blame yourselves' for the overwhelming proliferation of redundant data tools.

Hardest push from Matt ▶ 56:39 Challenging Industry Playbooks

Matt forcefully challenges Ben on why self-service analytics remains an unsolved goal, pressing why two decades of institutional knowledge in big data haven't resulted in a standard operational playbook.

Biggest teaching moment ▶ 13:19 Big Tech vs Traditional Corporate Reality

Ben educates the audience on the stark structural differences between Big Tech DE environments and typical corporate setups, contrasting Facebook's unified internal tooling with the multi-vendor chaos found in traditional enterprises.

Matt holds his own ▶ 41:10 Synthesizing Storage vs Compute Lock-in

Matt displays deep domain knowledge by clearly framing the strategic battle between Snowflake and Databricks, explaining how decoupled open storage standards challenge traditional single-vendor lock-in models.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Welcome and Ben Rogojan's Journey into Data Engineering 3500 Matt opens with a warm welcome and sets up 2025 trends before asking Ben about his career trajectory. Ben details his background transitioning from healthcare analytics to startups and Big Tech at Facebook.
Building a Personal Brand and Content Creation Insights 3600 Matt asks about building a brand and content creation, as well as core definitions of data engineering. Ben explains the importance of consistency and defines data engineering as making disparate data usable for humans and machines.
Distinguishing Data Engineers, Data Scientists, and Data Analysts 3600 Matt prompts Ben to differentiate between data analysts, data scientists, and data engineers. Ben breaks down the distinct scopes of work, noting how smaller companies often combine these roles.
Data Engineering Reality: Big Tech vs. Non-Silicon Valley Companies 4710 Matt asks how data engineering differs between Big Tech and traditional non-Silicon Valley companies. Ben explains that Big Tech has standardized internal infrastructure, whereas non-tech firms deal with fragmented tools and VP-driven software choices.
Career Pathways and Key Technical Skills for Data Engineers 3600 Matt asks about technical entry points and required skills for aspiring data engineers. Ben outlines core language requirements like Python and command line proficiency.
SQL Mastery and Foundations of Data Modeling 4700 Matt asks about SQL mastery and requests a definition of data modeling. Ben details the contrast between transactional OLTP models and analytical OLAP dimensional modeling.
Essential Data Engineering Frameworks, Cloud Platforms, and Orchestration 4600 Matt asks about secondary tech stack requirements like Spark, Kafka, and orchestration tools. Ben outlines cloud platform preferences and recommends tools like Airflow for orchestration.
Developing Soft Skills, Business Acumen, and Stakeholder Collaboration 5501 Matt explores soft skills and summarizes that technical teams must learn business speak while business leaders learn tech. Ben agrees and emphasizes how proactive data analysis creates strategic value.
AI Automation in Data Engineering: Practical Realities and System Risks 4610 Matt brings up AI code automation and asks if engineers will be replaced. Ben highlights that writing code faster isn't the main goal, sharing a story of Databricks Genie generating flawed queries that required human expertise to fix.
Ingestion Challenges, Automated Connectors, and Pipeline Economics 5541 Matt asks why the tooling ecosystem remains so fragmented, referencing his MAD landscape map. Ben playfully calls out venture capitalists for overfunding redundant tools, prompting Matt's humorous acceptance of responsibility.
Vendor Battles, Platform Consolidation, and the Apache Iceberg Standard 7502 Matt demonstrates high domain knowledge when contextualizing platform consolidation and vendor wars between Databricks and Snowflake over open table standards like Apache Iceberg.
SQL Server and On-Premises Cloud Migrations 5600 Matt references Ben's writing on real-world on-premise SQL Server migrations and fractional data teams. Ben details the limitations of older cloud services like Redshift and why companies seek practical cloud transitions.
Navigating Data Architecture for Early-Stage Companies 4600 Matt asks how early-stage startups should structure data infrastructure and hiring. Ben recommends leveraging fractional consultants for initial setup before hiring full-time embedded analysts.
Evaluating Data Warehousing and Orchestration Tools 4600 Matt prompts Ben to evaluate specific data platforms and tools. Ben outlines his preference for Snowflake and BigQuery over Azure or legacy architectures due to operational simplicity.
2025 Prediction 2: SQL Isn't Going Anywhere 4600 Ben predicts SQL will remain dominant despite text-to-SQL AI promises. Matt asks about natural language query interfaces, and Ben explains why complex queries still demand human engineering.
2025 Prediction 3: AI Moves From Press Releases to Production 6516 When Ben notes that self-service analytics remains an unfulfilled holy grail, Matt pushes back forcefully, questioning why decades of industry experience and big data hype haven't produced a standard playbook.
2025 Prediction 5: Vertical-Specific Data Solutions 3500 Ben shares his final prediction regarding vertical-specific data solutions like healthcare data standardization. Matt closes out the interview and thanks the guest.

Statements from this episode (26)

Insight
Rogojan: Organizations repeat foundational data mistakes in every AI hype cycle
“People Kind of go through the same iterations. New people come to the space or possibly maybe it's more the business side get really excited by the prospect of what AI or data science or neural networks, whatever iteration we're in can do. And then we kind of …”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 1:34
Assertion Not checkable as stated
Ben Rogojan observed ML engineers at Facebook accessing data directly from warehouses
“Some companies do have maybe the ML engineers access data more directly from maybe the data warehouse, which was things I saw like at Facebook.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 8:29
Assertion Supported
Riot Games deployed machine learning models using Apache Airflow
“One of the teams at Riot Games was like deploying their ML models and they were using Airflow.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 8:50
Assertion Not checkable as stated
Ben Rogojan: Data type issues still routinely break data pipelines
“But to this day, people are still having like a data type issue break a pipeline that then causes everyone to have to kind of spend a little bit of time fixing them.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 12:49
Insight
Ben Rogojan: Non-Big Tech data engineers must manage their own pipeline tooling
“At Facebook, you essentially have your internal airflow that a different team runs, and you can just push files to it and it picks up the data pipelines, you know, automatically. Whereas in most other companies, you might be the person that has to even manage,…”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 13:55
Assertion Not checkable as stated
Ben Rogojan: Most data engineers transition laterally from analyst or developer roles
“I see a lot of people move laterally into data engineering, often from either data analyst or software engineer. Data analyst, because I think there's just the sheer quantity of data analytics jobs are More.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 15:47
Insight
Rogojan: Understanding data flow takes longer than learning SQL syntax
“SQL you can learn quickly. What's going to take time is just getting a sense for how data Operates and flows and data sets in general.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 18:50
Opinion
Rogojan prefers AWS and GCP due to Azure usability issues
“I have my own qualms with Azure every time I use it. So I just like AWS and GCP slightly better.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 21:39
Prediction Not checkable as stated
Rogojan: Enterprise migration away from Spark and Kafka will take years
“I think initially just picking up Kafka, Spark, those are still so heavily in use that even if they go out in the next three to five years, it's going to take time to migrate and things of that nature.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 22:06
Insight
Ben Rogojan: Data teams bear primary responsibility for bridging business communication gaps
“I think in general, I found that if you're the data team, you're going to probably have to do a little more of the lift just because it's, Maybe a little easier for you to make that leap.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 27:51
Prediction Not checkable as stated
Rogojan: Fully replacing data engineers with AI is still years away
“So I think we're a few years away from just like replacing data engineers, which has been the goal for what it feels like a decade or plus now.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 32:00
Insight
Ben Rogojan: API connections are usually the hardest part of data extraction
“Cause that's usually the big thing. It's like the API connection tends to be the one that is the hardest.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 33:38
Assertion Supported
Rogojan: Legacy data software like Informatica cost six to seven figures
“Before it's like, okay, you want Informatica, you know, let's talk six, seven figures in, in contract costs.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 35:33
Opinion
Rogojan blames excessive VC funding for the bloated data engineering tool ecosystem
“I think the reason you see so many tools is like, I mean, VC funding was pretty high, so you can blame yourselves.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 36:33
Prediction Not checkable as stated
Rogojan: Apache Iceberg is becoming the de facto open table format standard
“Iceberg as the de facto, like if we're going to use an open, open table format, that's going to be likely the one people will, I assume, pick since everyone's now supporting it, right?”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 39:14
Insight
Rogojan: Databricks and Snowflake are fighting for workflows, not contract lock-in
“It's really been a white fight for workflows versus a fight for like contracts. Cause in the past it was like, okay, you sign a ten million dollar contract for Teradata. You're kind of stuck with it, right? Like there's no choice now for the next three to five…”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 41:45
Opinion
Ben Rogojan is not a fan of Amazon Redshift
“Not a big Redshift fan.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 43:14
Insight
Rogojan: Early startups should hire fractional data engineers and full-time analysts
“I do think like if you're early starting out, There's no problem in probably bringing on some sort of consultant to do maybe more of the data engine work and then bring on a full-time maybe analyst to kind of work on top of that is what I'd imagine would be go…”
Ben Rogojan Jan 23, 2025 ▶ 46:35
Insight
Ben Rogojan: Early startups can handle baseline reporting using transactional databases
“Like I think what usually I see happens is like initially people just report off whatever their transactional system is. So if you're using Postgres or something and you're, you know, cause you're you've got some application layer. You can report off that for …”
Ben Rogojan Jan 23, 2025 ▶ 47:18
Opinion
Rogojan: Snowflake is difficult to beat for lean, resource-constrained data teams
“Snowflake makes it really easy, and I'd say BigQuery does too, but Snowflake makes it really easy. Like, if you do not have a ton of people to manage things, I think it's hard to go against it.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 49:01
Opinion
Rogojan: Databricks is better suited for machine learning than data warehousing
“I think Databricks makes sense to me still more from an ML standpoint. I know they like the data warehousing side. I know they, that's where they're going to pitch themselves”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 49:14
Prediction Not checkable as stated
Rogojan: SQL will remain essential as LLM raw data-dumping approaches fail
“The next one I said SQL isn't going anywhere, and I kind of said that for more than one reason. One, I feel like someone's going to tell us it's time to revamp data lake v one again, where we'll just dump everything in there and we'll let an LLM figure it out …”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 51:39
Prediction Not checkable as stated
Ben Rogojan: Natural language-to-SQL AI will not take over anytime soon
“I don't think we're going to see that take over anytime soon, unless you have your data in such a well-formatted way that, like, yeah, of course it's easy to do.”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 53:12
Prediction Not checkable as stated
Rogojan: Companies will keep chasing data holy grails like self-service analytics
“I think we'll continue chasing the same holy grails. I feel like we want to see new things, but I think We're still trying to figure out some things like self-service analytics or some of the other holy grails, like even being like data driven kind of feels li…”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 55:07
Insight
Rogojan: Data teams get trampled by fickle business demands without strong leaders
“If you don't have a strong data leader that can either push back or make sure that you're doing things in a way that makes sense you kind of can get trampled by the business side and what they need and their kind of fickle nature where it's like, okay, now we'…”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 58:05
Prediction Not checkable as stated
Rogojan predicts a rise in vertical-specific data solutions for domain data
“Yeah, I think we'll probably start seeing some vertical specific data solutions come out whether that's like tooling or maybe just one example I gave in, in my article was like Tuva Health, which is trying to make it easier to just standardize healthcare speci…”
Ben Rogojan (Seattle Data Guy) Jan 23, 2025 ▶ 58:31
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.