Oct 27, 2023 · 43m · latent-space

Powering your Copilot for Data - with Artem Keydunov from Cube.dev

Artem Keydunov · 29m spoken Alessio Fanelli · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Cube co-founder Artem Keydunov joins hosts Swix and Alessio Fanelli to discuss how semantic layers bridge the gap between large language models and tabular databases. He details the technical architecture, historical lessons from StatsBot, and software engineering best practices necessary to build production-grade AI data copilots.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 21.1% of the talking time here. How this is scored →

The hosts as informed peer 6.1 Guest teaching 4.0 Guest disagreement 1.3 The hosts pushing back 2.3
05100:0015:0030:000:07–6:24 · The hosts as informed peer 6/10 The Origins of StatsBot and Early Text-to-SQL Limitations Alessio and Swix demonstrate solid contextual knowledge of early text-to-SQL products and the 2016 chatbot era. Swix bonds over using regular expressions for parsing financial data while Artem explains StatsBot's origins.6:25–9:53 · The hosts as informed peer 5/10 The Evolution of Cube and the Fundamentals of OLAP Cubes Swix prompts Artem to define OLAP and multidimensional cubes for the audience, establishing baseline definitions. Artem explains the distinction between 2D relational SQL tables and multidimensional metric cubes.9:53–16:27 · The hosts as informed peer 7/10 Bridging LLMs and Tabular Data via the Semantic Layer Artem gives an in-depth breakdown of how LLMs fail at complex SQL joins (fan traps, chasm traps) and how semantic layers act as structured context. Alessio contributes concrete architectural summaries to ground the explanation.16:27–19:18 · The hosts as informed peer 6/10 Treating Data Metrics as Code to Resolve Organizational Discrepancies Swix pushes back on the idealistic view of semantic layers by raising organizational dysfunction and conflicting departmental definitions of core metrics. Artem addresses this by arguing metrics should be governed as code with pull requests.19:18–23:59 · The hosts as informed peer 6/10 Data Copilots versus Natural Language BI Interfaces Alessio queries how BI vendors can differentiate if natural language query interfaces become commoditized. Artem explains that BI fragmentation will persist because natural language query is just another commodity visual interface like bar charts.23:59–30:14 · The hosts as informed peer 7/10 AI Agents, Smart Summaries, and the Modern Data Stack Evolution Alessio draws parallels to Linus's UI ideas at Notion and discusses modern data stack automation across dbt and ETL pipelines. Artem provides a measured assessment of where LLMs genuinely accelerate data engineering versus standard consolidation cycles.30:14–36:14 · The hosts as informed peer 6/10 Technical Nuances of Data Chatbots and Production Stacks Artem elaborates on practical production implementations, emphasizing that effective data bots ask clarifying follow-up questions and rely on external Python logic for accurate arithmetic rather than raw LLM generation.36:15–42:46 · The hosts as informed peer 6/10 Embedded Analytics Monetization and the Lightning Round Swix raises the monetization struggle in embedded analytics, prompting Artem to analyze market saturation and custom build dynamics. The segment transitions into rapid-fire lightning round questions on foundation models and AI engineering practices.0:07–6:24 · Guest teaching 3/10 The Origins of StatsBot and Early Text-to-SQL Limitations Alessio and Swix demonstrate solid contextual knowledge of early text-to-SQL products and the 2016 chatbot era. Swix bonds over using regular expressions for parsing financial data while Artem explains StatsBot's origins.6:25–9:53 · Guest teaching 5/10 The Evolution of Cube and the Fundamentals of OLAP Cubes Swix prompts Artem to define OLAP and multidimensional cubes for the audience, establishing baseline definitions. Artem explains the distinction between 2D relational SQL tables and multidimensional metric cubes.9:53–16:27 · Guest teaching 6/10 Bridging LLMs and Tabular Data via the Semantic Layer Artem gives an in-depth breakdown of how LLMs fail at complex SQL joins (fan traps, chasm traps) and how semantic layers act as structured context. Alessio contributes concrete architectural summaries to ground the explanation.16:27–19:18 · Guest teaching 4/10 Treating Data Metrics as Code to Resolve Organizational Discrepancies Swix pushes back on the idealistic view of semantic layers by raising organizational dysfunction and conflicting departmental definitions of core metrics. Artem addresses this by arguing metrics should be governed as code with pull requests.19:18–23:59 · Guest teaching 3/10 Data Copilots versus Natural Language BI Interfaces Alessio queries how BI vendors can differentiate if natural language query interfaces become commoditized. Artem explains that BI fragmentation will persist because natural language query is just another commodity visual interface like bar charts.23:59–30:14 · Guest teaching 4/10 AI Agents, Smart Summaries, and the Modern Data Stack Evolution Alessio draws parallels to Linus's UI ideas at Notion and discusses modern data stack automation across dbt and ETL pipelines. Artem provides a measured assessment of where LLMs genuinely accelerate data engineering versus standard consolidation cycles.30:14–36:14 · Guest teaching 4/10 Technical Nuances of Data Chatbots and Production Stacks Artem elaborates on practical production implementations, emphasizing that effective data bots ask clarifying follow-up questions and rely on external Python logic for accurate arithmetic rather than raw LLM generation.36:15–42:46 · Guest teaching 3/10 Embedded Analytics Monetization and the Lightning Round Swix raises the monetization struggle in embedded analytics, prompting Artem to analyze market saturation and custom build dynamics. The segment transitions into rapid-fire lightning round questions on foundation models and AI engineering practices.0:07–6:24 · Guest disagreement 1/10 The Origins of StatsBot and Early Text-to-SQL Limitations Alessio and Swix demonstrate solid contextual knowledge of early text-to-SQL products and the 2016 chatbot era. Swix bonds over using regular expressions for parsing financial data while Artem explains StatsBot's origins.6:25–9:53 · Guest disagreement 1/10 The Evolution of Cube and the Fundamentals of OLAP Cubes Swix prompts Artem to define OLAP and multidimensional cubes for the audience, establishing baseline definitions. Artem explains the distinction between 2D relational SQL tables and multidimensional metric cubes.9:53–16:27 · Guest disagreement 1/10 Bridging LLMs and Tabular Data via the Semantic Layer Artem gives an in-depth breakdown of how LLMs fail at complex SQL joins (fan traps, chasm traps) and how semantic layers act as structured context. Alessio contributes concrete architectural summaries to ground the explanation.16:27–19:18 · Guest disagreement 2/10 Treating Data Metrics as Code to Resolve Organizational Discrepancies Swix pushes back on the idealistic view of semantic layers by raising organizational dysfunction and conflicting departmental definitions of core metrics. Artem addresses this by arguing metrics should be governed as code with pull requests.19:18–23:59 · Guest disagreement 2/10 Data Copilots versus Natural Language BI Interfaces Alessio queries how BI vendors can differentiate if natural language query interfaces become commoditized. Artem explains that BI fragmentation will persist because natural language query is just another commodity visual interface like bar charts.23:59–30:14 · Guest disagreement 1/10 AI Agents, Smart Summaries, and the Modern Data Stack Evolution Alessio draws parallels to Linus's UI ideas at Notion and discusses modern data stack automation across dbt and ETL pipelines. Artem provides a measured assessment of where LLMs genuinely accelerate data engineering versus standard consolidation cycles.30:14–36:14 · Guest disagreement 1/10 Technical Nuances of Data Chatbots and Production Stacks Artem elaborates on practical production implementations, emphasizing that effective data bots ask clarifying follow-up questions and rely on external Python logic for accurate arithmetic rather than raw LLM generation.36:15–42:46 · Guest disagreement 1/10 Embedded Analytics Monetization and the Lightning Round Swix raises the monetization struggle in embedded analytics, prompting Artem to analyze market saturation and custom build dynamics. The segment transitions into rapid-fire lightning round questions on foundation models and AI engineering practices.0:07–6:24 · The hosts pushing back 1/10 The Origins of StatsBot and Early Text-to-SQL Limitations Alessio and Swix demonstrate solid contextual knowledge of early text-to-SQL products and the 2016 chatbot era. Swix bonds over using regular expressions for parsing financial data while Artem explains StatsBot's origins.6:25–9:53 · The hosts pushing back 1/10 The Evolution of Cube and the Fundamentals of OLAP Cubes Swix prompts Artem to define OLAP and multidimensional cubes for the audience, establishing baseline definitions. Artem explains the distinction between 2D relational SQL tables and multidimensional metric cubes.9:53–16:27 · The hosts pushing back 2/10 Bridging LLMs and Tabular Data via the Semantic Layer Artem gives an in-depth breakdown of how LLMs fail at complex SQL joins (fan traps, chasm traps) and how semantic layers act as structured context. Alessio contributes concrete architectural summaries to ground the explanation.16:27–19:18 · The hosts pushing back 6/10 Treating Data Metrics as Code to Resolve Organizational Discrepancies Swix pushes back on the idealistic view of semantic layers by raising organizational dysfunction and conflicting departmental definitions of core metrics. Artem addresses this by arguing metrics should be governed as code with pull requests.19:18–23:59 · The hosts pushing back 3/10 Data Copilots versus Natural Language BI Interfaces Alessio queries how BI vendors can differentiate if natural language query interfaces become commoditized. Artem explains that BI fragmentation will persist because natural language query is just another commodity visual interface like bar charts.23:59–30:14 · The hosts pushing back 2/10 AI Agents, Smart Summaries, and the Modern Data Stack Evolution Alessio draws parallels to Linus's UI ideas at Notion and discusses modern data stack automation across dbt and ETL pipelines. Artem provides a measured assessment of where LLMs genuinely accelerate data engineering versus standard consolidation cycles.30:14–36:14 · The hosts pushing back 1/10 Technical Nuances of Data Chatbots and Production Stacks Artem elaborates on practical production implementations, emphasizing that effective data bots ask clarifying follow-up questions and rely on external Python logic for accurate arithmetic rather than raw LLM generation.36:15–42:46 · The hosts pushing back 2/10 Embedded Analytics Monetization and the Lightning Round Swix raises the monetization struggle in embedded analytics, prompting Artem to analyze market saturation and custom build dynamics. The segment transitions into rapid-fire lightning round questions on foundation models and AI engineering practices.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 43.1% · guest 56.9%0:00 · the hosts 43.1% · guest 56.9%3:00 · the hosts 6.3% · guest 93.7%3:00 · the hosts 6.3% · guest 93.7%6:00 · the hosts 13.8% · guest 86.2%6:00 · the hosts 13.8% · guest 86.2%9:00 · the hosts 21.9% · guest 78.1%9:00 · the hosts 21.9% · guest 78.1%12:00 · the hosts 10.5% · guest 89.5%12:00 · the hosts 10.5% · guest 89.5%15:00 · the hosts 14.9% · guest 85.1%15:00 · the hosts 14.9% · guest 85.1%18:00 · the hosts 32.3% · guest 67.7%18:00 · the hosts 32.3% · guest 67.7%21:00 · the hosts 10.7% · guest 89.3%21:00 · the hosts 10.7% · guest 89.3%24:00 · the hosts 45.8% · guest 54.2%24:00 · the hosts 45.8% · guest 54.2%27:00 · the hosts 33.4% · guest 66.6%27:00 · the hosts 33.4% · guest 66.6%30:00 · the hosts 22.4% · guest 77.6%30:00 · the hosts 22.4% · guest 77.6%33:00 · the hosts 27.4% · guest 72.6%33:00 · the hosts 27.4% · guest 72.6%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 16.9% · guest 83.1%39:00 · the hosts 16.9% · guest 83.1%42:00 · the hosts 6.4% · guest 93.6%42:00 · the hosts 6.4% · guest 93.6%
Sharpest disagreement ▶ 16:27 Org chart conflict challenge

Swix directly confronts the clean narrative of semantic layers, arguing that conflicting cross-team metric definitions lead to shipping an org chart into the code.

Hardest push from the hosts ▶ 16:27 Swix presses on conflicting departmental metrics

Swix refuses to treat semantic modeling as purely technical, forcing Artem to address how real-world political friction between finance and sales impacts metric definitions.

Biggest teaching moment ▶ 13:30 Artem details SQL fan traps and semantic constraints

Artem breaks down the exact database pitfalls like fan traps and chasm traps that cause direct LLM-to-SQL generation to produce silent analytical errors.

The host holds their own ▶ 16:00 Alessio articulates metric abstraction mechanics

Alessio demonstrates his technical grasp of data modeling by succinctly illustrating how the semantic layer abstracts multi-table joins into clean single-metric queries.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Origins of StatsBot and Early Text-to-SQL Limitations 6311 Alessio and Swix demonstrate solid contextual knowledge of early text-to-SQL products and the 2016 chatbot era. Swix bonds over using regular expressions for parsing financial data while Artem explains StatsBot's origins.
The Evolution of Cube and the Fundamentals of OLAP Cubes 5511 Swix prompts Artem to define OLAP and multidimensional cubes for the audience, establishing baseline definitions. Artem explains the distinction between 2D relational SQL tables and multidimensional metric cubes.
Bridging LLMs and Tabular Data via the Semantic Layer 7612 Artem gives an in-depth breakdown of how LLMs fail at complex SQL joins (fan traps, chasm traps) and how semantic layers act as structured context. Alessio contributes concrete architectural summaries to ground the explanation.
Treating Data Metrics as Code to Resolve Organizational Discrepancies 6426 Swix pushes back on the idealistic view of semantic layers by raising organizational dysfunction and conflicting departmental definitions of core metrics. Artem addresses this by arguing metrics should be governed as code with pull requests.
Data Copilots versus Natural Language BI Interfaces 6323 Alessio queries how BI vendors can differentiate if natural language query interfaces become commoditized. Artem explains that BI fragmentation will persist because natural language query is just another commodity visual interface like bar charts.
AI Agents, Smart Summaries, and the Modern Data Stack Evolution 7412 Alessio draws parallels to Linus's UI ideas at Notion and discusses modern data stack automation across dbt and ETL pipelines. Artem provides a measured assessment of where LLMs genuinely accelerate data engineering versus standard consolidation cycles.
Technical Nuances of Data Chatbots and Production Stacks 6411 Artem elaborates on practical production implementations, emphasizing that effective data bots ask clarifying follow-up questions and rely on external Python logic for accurate arithmetic rather than raw LLM generation.
Embedded Analytics Monetization and the Lightning Round 6312 Swix raises the monetization struggle in embedded analytics, prompting Artem to analyze market saturation and custom build dynamics. The segment transitions into rapid-fire lightning round questions on foundation models and AI engineering practices.

Statements from this episode (10)

Insight
Keydunov: The semantic layer is the ideal repository for AI data context
“So our take of that and my take is that semantic layer is just really good place for this context to live because you need to give this context to the humans. You need to give that context to the AI system anyway, right? So that's why you define metric once an…”
Artem Keydunov Oct 27, 2023 ▶ 11:36
Insight
Keydunov: Querying a semantic layer over raw SQL reduces LLM errors
“Now, the query, I believe, should be created against semantic layer, because it reduces the room for the error, because what usually happens is that your query to semantic layer would be very simple. It would be like, give me that metric grouped by that dimens…”
Artem Keydunov Oct 27, 2023 ▶ 14:22
Insight
Keydunov: Semantic layer metrics should be managed as code via pull requests
“So I think we should treat our metrics and we are in semantic layer as a code, right? And then collaboration is a big part of it. You know, like if there are like a multiple teams that sort of have a different opinions, let them collaborate on the pull request…”
Artem Keydunov Oct 27, 2023 ▶ 17:29
Prediction Not checkable as stated
Keydunov: Natural language AI will not consolidate the fragmented BI market
“I don't think it's going to be something that will change the BI market, you know, like something that will can take The BI market and make it more consolidated rather than, you know, like what we have right now. I think it's still re will remain fragmented.”
Artem Keydunov Oct 27, 2023 ▶ 23:22
Opinion
Keydunov: Modern data stack will not be dramatically disrupted by AI
“I don't see a lot of being compacted by AI specifically. I think, you know, that space is being compacted as much as any other space in terms of, yes, we'll have all this copilot capabilities, some of AI capabilities here and there, but I don't see anything so…”
Artem Keydunov Oct 27, 2023 ▶ 27:31
Prediction Not checkable as stated
Keydunov: The modern data stack will enter a consolidation cycle
“I feel like right now we should go through the cycle of consolidation. And you know, like I mean, if Fivetrend and dbt merge, they can be alter-ex of a new generation or something like that. And you know, probably some ETL tool to there. But I feel it might ha…”
Artem Keydunov Oct 27, 2023 ▶ 28:06
Insight
Keydunov: Production AI data apps require extensive Python scaffolding around models
“When we're talking about production level use cases, it's quite a lot of Python code around, you know, like, your model to make it work, to be honest. It's, like, It's not that magic that you just throw the model in it, like it can give you all this answers. F…”
Artem Keydunov Oct 27, 2023 ▶ 33:54
Insight
Keydunov: Data-driven AI apps should start on a proper warehouse from day zero
“I would just recommend going through to warehouse as soon as possible. I think a lot of people feel that MySQL can be a warehouse, which can be maybe on like a lower scale, but you know, like, definitely not from a performance perspective. So just kind of havi…”
Artem Keydunov Oct 27, 2023 ▶ 34:56
Insight
Keydunov: Large enterprises abandon embedded analytics vendors to build in-house
“If you look at the embedded analytics market, the bigger organization, the big gets, They're really more custom, you know, like it becomes, and at some point I see many organizations, they just stop using any vendor, and they just kind of build most of the stu…”
Artem Keydunov Oct 27, 2023 ▶ 37:57
Prediction Not checkable as stated
Keydunov: AI will not make embedded analytics easier to monetize
“I don't think AI really going to change that just because it's using model, you just pay to open AI and that's it. Like everyone can do that, right? It's not much of a Competitive advantage. So it's going to be more like a commodity features that a lot of like…”
Artem Keydunov Oct 27, 2023 ▶ 38:46
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.