Jan 2, 2019 · 31m · a16z

a16z Podcast | Making the Most of the Data That Matters

Prat Moghe · 7m spoken Steven Sinofsky · 6m spoken Gaurav Dhillon · 6m spoken Roman Stanek · 6m spoken Michael Copeland · 44s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this a16z podcast panel moderated by Steven Sinofsky, industry leaders Prat Moghe, Gaurav Dhillon, and Roman Stanek discuss how modern enterprises can transform raw data into business outcomes through cloud migration, last-mile analytics delivery, and pragmatic predictive strategies.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 5.0 Guest teaching 3.7 Guest disagreement 3.2 The host pushing back 3.5
05100:0010:0020:0030:002:46–7:31 · The host as informed peer 4/10 Demystifying Big Data and Volume vs. Agility Host Steven Sinofsky opens by framing the central topic of big data and asks guests to demystify what makes data 'big'. Prat Moghe reframes the premise by arguing big data is about agility and decision speed rather than pure petabyte volume. Sinofsky gently intervenes when asking about Kazina to make sure Moghe keeps it focused rather than delivering a canned product pitch.7:31–13:11 · The host as informed peer 4/10 The Last Mile of Analytics and Field Delivery Roman Stanek explains GoodData's focus on the last mile of analytics for non-technical field workers. The guests engage in friendly banter, noting a internal bet about who would voice disagreement first. Sinofsky guides the discussion to explore how companies move from fixed weekly reporting to true data exploration.13:11–16:28 · The host as informed peer 4/10 Rate of Business Change and Full-Stack Experiences Stanek highlights that the primary friction with corporate data is the rapid pace of business change exceeding IT turnaround times. Moghe illustrates this dynamic with an example of a fast-growing restaurant chain using customer profiling. Sinofsky prompts the panel on whether this style of real-time custom analytics requires machine learning.16:28–19:17 · The host as informed peer 6/10 Predictive Analytics, Machine Learning, and Data Science Dhillon argues that predictive analytics and data scientists leveraging open-source tools like Berkeley's Spark represent the main shift in analytics. Sinofsky demonstrates specific domain knowledge by expanding on UC Berkeley's Amplab contributions. Stanek forcefully disagrees with Dhillon, arguing most businesses lack large enough data sets for true machine learning, leading to a direct argument between the two guests that Sinofsky humorously steps in to arbitrate.19:17–23:56 · The host as informed peer 7/10 The Persistence of Excel and Modern Data Pipelines Sinofsky uses Excel as a pivot topic, referencing his background leading Microsoft Office and making pivot tables accessible. Stanek notes that most modern analytics software effectively competes with Excel workbooks, prompting Dhillon to quickly clarify that Microsoft is a key partner and investor in SnapLogic.23:56–28:16 · The host as informed peer 5/10 On-Premise Data Realities, Cloud Migration, and Data Lakes Sinofsky asks how enterprises with on-premise systems of record can transition into modern cloud data architecture. Dhillon predicts that data lakes will eventually submerge traditional data warehouses, while Stanek and Moghe discuss regional compliance realities, data gravity, and hybrid cloud models.2:46–7:31 · Guest teaching 4/10 Demystifying Big Data and Volume vs. Agility Host Steven Sinofsky opens by framing the central topic of big data and asks guests to demystify what makes data 'big'. Prat Moghe reframes the premise by arguing big data is about agility and decision speed rather than pure petabyte volume. Sinofsky gently intervenes when asking about Kazina to make sure Moghe keeps it focused rather than delivering a canned product pitch.7:31–13:11 · Guest teaching 4/10 The Last Mile of Analytics and Field Delivery Roman Stanek explains GoodData's focus on the last mile of analytics for non-technical field workers. The guests engage in friendly banter, noting a internal bet about who would voice disagreement first. Sinofsky guides the discussion to explore how companies move from fixed weekly reporting to true data exploration.13:11–16:28 · Guest teaching 4/10 Rate of Business Change and Full-Stack Experiences Stanek highlights that the primary friction with corporate data is the rapid pace of business change exceeding IT turnaround times. Moghe illustrates this dynamic with an example of a fast-growing restaurant chain using customer profiling. Sinofsky prompts the panel on whether this style of real-time custom analytics requires machine learning.16:28–19:17 · Guest teaching 4/10 Predictive Analytics, Machine Learning, and Data Science Dhillon argues that predictive analytics and data scientists leveraging open-source tools like Berkeley's Spark represent the main shift in analytics. Sinofsky demonstrates specific domain knowledge by expanding on UC Berkeley's Amplab contributions. Stanek forcefully disagrees with Dhillon, arguing most businesses lack large enough data sets for true machine learning, leading to a direct argument between the two guests that Sinofsky humorously steps in to arbitrate.19:17–23:56 · Guest teaching 3/10 The Persistence of Excel and Modern Data Pipelines Sinofsky uses Excel as a pivot topic, referencing his background leading Microsoft Office and making pivot tables accessible. Stanek notes that most modern analytics software effectively competes with Excel workbooks, prompting Dhillon to quickly clarify that Microsoft is a key partner and investor in SnapLogic.23:56–28:16 · Guest teaching 3/10 On-Premise Data Realities, Cloud Migration, and Data Lakes Sinofsky asks how enterprises with on-premise systems of record can transition into modern cloud data architecture. Dhillon predicts that data lakes will eventually submerge traditional data warehouses, while Stanek and Moghe discuss regional compliance realities, data gravity, and hybrid cloud models.2:46–7:31 · Guest disagreement 2/10 Demystifying Big Data and Volume vs. Agility Host Steven Sinofsky opens by framing the central topic of big data and asks guests to demystify what makes data 'big'. Prat Moghe reframes the premise by arguing big data is about agility and decision speed rather than pure petabyte volume. Sinofsky gently intervenes when asking about Kazina to make sure Moghe keeps it focused rather than delivering a canned product pitch.7:31–13:11 · Guest disagreement 3/10 The Last Mile of Analytics and Field Delivery Roman Stanek explains GoodData's focus on the last mile of analytics for non-technical field workers. The guests engage in friendly banter, noting a internal bet about who would voice disagreement first. Sinofsky guides the discussion to explore how companies move from fixed weekly reporting to true data exploration.13:11–16:28 · Guest disagreement 2/10 Rate of Business Change and Full-Stack Experiences Stanek highlights that the primary friction with corporate data is the rapid pace of business change exceeding IT turnaround times. Moghe illustrates this dynamic with an example of a fast-growing restaurant chain using customer profiling. Sinofsky prompts the panel on whether this style of real-time custom analytics requires machine learning.16:28–19:17 · Guest disagreement 6/10 Predictive Analytics, Machine Learning, and Data Science Dhillon argues that predictive analytics and data scientists leveraging open-source tools like Berkeley's Spark represent the main shift in analytics. Sinofsky demonstrates specific domain knowledge by expanding on UC Berkeley's Amplab contributions. Stanek forcefully disagrees with Dhillon, arguing most businesses lack large enough data sets for true machine learning, leading to a direct argument between the two guests that Sinofsky humorously steps in to arbitrate.19:17–23:56 · Guest disagreement 4/10 The Persistence of Excel and Modern Data Pipelines Sinofsky uses Excel as a pivot topic, referencing his background leading Microsoft Office and making pivot tables accessible. Stanek notes that most modern analytics software effectively competes with Excel workbooks, prompting Dhillon to quickly clarify that Microsoft is a key partner and investor in SnapLogic.23:56–28:16 · Guest disagreement 2/10 On-Premise Data Realities, Cloud Migration, and Data Lakes Sinofsky asks how enterprises with on-premise systems of record can transition into modern cloud data architecture. Dhillon predicts that data lakes will eventually submerge traditional data warehouses, while Stanek and Moghe discuss regional compliance realities, data gravity, and hybrid cloud models.2:46–7:31 · The host pushing back 3/10 Demystifying Big Data and Volume vs. Agility Host Steven Sinofsky opens by framing the central topic of big data and asks guests to demystify what makes data 'big'. Prat Moghe reframes the premise by arguing big data is about agility and decision speed rather than pure petabyte volume. Sinofsky gently intervenes when asking about Kazina to make sure Moghe keeps it focused rather than delivering a canned product pitch.7:31–13:11 · The host pushing back 3/10 The Last Mile of Analytics and Field Delivery Roman Stanek explains GoodData's focus on the last mile of analytics for non-technical field workers. The guests engage in friendly banter, noting a internal bet about who would voice disagreement first. Sinofsky guides the discussion to explore how companies move from fixed weekly reporting to true data exploration.13:11–16:28 · The host pushing back 3/10 Rate of Business Change and Full-Stack Experiences Stanek highlights that the primary friction with corporate data is the rapid pace of business change exceeding IT turnaround times. Moghe illustrates this dynamic with an example of a fast-growing restaurant chain using customer profiling. Sinofsky prompts the panel on whether this style of real-time custom analytics requires machine learning.16:28–19:17 · The host pushing back 4/10 Predictive Analytics, Machine Learning, and Data Science Dhillon argues that predictive analytics and data scientists leveraging open-source tools like Berkeley's Spark represent the main shift in analytics. Sinofsky demonstrates specific domain knowledge by expanding on UC Berkeley's Amplab contributions. Stanek forcefully disagrees with Dhillon, arguing most businesses lack large enough data sets for true machine learning, leading to a direct argument between the two guests that Sinofsky humorously steps in to arbitrate.19:17–23:56 · The host pushing back 5/10 The Persistence of Excel and Modern Data Pipelines Sinofsky uses Excel as a pivot topic, referencing his background leading Microsoft Office and making pivot tables accessible. Stanek notes that most modern analytics software effectively competes with Excel workbooks, prompting Dhillon to quickly clarify that Microsoft is a key partner and investor in SnapLogic.23:56–28:16 · The host pushing back 3/10 On-Premise Data Realities, Cloud Migration, and Data Lakes Sinofsky asks how enterprises with on-premise systems of record can transition into modern cloud data architecture. Dhillon predicts that data lakes will eventually submerge traditional data warehouses, while Stanek and Moghe discuss regional compliance realities, data gravity, and hybrid cloud models.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 18:26 Guest clash on data science and machine learning relevance

Roman Stanek directly interrupts and disagrees with Gaurav Dhillon's premise that data scientists are essential, prompting Dhillon to sharply respond that Stanek's view reduces analytics to simple reports.

Hardest push from the host ▶ 4:25 Host stopping a guest's company pitch

Steven Sinofsky steps in as Prat Moghe starts explaining Kazina, cautioning him not to turn his answer into a founder elevator pitch.

Biggest teaching moment ▶ 3:14 Prat reframing the definition of Big Data

Prat Moghe corrects the popular notion that big data is defined by petabyte volume, educating the room that it is actually about data agility and decision-making speed.

The host holds their own ▶ 20:42 Host bringing personal authority on Excel pivot tables

Steven Sinofsky asserts his deep technical background by reminding the guests of his former work managing Microsoft's development of Excel pivot table usability.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Demystifying Big Data and Volume vs. Agility 4423 Host Steven Sinofsky opens by framing the central topic of big data and asks guests to demystify what makes data 'big'. Prat Moghe reframes the premise by arguing big data is about agility and decision speed rather than pure petabyte volume. Sinofsky gently intervenes when asking about Kazina to make sure Moghe keeps it focused rather than delivering a canned product pitch.
The Last Mile of Analytics and Field Delivery 4433 Roman Stanek explains GoodData's focus on the last mile of analytics for non-technical field workers. The guests engage in friendly banter, noting a internal bet about who would voice disagreement first. Sinofsky guides the discussion to explore how companies move from fixed weekly reporting to true data exploration.
Rate of Business Change and Full-Stack Experiences 4423 Stanek highlights that the primary friction with corporate data is the rapid pace of business change exceeding IT turnaround times. Moghe illustrates this dynamic with an example of a fast-growing restaurant chain using customer profiling. Sinofsky prompts the panel on whether this style of real-time custom analytics requires machine learning.
Predictive Analytics, Machine Learning, and Data Science 6464 Dhillon argues that predictive analytics and data scientists leveraging open-source tools like Berkeley's Spark represent the main shift in analytics. Sinofsky demonstrates specific domain knowledge by expanding on UC Berkeley's Amplab contributions. Stanek forcefully disagrees with Dhillon, arguing most businesses lack large enough data sets for true machine learning, leading to a direct argument between the two guests that Sinofsky humorously steps in to arbitrate.
The Persistence of Excel and Modern Data Pipelines 7345 Sinofsky uses Excel as a pivot topic, referencing his background leading Microsoft Office and making pivot tables accessible. Stanek notes that most modern analytics software effectively competes with Excel workbooks, prompting Dhillon to quickly clarify that Microsoft is a key partner and investor in SnapLogic.
On-Premise Data Realities, Cloud Migration, and Data Lakes 5323 Sinofsky asks how enterprises with on-premise systems of record can transition into modern cloud data architecture. Dhillon predicts that data lakes will eventually submerge traditional data warehouses, while Stanek and Moghe discuss regional compliance realities, data gravity, and hybrid cloud models.

Statements from this episode (16)

Insight
Moghe: Big data is defined by decision speed, not storage volume
“I sort of define big data as it's a mindset. It's about being really fast about using data to make decisions. So it's not just about petabytes of data. It's about, you know, how fast can you leverage data to create business outcomes and so that mindset is what…”
Prat Moghe Jan 2, 2019 ▶ 4:02
Insight
Dhillon: Big data's key innovation is automated cross-source data correlation
“In big data, to me, the fundamental breakthrough is providing information from multiple places and producing insights where the data finds the data.”
Gaurav Dhillon Jan 2, 2019 ▶ 6:09
Opinion
Stanek: Hadoop and data warehouses are where data goes to die
“With all the investment in Hadoop and this and Hadoop that, you know, most companies are still data bankrupt. You know, Hadoop or Data Warehouse or whatever is a place where data goes to die”
Roman Stanek Jan 2, 2019 ▶ 7:47
Assertion Not checkable as stated
Stanek: GoodData serves 500,000 white-labeled users
“We have about half million users, and very few of them actually know they use good data because they see somebody else's logo”
Roman Stanek Jan 2, 2019 ▶ 9:09
Insight
Moghe: Big data projects fail when data collection precedes business goals
“When you looked at many big data projects, the ones that fail, Are ones where people have taken this approach of saying, I want to collect all the data, and then I want to figure out what questions I can ask. I want to look for hidden patterns, as opposed to p…”
Prat Moghe Jan 2, 2019 ▶ 10:40
Insight
Stanek: Business rate of change is enterprise data's biggest challenge
“I actually believe that the biggest problem of data, not big data, small data, any data, is the rate of change of business.”
Roman Stanek Jan 2, 2019 ▶ 13:13
Prediction Not checkable as stated
Moghe: Enterprise data is shifting to verticalized full-stack applications
“So there's a whole new breed of, ah, we heard this morning, like the full stack, you know, sort of the full stack app. Like verticalized, experiences, everything that matters. I think that's where it's going. I think where it's going is all that data gets surf…”
Prat Moghe Jan 2, 2019 ▶ 16:03
Insight
Dhillon: Traditional business intelligence yields diminishing marginal returns
“The traditional rear view mirror view of business intelligence has some element of return, but it's also at some point been well done. There are ways to improve that, but we're getting to a point of diminishing marginal returns on that.”
Gaurav Dhillon Jan 2, 2019 ▶ 17:00
Assertion Not checkable as stated
Dhillon: Data science roles have gone mainstream beyond finance
“This is now widespread outside of financial services. You know, Goldman Sachs, Morgan Stanley always had quant jocks. Now everybody has quant jocks.”
Gaurav Dhillon Jan 2, 2019 ▶ 18:05
Insight
Stanek: Most companies lack sufficient data scale for in-house machine learning
“Most companies are not big enough to have big enough sample for machine learning. You know, big banks. What makes Google Google? What makes Amazon Amazon is that the data sample is so big that you can actually really learn from it. Typical company would look a…”
Roman Stanek Jan 2, 2019 ▶ 18:37
Insight
Stanek: Cloud analytics unlocks machine learning by aggregating cross-company data
“I actually believe that that's why analytics done in a cloud is actually a lot of value, because we see data across tens of thousands of companies, and we can actually do machine learning from, you know, massive data sets that individually don't actually mean …”
Roman Stanek Jan 2, 2019 ▶ 18:54
Insight
Stanek: Business users prefer spreadsheet interfaces over Spark and Hadoop
“Some of the most frequently used kind of data analytics tools extremely basic, because they actually look and feel like sheet of paper, like two-dimensional sheet of paper, and you know, so that's the problem with analytics, that on one hand, we have, you know…”
Roman Stanek Jan 2, 2019 ▶ 20:55
Assertion Not checkable as stated
Moghe: Spark and Hadoop do not replace existing data warehouses
“Spark doesn't subsume data warehousing. Hadoop doesn't subsume, you know, streaming. So they're just like different technologies for different jobs.”
Prat Moghe Jan 2, 2019 ▶ 23:39
Prediction Not checkable as stated
Dhillon: Data lakes will eventually drown out traditional data warehouses
“The rising tide of the data lake, we think, will drown out the data warehouse in the fullness of time.”
Gaurav Dhillon Jan 2, 2019 ▶ 25:17
Opinion
Stanek: Primary enterprise data may never move to the cloud
“We don't have time to wait for companies to move the data, the primary data to the cloud, that may not even happen”
Roman Stanek Jan 2, 2019 ▶ 26:35
Prediction Not checkable as stated
Stanek: European data balkanization will create severe challenges for cloud vendors
“Instead of putting one data center for, you know, My, all, all audience, you know, user base, we need to build multiple data centers, and that kind of fragmentation and balkanization of data will continue, and that's going to be more and more difficult for clo…”
Roman Stanek Jan 2, 2019 ▶ 28:34
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.