Jan 2, 2019 · 27m · a16z

a16z Podcast | Making Sense of Big Data, Machine Learning, and Deep Learning

Christopher Nguyen · 18m spoken Sonal Chokshi · 6m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, host Sonal converses with Adetao CEO Christopher Nguyen about the true definitions and business impacts of big data, machine learning, and deep learning. They explore the evolution of the enterprise technology stack, the transition to predictive intelligence, and how machine learning mirrors human cognitive development.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The host holds 23.8% of the talking time here. How this is scored →

The host as informed peer 5.6 Guest teaching 4.9 Guest disagreement 1.9 The host pushing back 5.6
05100:0010:0020:000:34–3:32 · The host as informed peer 5/10 Redefining Big Data Through Machine Learning Sonal challenges Christopher's assertion that machine learning is the primary reason for big data, noting it sounds counterintuitive to industry consensus. Christopher counters by rejecting the standard volume and velocity definitions in favor of learning thresholds.3:32–7:54 · The host as informed peer 4/10 Comparing Machine Learning to Human Experience and Wisdom Sonal asks clarifying questions comparing machine learning exception handling to human child development and business intelligence. Christopher explains how predictive machine learning differs from traditional backward-looking BI aggregations.7:54–11:23 · The host as informed peer 7/10 The Big Data Storage Stack: Hadoop and Distributed Infrastructure Sonal demonstrates strong technical awareness by citing the Berkeley Data Analytics Stack (BADAS) and a16z's investment portfolio. She pushes Christopher to clarify whether commodity hardware adoption was driven by software architecture or declining hardware costs.11:23–15:27 · The host as informed peer 6/10 Big Compute: MapReduce vs. Apache Spark Sonal posits that MapReduce's slowness stems from its two-function constraint, but Christopher corrects her by explaining it was intentionally engineered for fault tolerance. Sonal acknowledges the insight and pivots to comparing Spark's in-memory latency advantages.15:27–17:52 · The host as informed peer 5/10 Why Speed Matters and the Missing Big Apps Layer Sonal explicitly plays devil's advocate, asking why fast processing speed actually matters for businesses. Christopher explains the ergonomic five-second rule and how real-time speed fundamentally alters decision workflows.17:52–23:14 · The host as informed peer 6/10 Machine Learning as an Application Property and Negative Latency Sonal presses Christopher multiple times when his explanations of application-level machine learning remain abstract, demanding concrete value. Christopher responds with Larry Page's vision of negative latency and predictive anticipation.23:14–26:55 · The host as informed peer 6/10 Deep Learning, Species Intelligence, and Enterprise Advantage Sonal contextualizes deep learning within broader machine learning discourse and repeatedly demands concrete commercial outcomes beyond academic fascination. Christopher highlights enterprise data competitiveness alongside long-term species exploration.0:34–3:32 · Guest teaching 5/10 Redefining Big Data Through Machine Learning Sonal challenges Christopher's assertion that machine learning is the primary reason for big data, noting it sounds counterintuitive to industry consensus. Christopher counters by rejecting the standard volume and velocity definitions in favor of learning thresholds.3:32–7:54 · Guest teaching 4/10 Comparing Machine Learning to Human Experience and Wisdom Sonal asks clarifying questions comparing machine learning exception handling to human child development and business intelligence. Christopher explains how predictive machine learning differs from traditional backward-looking BI aggregations.7:54–11:23 · Guest teaching 4/10 The Big Data Storage Stack: Hadoop and Distributed Infrastructure Sonal demonstrates strong technical awareness by citing the Berkeley Data Analytics Stack (BADAS) and a16z's investment portfolio. She pushes Christopher to clarify whether commodity hardware adoption was driven by software architecture or declining hardware costs.11:23–15:27 · Guest teaching 7/10 Big Compute: MapReduce vs. Apache Spark Sonal posits that MapReduce's slowness stems from its two-function constraint, but Christopher corrects her by explaining it was intentionally engineered for fault tolerance. Sonal acknowledges the insight and pivots to comparing Spark's in-memory latency advantages.15:27–17:52 · Guest teaching 5/10 Why Speed Matters and the Missing Big Apps Layer Sonal explicitly plays devil's advocate, asking why fast processing speed actually matters for businesses. Christopher explains the ergonomic five-second rule and how real-time speed fundamentally alters decision workflows.17:52–23:14 · Guest teaching 5/10 Machine Learning as an Application Property and Negative Latency Sonal presses Christopher multiple times when his explanations of application-level machine learning remain abstract, demanding concrete value. Christopher responds with Larry Page's vision of negative latency and predictive anticipation.23:14–26:55 · Guest teaching 4/10 Deep Learning, Species Intelligence, and Enterprise Advantage Sonal contextualizes deep learning within broader machine learning discourse and repeatedly demands concrete commercial outcomes beyond academic fascination. Christopher highlights enterprise data competitiveness alongside long-term species exploration.0:34–3:32 · Guest disagreement 4/10 Redefining Big Data Through Machine Learning Sonal challenges Christopher's assertion that machine learning is the primary reason for big data, noting it sounds counterintuitive to industry consensus. Christopher counters by rejecting the standard volume and velocity definitions in favor of learning thresholds.3:32–7:54 · Guest disagreement 1/10 Comparing Machine Learning to Human Experience and Wisdom Sonal asks clarifying questions comparing machine learning exception handling to human child development and business intelligence. Christopher explains how predictive machine learning differs from traditional backward-looking BI aggregations.7:54–11:23 · Guest disagreement 1/10 The Big Data Storage Stack: Hadoop and Distributed Infrastructure Sonal demonstrates strong technical awareness by citing the Berkeley Data Analytics Stack (BADAS) and a16z's investment portfolio. She pushes Christopher to clarify whether commodity hardware adoption was driven by software architecture or declining hardware costs.11:23–15:27 · Guest disagreement 2/10 Big Compute: MapReduce vs. Apache Spark Sonal posits that MapReduce's slowness stems from its two-function constraint, but Christopher corrects her by explaining it was intentionally engineered for fault tolerance. Sonal acknowledges the insight and pivots to comparing Spark's in-memory latency advantages.15:27–17:52 · Guest disagreement 2/10 Why Speed Matters and the Missing Big Apps Layer Sonal explicitly plays devil's advocate, asking why fast processing speed actually matters for businesses. Christopher explains the ergonomic five-second rule and how real-time speed fundamentally alters decision workflows.17:52–23:14 · Guest disagreement 2/10 Machine Learning as an Application Property and Negative Latency Sonal presses Christopher multiple times when his explanations of application-level machine learning remain abstract, demanding concrete value. Christopher responds with Larry Page's vision of negative latency and predictive anticipation.23:14–26:55 · Guest disagreement 1/10 Deep Learning, Species Intelligence, and Enterprise Advantage Sonal contextualizes deep learning within broader machine learning discourse and repeatedly demands concrete commercial outcomes beyond academic fascination. Christopher highlights enterprise data competitiveness alongside long-term species exploration.0:34–3:32 · The host pushing back 6/10 Redefining Big Data Through Machine Learning Sonal challenges Christopher's assertion that machine learning is the primary reason for big data, noting it sounds counterintuitive to industry consensus. Christopher counters by rejecting the standard volume and velocity definitions in favor of learning thresholds.3:32–7:54 · The host pushing back 3/10 Comparing Machine Learning to Human Experience and Wisdom Sonal asks clarifying questions comparing machine learning exception handling to human child development and business intelligence. Christopher explains how predictive machine learning differs from traditional backward-looking BI aggregations.7:54–11:23 · The host pushing back 5/10 The Big Data Storage Stack: Hadoop and Distributed Infrastructure Sonal demonstrates strong technical awareness by citing the Berkeley Data Analytics Stack (BADAS) and a16z's investment portfolio. She pushes Christopher to clarify whether commodity hardware adoption was driven by software architecture or declining hardware costs.11:23–15:27 · The host pushing back 4/10 Big Compute: MapReduce vs. Apache Spark Sonal posits that MapReduce's slowness stems from its two-function constraint, but Christopher corrects her by explaining it was intentionally engineered for fault tolerance. Sonal acknowledges the insight and pivots to comparing Spark's in-memory latency advantages.15:27–17:52 · The host pushing back 7/10 Why Speed Matters and the Missing Big Apps Layer Sonal explicitly plays devil's advocate, asking why fast processing speed actually matters for businesses. Christopher explains the ergonomic five-second rule and how real-time speed fundamentally alters decision workflows.17:52–23:14 · The host pushing back 8/10 Machine Learning as an Application Property and Negative Latency Sonal presses Christopher multiple times when his explanations of application-level machine learning remain abstract, demanding concrete value. Christopher responds with Larry Page's vision of negative latency and predictive anticipation.23:14–26:55 · The host pushing back 6/10 Deep Learning, Species Intelligence, and Enterprise Advantage Sonal contextualizes deep learning within broader machine learning discourse and repeatedly demands concrete commercial outcomes beyond academic fascination. Christopher highlights enterprise data competitiveness alongside long-term species exploration.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 38.6% · guest 61.4%0:00 · the host 38.6% · guest 61.4%3:00 · the host 20.6% · guest 79.4%3:00 · the host 20.6% · guest 79.4%6:00 · the host 31% · guest 69%6:00 · the host 31% · guest 69%9:00 · the host 25.1% · guest 74.9%9:00 · the host 25.1% · guest 74.9%12:00 · the host 10% · guest 90%12:00 · the host 10% · guest 90%15:00 · the host 23.4% · guest 76.6%15:00 · the host 23.4% · guest 76.6%18:00 · the host 35.8% · guest 64.2%18:00 · the host 35.8% · guest 64.2%21:00 · the host 10.8% · guest 89.2%21:00 · the host 10.8% · guest 89.2%24:00 · the host 18.6% · guest 81.4%24:00 · the host 18.6% · guest 81.4%27:00 · the host 100% · guest 0%27:00 · the host 100% · guest 0%
Sharpest disagreement ▶ 0:52 Rejecting standard Big Data definitions

Christopher explicitly rejects the widely accepted V's definition of big data, arguing that focusing on volume and velocity misses the true purpose.

Hardest push from the host ▶ 20:05 Demanding practical clarity over abstraction

Sonal refuses to accept high-level generalities about machine learning in applications, explicitly telling the guest that the business utility remains unclear.

Biggest teaching moment ▶ 12:29 Explaining MapReduce intentional slowness

When Sonal guesses MapReduce slowness was caused by Map and Reduce functional constraints, Christopher corrects her, explaining it was intentionally engineered for disk-bound fault tolerance.

The host holds their own ▶ 7:54 Demonstrating deep analytics stack knowledge

Sonal demonstrates her domain knowledge by referencing the Berkeley Data Analytics Stack (BADAS) and asking specific structural questions about infrastructure layers.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Redefining Big Data Through Machine Learning 5546 Sonal challenges Christopher's assertion that machine learning is the primary reason for big data, noting it sounds counterintuitive to industry consensus. Christopher counters by rejecting the standard volume and velocity definitions in favor of learning thresholds.
Comparing Machine Learning to Human Experience and Wisdom 4413 Sonal asks clarifying questions comparing machine learning exception handling to human child development and business intelligence. Christopher explains how predictive machine learning differs from traditional backward-looking BI aggregations.
The Big Data Storage Stack: Hadoop and Distributed Infrastructure 7415 Sonal demonstrates strong technical awareness by citing the Berkeley Data Analytics Stack (BADAS) and a16z's investment portfolio. She pushes Christopher to clarify whether commodity hardware adoption was driven by software architecture or declining hardware costs.
Big Compute: MapReduce vs. Apache Spark 6724 Sonal posits that MapReduce's slowness stems from its two-function constraint, but Christopher corrects her by explaining it was intentionally engineered for fault tolerance. Sonal acknowledges the insight and pivots to comparing Spark's in-memory latency advantages.
Why Speed Matters and the Missing Big Apps Layer 5527 Sonal explicitly plays devil's advocate, asking why fast processing speed actually matters for businesses. Christopher explains the ergonomic five-second rule and how real-time speed fundamentally alters decision workflows.
Machine Learning as an Application Property and Negative Latency 6528 Sonal presses Christopher multiple times when his explanations of application-level machine learning remain abstract, demanding concrete value. Christopher responds with Larry Page's vision of negative latency and predictive anticipation.
Deep Learning, Species Intelligence, and Enterprise Advantage 6416 Sonal contextualizes deep learning within broader machine learning discourse and repeatedly demands concrete commercial outcomes beyond academic fascination. Christopher highlights enterprise data competitiveness alongside long-term species exploration.

Statements from this episode (11)

Insight
Nguyen: Machine Learning Is the Primary Purpose of Big Data
“So it turns out the reason for big data is machine learning.”
Christopher Nguyen Jan 2, 2019 ▶ 1:30
Insight
Nguyen: Big Data Is Defined by Learning Thresholds, Not Volume
“So I like something that Peter Norvig, the director of research at Google said when he referred to big data, he says, big data is not just quantitatively different, but it's qualitatively different. In other words, there's something that happens when you have …”
Christopher Nguyen Jan 2, 2019 ▶ 1:42
Insight
Nguyen: Machine Learning Mirrors Human Learning From Experience
“And the way I think about big data is when machines learn from big data is very much like human beings learn from life experiences.”
Christopher Nguyen Jan 2, 2019 ▶ 3:24
Insight
Nguyen: Human Intuition Functions Like Parameters in a Machine Learning Model
“Well, what we think of as intuition are actually, you can think of as parameters inside a machine learning model.”
Christopher Nguyen Jan 2, 2019 ▶ 5:09
Insight
Nguyen: Modern Business Intelligence Uses ML to Predict Unknowns
“You can think of business intelligence going forward as the ability to apply machine learning algorithms to big data, and not just look at past questions, but also future questions, or asking to predict the unknowns from the knowns.”
Christopher Nguyen Jan 2, 2019 ▶ 6:33
Insight
Nguyen: Big Data Progress Is Driven by Cheaper Tech, Not Smarter People
“We don't necessarily get smarter over time. It's just that certain technologies get cheaper. They get, they become more available. So machine learning algorithms have always been around. The data that exists that you could collect has always been around. But i…”
Christopher Nguyen Jan 2, 2019 ▶ 7:20
Assertion Supported
Nguyen: MapReduce Was Intentionally Designed for Reliability Over Speed
“Interestingly, a lot of people may not realize that MapReduce was designed to be slow.”
Christopher Nguyen Jan 2, 2019 ▶ 12:37
What-if
Nguyen: Apache Spark Would Have Failed Earlier Due to Memory Costs
“Now Spark, if it was created six, five, six years before its time would have completely failed because memory was so much more expensive.”
Christopher Nguyen Jan 2, 2019 ▶ 14:45
Insight
Nguyen: Users Abandon Software Tasks if Latency Exceeds Five Seconds
“We had a phrase we call the five second barrier. And if the user can't get something done, you know, within five seconds, they won't ever do it. It's not like they'll do it at, you know, at twice the latency.”
Christopher Nguyen Jan 2, 2019 ▶ 16:19
Prediction Not checkable as stated
Nguyen: Machine Learning Will Become a Feature of Every Application
“What you will see is that all of this machine learning will be a property of every application.”
Christopher Nguyen Jan 2, 2019 ▶ 19:09
Prediction Not checkable as stated
Nguyen: Users Will Soon Expect Machines to Learn Like Human Colleagues
“Well, I claim that there will be a day very soon when then you will feel that about the machines you work with. In other words, you would expect that to be a property of all these machines.”
Christopher Nguyen Jan 2, 2019 ▶ 19:57
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.