Dec 17, 2015 · 26m · mad

The Acceleration of Innovation in Big Data w/ Stefan Groschupf, Datameer

Stefan Groschupf · 21m spoken Matt Turck · 52s spoken Tony Baer · 10s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Datameer Founder and CEO Stefan Groschupf presents a compelling keynote on how exponential hardware acceleration and open-source innovation are disrupting enterprise big data architectures. He argues that modern infrastructure must transition away from rigid schemas and static software stacks toward dynamic resource orchestration, automated deep learning, and user-driven business intelligence.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.6% of the talking time here. How this is scored →

Matt as informed peer 0.8 Guest teaching 1.8 Guest disagreement 2.8 Matt pushing back 0.4
05100:0010:0020:000:31–3:39 · Matt as informed peer 0/10 Founder Background and the History of Nutch This is a solo presentation segment where the host does not speak. Stefan introduces himself with a playful and slightly provocative tone, warning the audience that he intends to challenge big data assumptions.3:39–9:17 · Matt as informed peer 0/10 Acceleration of Complexity in Big Data Ecosystems The guest delivers a presentation monologue outlining Moore's Law and the history of Hadoop. He forcefully argues that introducing SQL on top of Hadoop was one of the most unfortunate mistakes in the industry.9:17–11:27 · Matt as informed peer 0/10 Analyzing Disruption: Hadoop, Spark, and Flink In this presentation segment, Stefan challenges the audience's perception of Spark's disruptiveness compared to Hadoop and highlights Flink's performance advantages.11:27–16:01 · Matt as informed peer 0/10 Rethinking the Data Stack: Data Center OS as the Next Shift Stefan presents his architectural thesis about Mesos and Yarn acting as data center operating systems. The host remains off-mic during this monologue section.16:01–26:40 · Matt as informed peer 4/10 Deep Learning and the Future Role of Data Scientists Host Matt Turck joins the conversation, showing solid research by referencing specific white papers and case studies from Datameer's website. The dynamic remains collaborative and polite during the Q&A session.0:31–3:39 · Guest teaching 1/10 Founder Background and the History of Nutch This is a solo presentation segment where the host does not speak. Stefan introduces himself with a playful and slightly provocative tone, warning the audience that he intends to challenge big data assumptions.3:39–9:17 · Guest teaching 2/10 Acceleration of Complexity in Big Data Ecosystems The guest delivers a presentation monologue outlining Moore's Law and the history of Hadoop. He forcefully argues that introducing SQL on top of Hadoop was one of the most unfortunate mistakes in the industry.9:17–11:27 · Guest teaching 2/10 Analyzing Disruption: Hadoop, Spark, and Flink In this presentation segment, Stefan challenges the audience's perception of Spark's disruptiveness compared to Hadoop and highlights Flink's performance advantages.11:27–16:01 · Guest teaching 2/10 Rethinking the Data Stack: Data Center OS as the Next Shift Stefan presents his architectural thesis about Mesos and Yarn acting as data center operating systems. The host remains off-mic during this monologue section.16:01–26:40 · Guest teaching 2/10 Deep Learning and the Future Role of Data Scientists Host Matt Turck joins the conversation, showing solid research by referencing specific white papers and case studies from Datameer's website. The dynamic remains collaborative and polite during the Q&A session.0:31–3:39 · Guest disagreement 3/10 Founder Background and the History of Nutch This is a solo presentation segment where the host does not speak. Stefan introduces himself with a playful and slightly provocative tone, warning the audience that he intends to challenge big data assumptions.3:39–9:17 · Guest disagreement 4/10 Acceleration of Complexity in Big Data Ecosystems The guest delivers a presentation monologue outlining Moore's Law and the history of Hadoop. He forcefully argues that introducing SQL on top of Hadoop was one of the most unfortunate mistakes in the industry.9:17–11:27 · Guest disagreement 3/10 Analyzing Disruption: Hadoop, Spark, and Flink In this presentation segment, Stefan challenges the audience's perception of Spark's disruptiveness compared to Hadoop and highlights Flink's performance advantages.11:27–16:01 · Guest disagreement 2/10 Rethinking the Data Stack: Data Center OS as the Next Shift Stefan presents his architectural thesis about Mesos and Yarn acting as data center operating systems. The host remains off-mic during this monologue section.16:01–26:40 · Guest disagreement 2/10 Deep Learning and the Future Role of Data Scientists Host Matt Turck joins the conversation, showing solid research by referencing specific white papers and case studies from Datameer's website. The dynamic remains collaborative and polite during the Q&A session.0:31–3:39 · Matt pushing back 0/10 Founder Background and the History of Nutch This is a solo presentation segment where the host does not speak. Stefan introduces himself with a playful and slightly provocative tone, warning the audience that he intends to challenge big data assumptions.3:39–9:17 · Matt pushing back 0/10 Acceleration of Complexity in Big Data Ecosystems The guest delivers a presentation monologue outlining Moore's Law and the history of Hadoop. He forcefully argues that introducing SQL on top of Hadoop was one of the most unfortunate mistakes in the industry.9:17–11:27 · Matt pushing back 0/10 Analyzing Disruption: Hadoop, Spark, and Flink In this presentation segment, Stefan challenges the audience's perception of Spark's disruptiveness compared to Hadoop and highlights Flink's performance advantages.11:27–16:01 · Matt pushing back 0/10 Rethinking the Data Stack: Data Center OS as the Next Shift Stefan presents his architectural thesis about Mesos and Yarn acting as data center operating systems. The host remains off-mic during this monologue section.16:01–26:40 · Matt pushing back 2/10 Deep Learning and the Future Role of Data Scientists Host Matt Turck joins the conversation, showing solid research by referencing specific white papers and case studies from Datameer's website. The dynamic remains collaborative and polite during the Q&A session.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 11% · guest 89%15:00 · Matt 11% · guest 89%18:00 · Matt 6.2% · guest 93.8%18:00 · Matt 6.2% · guest 93.8%21:00 · Matt 14.7% · guest 85.3%21:00 · Matt 14.7% · guest 85.3%24:00 · Matt 2% · guest 98%24:00 · Matt 2% · guest 98%
Sharpest disagreement ▶ 16:30 Advising against studying data science

Stefan forcefully challenges conventional career wisdom by telling the audience to rethink choosing data science due to deep learning automation.

Hardest push from Matt ▶ 21:20 Challenging the IoT data narrative

Matt presses the guest on whether IoT data is actually structurally unique or just standard data with elevated buzzword marketing.

Biggest teaching moment ▶ 21:45 Explaining IoT schema complexity

Stefan details how firmware updates create evolving JSON structures, illustrating why schema-on-read architectures are necessary.

Matt holds his own ▶ 21:20 Citing IoT white paper

Matt demonstrates preparation and domain expertise by citing a specific white paper from Datameer's website on IoT analytics.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Founder Background and the History of Nutch 0130 This is a solo presentation segment where the host does not speak. Stefan introduces himself with a playful and slightly provocative tone, warning the audience that he intends to challenge big data assumptions.
Acceleration of Complexity in Big Data Ecosystems 0240 The guest delivers a presentation monologue outlining Moore's Law and the history of Hadoop. He forcefully argues that introducing SQL on top of Hadoop was one of the most unfortunate mistakes in the industry.
Analyzing Disruption: Hadoop, Spark, and Flink 0230 In this presentation segment, Stefan challenges the audience's perception of Spark's disruptiveness compared to Hadoop and highlights Flink's performance advantages.
Rethinking the Data Stack: Data Center OS as the Next Shift 0220 Stefan presents his architectural thesis about Mesos and Yarn acting as data center operating systems. The host remains off-mic during this monologue section.
Deep Learning and the Future Role of Data Scientists 4222 Host Matt Turck joins the conversation, showing solid research by referencing specific white papers and case studies from Datameer's website. The dynamic remains collaborative and polite during the Q&A session.

Statements from this episode (7)

Opinion
Groschupf: SQL on top of Hadoop was an unfortunate development
“Well, it's unfortunate, what I think is one of the most unfortunate thing that happened in the Hadoop space is kind of the introduction of SQL on top of Hadoop.”
Stefan Groschupf Dec 17, 2015 ▶ 5:42
Insight
Groschupf: Any database schema built today is outdated tomorrow
“Whatever schema you're building today, it's outdated tomorrow.”
Stefan Groschupf Dec 17, 2015 ▶ 9:04
Assertion Not checkable as stated
Stefan Groschupf: Apache Flink is already faster than Spark
“And what's really interesting is Flink is already faster than Spark.”
Stefan Groschupf Dec 17, 2015 ▶ 10:35
Prediction Not checkable as stated
Groschupf: Data center operating systems will be the next Hadoop killer
“That's a data center OS, and I really think that's the next Hadoop killer.”
Stefan Groschupf Dec 17, 2015 ▶ 13:21
What-if
Groschupf: New computation frameworks should be built on Mesosphere, not YARN
“And if I would have to rewrite kind of the code I wrote in 2006, I would not necessarily write it on Yarn. I would maybe write a whole new computation framework on Mesosphere.”
Stefan Groschupf Dec 17, 2015 ▶ 14:21
Insight
Groschupf: Aspiring students should rethink choosing data science as a career
“So if you have to make a career choice to study data science today, I would rethink that. And I'm serious.”
Stefan Groschupf Dec 17, 2015 ▶ 16:52
Opinion
Groschupf: 'Big data' is a corporate buzzword created to make money
“I think big data is a big buzzword from big company to make big money.”
Stefan Groschupf Dec 17, 2015 ▶ 21:29
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.