Nov 9, 2016 · 23m · mad

Lessons Learned from Advanced Data Science Orgs // Domino Data Lab [FirstMark's Data Driven]

Nick (Domino Data Lab) · 17m spoken Matt Turck · 39s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this presentation at FirstMark's Data Driven NYC, Domino Data Lab CEO Nick Elprin outlines best practices for scaling enterprise data science organizations through collaboration, agility, and reproducibility. He combines theoretical organizational insights, product strategy, a live platform demonstration, and an interactive Q&A session on modern quantitative research tooling.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.6% of the talking time here. How this is scored →

Matt as informed peer 0.9 Guest teaching 2.9 Guest disagreement 1.0 Matt pushing back 0.1
05100:0010:0020:000:16–5:08 · Matt as informed peer 0/10 Presentation Objectives and Goals Nick presents a solo talk outlining data science myths versus best practices in advanced analytics organizations. The host does not speak during this monologue segment, so host scores are zero.5:08–8:10 · Matt as informed peer 0/10 Core Themes: Collaboration, Reproducibility, and Agility Nick continues his presentation on collaboration, reproducibility, and agility, explaining why code versioning alone is insufficient for quantitative research. Host-side scores remain zero for this uninterrupted monologue.8:10–14:12 · Matt as informed peer 0/10 Product Strategy: Incentivizing Best Practices Bottom-Up Nick demonstrates the Domino platform and explains how to incentivize best practices bottom-up. As another presentation monologue, host-side metrics are strictly zero.14:12–16:56 · Matt as informed peer 3/10 Analytical Life Cycle and Data Science Tooling Trends Matt Turck steps in to praise the demo and ask informed questions about the analytical lifecycle and emerging data science tooling trends. Nick reframes market hype around Spark versus actual value extraction.16:56–19:53 · Matt as informed peer 1/10 Q&A: Data Versioning and Aggregation Strategies Nick answers audience questions regarding file versioning scale and compliance walls in financial institutions. The host simply moderates the Q&A.19:53–21:59 · Matt as informed peer 1/10 Q&A: General Adoption vs. Specialized Workflows Nick addresses an audience question regarding non-data users adopting Domino, contrasting general tools like SharePoint with specialized quantitative research workflows.21:59–23:29 · Matt as informed peer 1/10 Q&A: Scientific Reproducibility and Event Conclusion Nick discusses reproducibility as a commercial driver driven by regulatory compliance before the host wraps up the session.0:16–5:08 · Guest teaching 2/10 Presentation Objectives and Goals Nick presents a solo talk outlining data science myths versus best practices in advanced analytics organizations. The host does not speak during this monologue segment, so host scores are zero.5:08–8:10 · Guest teaching 3/10 Core Themes: Collaboration, Reproducibility, and Agility Nick continues his presentation on collaboration, reproducibility, and agility, explaining why code versioning alone is insufficient for quantitative research. Host-side scores remain zero for this uninterrupted monologue.8:10–14:12 · Guest teaching 2/10 Product Strategy: Incentivizing Best Practices Bottom-Up Nick demonstrates the Domino platform and explains how to incentivize best practices bottom-up. As another presentation monologue, host-side metrics are strictly zero.14:12–16:56 · Guest teaching 3/10 Analytical Life Cycle and Data Science Tooling Trends Matt Turck steps in to praise the demo and ask informed questions about the analytical lifecycle and emerging data science tooling trends. Nick reframes market hype around Spark versus actual value extraction.16:56–19:53 · Guest teaching 3/10 Q&A: Data Versioning and Aggregation Strategies Nick answers audience questions regarding file versioning scale and compliance walls in financial institutions. The host simply moderates the Q&A.19:53–21:59 · Guest teaching 4/10 Q&A: General Adoption vs. Specialized Workflows Nick addresses an audience question regarding non-data users adopting Domino, contrasting general tools like SharePoint with specialized quantitative research workflows.21:59–23:29 · Guest teaching 3/10 Q&A: Scientific Reproducibility and Event Conclusion Nick discusses reproducibility as a commercial driver driven by regulatory compliance before the host wraps up the session.0:16–5:08 · Guest disagreement 1/10 Presentation Objectives and Goals Nick presents a solo talk outlining data science myths versus best practices in advanced analytics organizations. The host does not speak during this monologue segment, so host scores are zero.5:08–8:10 · Guest disagreement 1/10 Core Themes: Collaboration, Reproducibility, and Agility Nick continues his presentation on collaboration, reproducibility, and agility, explaining why code versioning alone is insufficient for quantitative research. Host-side scores remain zero for this uninterrupted monologue.8:10–14:12 · Guest disagreement 0/10 Product Strategy: Incentivizing Best Practices Bottom-Up Nick demonstrates the Domino platform and explains how to incentivize best practices bottom-up. As another presentation monologue, host-side metrics are strictly zero.14:12–16:56 · Guest disagreement 2/10 Analytical Life Cycle and Data Science Tooling Trends Matt Turck steps in to praise the demo and ask informed questions about the analytical lifecycle and emerging data science tooling trends. Nick reframes market hype around Spark versus actual value extraction.16:56–19:53 · Guest disagreement 1/10 Q&A: Data Versioning and Aggregation Strategies Nick answers audience questions regarding file versioning scale and compliance walls in financial institutions. The host simply moderates the Q&A.19:53–21:59 · Guest disagreement 2/10 Q&A: General Adoption vs. Specialized Workflows Nick addresses an audience question regarding non-data users adopting Domino, contrasting general tools like SharePoint with specialized quantitative research workflows.21:59–23:29 · Guest disagreement 0/10 Q&A: Scientific Reproducibility and Event Conclusion Nick discusses reproducibility as a commercial driver driven by regulatory compliance before the host wraps up the session.0:16–5:08 · Matt pushing back 0/10 Presentation Objectives and Goals Nick presents a solo talk outlining data science myths versus best practices in advanced analytics organizations. The host does not speak during this monologue segment, so host scores are zero.5:08–8:10 · Matt pushing back 0/10 Core Themes: Collaboration, Reproducibility, and Agility Nick continues his presentation on collaboration, reproducibility, and agility, explaining why code versioning alone is insufficient for quantitative research. Host-side scores remain zero for this uninterrupted monologue.8:10–14:12 · Matt pushing back 0/10 Product Strategy: Incentivizing Best Practices Bottom-Up Nick demonstrates the Domino platform and explains how to incentivize best practices bottom-up. As another presentation monologue, host-side metrics are strictly zero.14:12–16:56 · Matt pushing back 1/10 Analytical Life Cycle and Data Science Tooling Trends Matt Turck steps in to praise the demo and ask informed questions about the analytical lifecycle and emerging data science tooling trends. Nick reframes market hype around Spark versus actual value extraction.16:56–19:53 · Matt pushing back 0/10 Q&A: Data Versioning and Aggregation Strategies Nick answers audience questions regarding file versioning scale and compliance walls in financial institutions. The host simply moderates the Q&A.19:53–21:59 · Matt pushing back 0/10 Q&A: General Adoption vs. Specialized Workflows Nick addresses an audience question regarding non-data users adopting Domino, contrasting general tools like SharePoint with specialized quantitative research workflows.21:59–23:29 · Matt pushing back 0/10 Q&A: Scientific Reproducibility and Event Conclusion Nick discusses reproducibility as a commercial driver driven by regulatory compliance before the host wraps up the session.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 8.7% · guest 91.3%12:00 · Matt 8.7% · guest 91.3%15:00 · Matt 10.8% · guest 89.2%15:00 · Matt 10.8% · guest 89.2%18:00 · Matt 0.3% · guest 99.7%18:00 · Matt 0.3% · guest 99.7%21:00 · Matt 11.3% · guest 88.7%21:00 · Matt 11.3% · guest 88.7%
Sharpest disagreement ▶ 15:56 Calling out Spark hype

Nick directly pushes back against industry hype, noting that many companies ask for Spark support without actually using or understanding their own Spark clusters.

Hardest push from Matt ▶ 15:38 Prompting on emerging tooling trends

Matt asks Nick to comment on specific industry shifts, asking what surprised him regarding data scientist preferences between tools like Zeppelin and Jupyter.

Biggest teaching moment ▶ 6:20 Defining quantitative reproducibility

Nick educates the audience on why standard software engineering code snapshots are insufficient for quantitative research without data, environment, and results.

Matt holds his own ▶ 15:38 Referencing the analytical lifecycle and competing tools

Matt demonstrates subject matter expertise by citing specific tool pairings like Zeppelin vs. Jupyter and probing the analytical lifecycle.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Presentation Objectives and Goals 0210 Nick presents a solo talk outlining data science myths versus best practices in advanced analytics organizations. The host does not speak during this monologue segment, so host scores are zero.
Core Themes: Collaboration, Reproducibility, and Agility 0310 Nick continues his presentation on collaboration, reproducibility, and agility, explaining why code versioning alone is insufficient for quantitative research. Host-side scores remain zero for this uninterrupted monologue.
Product Strategy: Incentivizing Best Practices Bottom-Up 0200 Nick demonstrates the Domino platform and explains how to incentivize best practices bottom-up. As another presentation monologue, host-side metrics are strictly zero.
Analytical Life Cycle and Data Science Tooling Trends 3321 Matt Turck steps in to praise the demo and ask informed questions about the analytical lifecycle and emerging data science tooling trends. Nick reframes market hype around Spark versus actual value extraction.
Q&A: Data Versioning and Aggregation Strategies 1310 Nick answers audience questions regarding file versioning scale and compliance walls in financial institutions. The host simply moderates the Q&A.
Q&A: General Adoption vs. Specialized Workflows 1420 Nick addresses an audience question regarding non-data users adopting Domino, contrasting general tools like SharePoint with specialized quantitative research workflows.
Q&A: Scientific Reproducibility and Event Conclusion 1300 Nick discusses reproducibility as a commercial driver driven by regulatory compliance before the host wraps up the session.

Statements from this episode (9)

Insight
Elprin: Elite data science organizations prioritize collective knowledge over solo practitioners
“The best organizations we've seen think of their work as contributing to collective knowledge.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 2:23
Insight
Elprin: Data science progress relies on compounding small insights, not epiphanies
“Progress actually comes from lots of little insights that compound over time”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 2:49
Insight
Elprin: Agility to experiment with new tools beats single-platform lock-in
“That's never the answer. It's much more about having agility to let people rapidly experiment with whatever the next thing is that comes out”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 3:45
Insight
Elprin: True data science reproducibility requires code, data, results, and environment
“What we've learned is that for this kind of work, this advanced analytics quantitative research work, reproducibility is more than just, hey, we have a snapshot of the code. You know, like, that's, you know, it's great for software engineering. GitHub does tha…”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 6:11
Assertion Not checkable as stated
Elprin: Regulatory compliance drives enterprise demand for data science reproducibility tools
“One of the really interesting things about reproducibility we've learned is that it's been extremely resonant in industries where there are regulatory and compliance concerns, and we're sort of seeing that more and more, especially in financial services, they'…”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 7:03
Insight
Elprin: Rapid model deployment creates feedback loops that sustain team funding
“If you've actually built something that works, how quickly can you get it out into the business? Because that, that creates the feedback loop that that creates credibility and buy-in to continue sort of investing in the work that's going on.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 7:52
Insight
Elprin: Drive organizational best practices by packaging them inside individual productivity tools
“What those people seem to want is the ability to test more ideas faster, that experimental agility some ways to sort of expose their work more out into the business. But let's package that in a way that automates or incentivizes best practices.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 8:51
Opinion
Elprin: Apache Spark still generates more industry hype than actual business value
“I think there's still more hype around Spark than actual value extraction from it.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 15:56
Assertion Not checkable as stated
Elprin: Highly sophisticated enterprise data teams often operate outside strict compliance walls
“I think that the folks who are doing sufficient, sufficiently sophisticated work to really want to use what we're doing tend to be in groups where they don't have those restrictions.”
Nick (Domino Data Lab) Nov 9, 2016 ▶ 19:23
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.