Nov 9, 2016 · 23m · mad
Lessons Learned from Advanced Data Science Orgs // Domino Data Lab [FirstMark's Data Driven]
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this presentation at FirstMark's Data Driven NYC, Domino Data Lab CEO Nick Elprin outlines best practices for scaling enterprise data science organizations through collaboration, agility, and reproducibility. He combines theoretical organizational insights, product strategy, a live platform demonstration, and an interactive Q&A session on modern quantitative research tooling.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.6% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Nick directly pushes back against industry hype, noting that many companies ask for Spark support without actually using or understanding their own Spark clusters.
Hardest push from Matt ▶ 15:38 Prompting on emerging tooling trendsMatt asks Nick to comment on specific industry shifts, asking what surprised him regarding data scientist preferences between tools like Zeppelin and Jupyter.
Biggest teaching moment ▶ 6:20 Defining quantitative reproducibilityNick educates the audience on why standard software engineering code snapshots are insufficient for quantitative research without data, environment, and results.
Matt holds his own ▶ 15:38 Referencing the analytical lifecycle and competing toolsMatt demonstrates subject matter expertise by citing specific tool pairings like Zeppelin vs. Jupyter and probing the analytical lifecycle.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Presentation Objectives and Goals | 0 | 2 | 1 | 0 | Nick presents a solo talk outlining data science myths versus best practices in advanced analytics organizations. The host does not speak during this monologue segment, so host scores are zero. | |
| Core Themes: Collaboration, Reproducibility, and Agility | 0 | 3 | 1 | 0 | Nick continues his presentation on collaboration, reproducibility, and agility, explaining why code versioning alone is insufficient for quantitative research. Host-side scores remain zero for this uninterrupted monologue. | |
| Product Strategy: Incentivizing Best Practices Bottom-Up | 0 | 2 | 0 | 0 | Nick demonstrates the Domino platform and explains how to incentivize best practices bottom-up. As another presentation monologue, host-side metrics are strictly zero. | |
| Analytical Life Cycle and Data Science Tooling Trends | 3 | 3 | 2 | 1 | Matt Turck steps in to praise the demo and ask informed questions about the analytical lifecycle and emerging data science tooling trends. Nick reframes market hype around Spark versus actual value extraction. | |
| Q&A: Data Versioning and Aggregation Strategies | 1 | 3 | 1 | 0 | Nick answers audience questions regarding file versioning scale and compliance walls in financial institutions. The host simply moderates the Q&A. | |
| Q&A: General Adoption vs. Specialized Workflows | 1 | 4 | 2 | 0 | Nick addresses an audience question regarding non-data users adopting Domino, contrasting general tools like SharePoint with specialized quantitative research workflows. | |
| Q&A: Scientific Reproducibility and Event Conclusion | 1 | 3 | 0 | 0 | Nick discusses reproducibility as a commercial driver driven by regulatory compliance before the host wraps up the session. |