May 27, 2014 · 22m · mad
Stephen Purpura, Context Relevant // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
At a Data Driven NYC event, Context Relevant CEO Stephen Purpura demonstrates how Big Data 2.0 and automated machine learning can eliminate data science bottlenecks by rapidly discovering mathematical formulas and predictive models at scale. Through a live demonstration solving Hero of Alexandria's triangle area problem across 120 distributed servers, Purpura illustrates how automated feature engineering and cloud infrastructure transform enterprise analytics.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 0.7% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Audience member Nick Hartman politely challenges the premise of automated curve-fitting by pointing out that statistical correlation does not equal causality.
Hardest push from Matt ▶ 16:58 Host limits Q&A durationHost Matt Turck steps in to restrict the Q&A session to a single question due to time limits.
Biggest teaching moment ▶ 18:30 Explaining model validation mechanicsPurpura educates the audience on using split datasets (train, tune, test) and detecting sub-linear learning curves to prevent spurious correlations.
Matt holds his own ▶ 16:58 Host maintains event controlMatt Turck briefly asserts control over the event schedule by calling time and selecting the final questioner.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Framing the Triangle Area Problem in Big Data Terms | 0 | 3 | 0 | 0 | Purpura presents a monologue introducing Big Data 2.0 and automated feature engineering using Hero of Alexandria's triangle area formula. The host does not speak during this presentation segment. | |
| Overview of the 3-Step Automated Process and Code Setup | 0 | 3 | 0 | 0 | Purpura walks through the Python script setup, explaining how 120 distributed servers process 100 million synthetic triangles to discover mathematical relationships automatically. | |
| Mathematical Algebra and Feature Transformation Search | 0 | 4 | 0 | 0 | Purpura explains mathematical feature transformations, contrasting basic polynomial features against semi-perimeter derivations to eliminate error. The presentation remains a non-interactive monologue. | |
| Live Demo Execution Results and the Value of Big Data 2.0 | 0 | 3 | 0 | 0 | Purpura presents the live model execution output and summarizes commercial applications in financial services and trading spreads. The host does not intervene. | |
| Audience Q&A on Correlation, Causality, and Data Science Productivity | 1 | 4 | 1 | 1 | Host Matt Turck briefly limits audience Q&A to one question. Audience member Nick Hartman asks about correlation vs causality, prompting Purpura to explain train/tune/test validation methodologies. |