May 27, 2014 · 22m · mad

Stephen Purpura, Context Relevant // Data Driven #26 // April 2014 (Hosted by FirstMark Capital)

Stephen Purpura · 17m spoken Nick Hartman · 43s spoken Matt Turck · 9s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At a Data Driven NYC event, Context Relevant CEO Stephen Purpura demonstrates how Big Data 2.0 and automated machine learning can eliminate data science bottlenecks by rapidly discovering mathematical formulas and predictive models at scale. Through a live demonstration solving Hero of Alexandria's triangle area problem across 120 distributed servers, Purpura illustrates how automated feature engineering and cloud infrastructure transform enterprise analytics.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 0.7% of the talking time here. How this is scored →

Matt as informed peer 0.2 Guest teaching 3.4 Guest disagreement 0.2 Matt pushing back 0.2
05100:0010:0020:001:36–4:47 · Matt as informed peer 0/10 Framing the Triangle Area Problem in Big Data Terms Purpura presents a monologue introducing Big Data 2.0 and automated feature engineering using Hero of Alexandria's triangle area formula. The host does not speak during this presentation segment.4:47–10:39 · Matt as informed peer 0/10 Overview of the 3-Step Automated Process and Code Setup Purpura walks through the Python script setup, explaining how 120 distributed servers process 100 million synthetic triangles to discover mathematical relationships automatically.10:39–14:03 · Matt as informed peer 0/10 Mathematical Algebra and Feature Transformation Search Purpura explains mathematical feature transformations, contrasting basic polynomial features against semi-perimeter derivations to eliminate error. The presentation remains a non-interactive monologue.14:03–16:52 · Matt as informed peer 0/10 Live Demo Execution Results and the Value of Big Data 2.0 Purpura presents the live model execution output and summarizes commercial applications in financial services and trading spreads. The host does not intervene.16:52–22:36 · Matt as informed peer 1/10 Audience Q&A on Correlation, Causality, and Data Science Productivity Host Matt Turck briefly limits audience Q&A to one question. Audience member Nick Hartman asks about correlation vs causality, prompting Purpura to explain train/tune/test validation methodologies.1:36–4:47 · Guest teaching 3/10 Framing the Triangle Area Problem in Big Data Terms Purpura presents a monologue introducing Big Data 2.0 and automated feature engineering using Hero of Alexandria's triangle area formula. The host does not speak during this presentation segment.4:47–10:39 · Guest teaching 3/10 Overview of the 3-Step Automated Process and Code Setup Purpura walks through the Python script setup, explaining how 120 distributed servers process 100 million synthetic triangles to discover mathematical relationships automatically.10:39–14:03 · Guest teaching 4/10 Mathematical Algebra and Feature Transformation Search Purpura explains mathematical feature transformations, contrasting basic polynomial features against semi-perimeter derivations to eliminate error. The presentation remains a non-interactive monologue.14:03–16:52 · Guest teaching 3/10 Live Demo Execution Results and the Value of Big Data 2.0 Purpura presents the live model execution output and summarizes commercial applications in financial services and trading spreads. The host does not intervene.16:52–22:36 · Guest teaching 4/10 Audience Q&A on Correlation, Causality, and Data Science Productivity Host Matt Turck briefly limits audience Q&A to one question. Audience member Nick Hartman asks about correlation vs causality, prompting Purpura to explain train/tune/test validation methodologies.1:36–4:47 · Guest disagreement 0/10 Framing the Triangle Area Problem in Big Data Terms Purpura presents a monologue introducing Big Data 2.0 and automated feature engineering using Hero of Alexandria's triangle area formula. The host does not speak during this presentation segment.4:47–10:39 · Guest disagreement 0/10 Overview of the 3-Step Automated Process and Code Setup Purpura walks through the Python script setup, explaining how 120 distributed servers process 100 million synthetic triangles to discover mathematical relationships automatically.10:39–14:03 · Guest disagreement 0/10 Mathematical Algebra and Feature Transformation Search Purpura explains mathematical feature transformations, contrasting basic polynomial features against semi-perimeter derivations to eliminate error. The presentation remains a non-interactive monologue.14:03–16:52 · Guest disagreement 0/10 Live Demo Execution Results and the Value of Big Data 2.0 Purpura presents the live model execution output and summarizes commercial applications in financial services and trading spreads. The host does not intervene.16:52–22:36 · Guest disagreement 1/10 Audience Q&A on Correlation, Causality, and Data Science Productivity Host Matt Turck briefly limits audience Q&A to one question. Audience member Nick Hartman asks about correlation vs causality, prompting Purpura to explain train/tune/test validation methodologies.1:36–4:47 · Matt pushing back 0/10 Framing the Triangle Area Problem in Big Data Terms Purpura presents a monologue introducing Big Data 2.0 and automated feature engineering using Hero of Alexandria's triangle area formula. The host does not speak during this presentation segment.4:47–10:39 · Matt pushing back 0/10 Overview of the 3-Step Automated Process and Code Setup Purpura walks through the Python script setup, explaining how 120 distributed servers process 100 million synthetic triangles to discover mathematical relationships automatically.10:39–14:03 · Matt pushing back 0/10 Mathematical Algebra and Feature Transformation Search Purpura explains mathematical feature transformations, contrasting basic polynomial features against semi-perimeter derivations to eliminate error. The presentation remains a non-interactive monologue.14:03–16:52 · Matt pushing back 0/10 Live Demo Execution Results and the Value of Big Data 2.0 Purpura presents the live model execution output and summarizes commercial applications in financial services and trading spreads. The host does not intervene.16:52–22:36 · Matt pushing back 1/10 Audience Q&A on Correlation, Causality, and Data Science Productivity Host Matt Turck briefly limits audience Q&A to one question. Audience member Nick Hartman asks about correlation vs causality, prompting Purpura to explain train/tune/test validation methodologies.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 4.8% · guest 95.2%15:00 · Matt 4.8% · guest 95.2%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 1% · guest 99%21:00 · Matt 1% · guest 99%
Sharpest disagreement ▶ 17:10 Challenging automated model validity

Audience member Nick Hartman politely challenges the premise of automated curve-fitting by pointing out that statistical correlation does not equal causality.

Hardest push from Matt ▶ 16:58 Host limits Q&A duration

Host Matt Turck steps in to restrict the Q&A session to a single question due to time limits.

Biggest teaching moment ▶ 18:30 Explaining model validation mechanics

Purpura educates the audience on using split datasets (train, tune, test) and detecting sub-linear learning curves to prevent spurious correlations.

Matt holds his own ▶ 16:58 Host maintains event control

Matt Turck briefly asserts control over the event schedule by calling time and selecting the final questioner.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Framing the Triangle Area Problem in Big Data Terms 0300 Purpura presents a monologue introducing Big Data 2.0 and automated feature engineering using Hero of Alexandria's triangle area formula. The host does not speak during this presentation segment.
Overview of the 3-Step Automated Process and Code Setup 0300 Purpura walks through the Python script setup, explaining how 120 distributed servers process 100 million synthetic triangles to discover mathematical relationships automatically.
Mathematical Algebra and Feature Transformation Search 0400 Purpura explains mathematical feature transformations, contrasting basic polynomial features against semi-perimeter derivations to eliminate error. The presentation remains a non-interactive monologue.
Live Demo Execution Results and the Value of Big Data 2.0 0300 Purpura presents the live model execution output and summarizes commercial applications in financial services and trading spreads. The host does not intervene.
Audience Q&A on Correlation, Causality, and Data Science Productivity 1411 Host Matt Turck briefly limits audience Q&A to one question. Audience member Nick Hartman asks about correlation vs causality, prompting Purpura to explain train/tune/test validation methodologies.

Statements from this episode (6)

Assertion Not checkable as stated
Stephen Purpura claims Context Relevant derives mathematical formulas from data in seconds
“And now our software does this automatically from data in a few seconds.”
Stephen Purpura May 27, 2014 ▶ 2:13
Disclosure
Context Relevant ran a 120-server demo to automatically discover triangle area formulas
“I just started it, and what this is going to do is start up a 120 servers to analyze the, about a hundred million triangles to figure out what the generalizing equation is for calculating the area.”
Stephen Purpura May 27, 2014 ▶ 6:02
Assertion Not checkable as stated
Context Relevant analyzes millions of ML models while competitors analyze just one
“And then another part of our company is the data science group, which has optimized, ah, a pipeline to run on top of that infrastructure, so, such that you can do analysis of millions of models in the time that it takes to, most companies to frankly do one.”
Stephen Purpura May 27, 2014 ▶ 9:48
Insight
Stephen Purpura argues 90-percent accurate predictive models suffice for commercial monetization
“You know, you can make money with a 90% answer. You know, especially on large, on thing, on problems where there's a large data volume.”
Stephen Purpura May 27, 2014 ▶ 12:35
Disclosure
Context Relevant positions its product to boost productivity, not replace data scientists
“And so, we actually do not sell our technology as a complete replacement for data scientists. We encourage you to think of it as a huge productivity game.”
Stephen Purpura May 27, 2014 ▶ 21:12
Opinion
Stephen Purpura says automated ML is a generational leap lacking full autonomy
“The technology is not at the phase yet where it's completely automated and can be used by anyone, but it's certainly made a major generational leap over where, when I was training graduate students at Cornell only a few years ago.”
Stephen Purpura May 27, 2014 ▶ 22:02
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.