Dec 5, 2013 · 44m · mad

Panel: Continuuity, Sailthru and Visual Revenue // Data Driven NYC #7 // June 2012

Todd Papaioannou · 13m spoken Dennis Mortensen · 9m spoken Neil Capel · 4m spoken Matt Turck · 1m spoken David Crawford · 45s spoken Tony Bairwood · 43s spoken Joe Eisenberg · 42s spoken Nicholas Dandenberg · 42s spoken Gary Vidal · 38s spoken Dan Eldridge · 25s spoken Joel Natividad · 21s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC panel, tech leaders from Continuuity, Sailthru, and Visual Revenue discuss the evolution of big data platforms, real-time predictive analytics, and developer tools. The speakers highlight the transition toward intelligent software agents, developer-friendly infrastructure abstractions, and the operational balance between automated algorithms and human domain expertise.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 4.1% of the talking time here. How this is scored →

Matt as informed peer 0.7 Guest teaching 3.5 Guest disagreement 1.3 Matt pushing back 0.3
05100:0015:0030:000:00–8:40 · Matt as informed peer 3/10 Panelist Origin Stories and Recruitment Challenges Host Matt Turck opens the panel by framing startup origin questions and recruiting challenges. Guests share founding stories and recruitment realities, with Matt interjecting to press on developer hiring.8:40–13:59 · Matt as informed peer 1/10 The Future of Big Data Applications and Intelligent Agents Audience Q&A begins while host moderates. Dennis and Todd explain how intelligent agents will replace static dashboards across consumer intelligence applications.13:59–16:32 · Matt as informed peer 1/10 Simplifying Hadoop and Democratizing Big Data Development Todd responds to an audience question about simplifying Hadoop, drawing parallels to kernel development and the Spring framework.16:32–18:49 · Matt as informed peer 1/10 Real-Time Relevance, Algorithmic Anomalies, and Infrastructure Scaling Neil Capel addresses audience questions about real-time recommendations, infrastructure scaling, and filtering behavioral anomalies.18:49–22:14 · Matt as informed peer 1/10 Open Source Infrastructure vs Proprietary Tech Stack Strategy Todd explains why web-scale tech giants build and open-source infrastructure layers rather than selling proprietary stacks.22:14–24:38 · Matt as informed peer 0/10 Predictive Accuracy, Data Quality, and Modeling Realities Daniel clarifies that predictive success stems from high-quality data rather than complex modeling tricks, reframing the questioner's assumptions.24:38–30:38 · Matt as informed peer 0/10 Intelligent Agents, Closed Feedback Loops, and Editorial Control Todd and Dennis describe closed feedback loops in editorial systems where algorithms suggest content while respecting human editorial overrides.30:38–34:33 · Matt as informed peer 0/10 Real-Time Model Updates versus Batch Retraining Data scientists on the panel break down continuous model updating versus periodic batch retraining depending on concept drift.34:33–38:21 · Matt as informed peer 0/10 Data Privacy, Open Business Data, and Healthcare Potential Dennis rejects the premise of selling client data while Todd proposes opening anonymized healthcare data for societal benefit.38:21–41:43 · Matt as informed peer 1/10 Limitations of Legacy BI Tools and the Rise of Schema-at-Read Todd corrects an audience member's summary of his remarks before highlighting the industry shift from schema-at-write to schema-at-read.41:43–44:34 · Matt as informed peer 0/10 Enterprise Client Data Integration and High-Touch Sales Models Neil and Dennis detail enterprise client integration, balancing automated data collection with high-touch editorial discovery.44:34–44:59 · Matt as informed peer 0/10 Event Conclusion and Social Media Presenter Information Host Matt Turck wraps up the panel session and invites attendees to network and drink. Scores are 0 for monologue housekeeping.0:00–8:40 · Guest teaching 3/10 Panelist Origin Stories and Recruitment Challenges Host Matt Turck opens the panel by framing startup origin questions and recruiting challenges. Guests share founding stories and recruitment realities, with Matt interjecting to press on developer hiring.8:40–13:59 · Guest teaching 4/10 The Future of Big Data Applications and Intelligent Agents Audience Q&A begins while host moderates. Dennis and Todd explain how intelligent agents will replace static dashboards across consumer intelligence applications.13:59–16:32 · Guest teaching 4/10 Simplifying Hadoop and Democratizing Big Data Development Todd responds to an audience question about simplifying Hadoop, drawing parallels to kernel development and the Spring framework.16:32–18:49 · Guest teaching 3/10 Real-Time Relevance, Algorithmic Anomalies, and Infrastructure Scaling Neil Capel addresses audience questions about real-time recommendations, infrastructure scaling, and filtering behavioral anomalies.18:49–22:14 · Guest teaching 4/10 Open Source Infrastructure vs Proprietary Tech Stack Strategy Todd explains why web-scale tech giants build and open-source infrastructure layers rather than selling proprietary stacks.22:14–24:38 · Guest teaching 5/10 Predictive Accuracy, Data Quality, and Modeling Realities Daniel clarifies that predictive success stems from high-quality data rather than complex modeling tricks, reframing the questioner's assumptions.24:38–30:38 · Guest teaching 4/10 Intelligent Agents, Closed Feedback Loops, and Editorial Control Todd and Dennis describe closed feedback loops in editorial systems where algorithms suggest content while respecting human editorial overrides.30:38–34:33 · Guest teaching 4/10 Real-Time Model Updates versus Batch Retraining Data scientists on the panel break down continuous model updating versus periodic batch retraining depending on concept drift.34:33–38:21 · Guest teaching 4/10 Data Privacy, Open Business Data, and Healthcare Potential Dennis rejects the premise of selling client data while Todd proposes opening anonymized healthcare data for societal benefit.38:21–41:43 · Guest teaching 3/10 Limitations of Legacy BI Tools and the Rise of Schema-at-Read Todd corrects an audience member's summary of his remarks before highlighting the industry shift from schema-at-write to schema-at-read.41:43–44:34 · Guest teaching 4/10 Enterprise Client Data Integration and High-Touch Sales Models Neil and Dennis detail enterprise client integration, balancing automated data collection with high-touch editorial discovery.44:34–44:59 · Guest teaching 0/10 Event Conclusion and Social Media Presenter Information Host Matt Turck wraps up the panel session and invites attendees to network and drink. Scores are 0 for monologue housekeeping.0:00–8:40 · Guest disagreement 1/10 Panelist Origin Stories and Recruitment Challenges Host Matt Turck opens the panel by framing startup origin questions and recruiting challenges. Guests share founding stories and recruitment realities, with Matt interjecting to press on developer hiring.8:40–13:59 · Guest disagreement 1/10 The Future of Big Data Applications and Intelligent Agents Audience Q&A begins while host moderates. Dennis and Todd explain how intelligent agents will replace static dashboards across consumer intelligence applications.13:59–16:32 · Guest disagreement 1/10 Simplifying Hadoop and Democratizing Big Data Development Todd responds to an audience question about simplifying Hadoop, drawing parallels to kernel development and the Spring framework.16:32–18:49 · Guest disagreement 1/10 Real-Time Relevance, Algorithmic Anomalies, and Infrastructure Scaling Neil Capel addresses audience questions about real-time recommendations, infrastructure scaling, and filtering behavioral anomalies.18:49–22:14 · Guest disagreement 1/10 Open Source Infrastructure vs Proprietary Tech Stack Strategy Todd explains why web-scale tech giants build and open-source infrastructure layers rather than selling proprietary stacks.22:14–24:38 · Guest disagreement 2/10 Predictive Accuracy, Data Quality, and Modeling Realities Daniel clarifies that predictive success stems from high-quality data rather than complex modeling tricks, reframing the questioner's assumptions.24:38–30:38 · Guest disagreement 1/10 Intelligent Agents, Closed Feedback Loops, and Editorial Control Todd and Dennis describe closed feedback loops in editorial systems where algorithms suggest content while respecting human editorial overrides.30:38–34:33 · Guest disagreement 1/10 Real-Time Model Updates versus Batch Retraining Data scientists on the panel break down continuous model updating versus periodic batch retraining depending on concept drift.34:33–38:21 · Guest disagreement 2/10 Data Privacy, Open Business Data, and Healthcare Potential Dennis rejects the premise of selling client data while Todd proposes opening anonymized healthcare data for societal benefit.38:21–41:43 · Guest disagreement 3/10 Limitations of Legacy BI Tools and the Rise of Schema-at-Read Todd corrects an audience member's summary of his remarks before highlighting the industry shift from schema-at-write to schema-at-read.41:43–44:34 · Guest disagreement 1/10 Enterprise Client Data Integration and High-Touch Sales Models Neil and Dennis detail enterprise client integration, balancing automated data collection with high-touch editorial discovery.44:34–44:59 · Guest disagreement 0/10 Event Conclusion and Social Media Presenter Information Host Matt Turck wraps up the panel session and invites attendees to network and drink. Scores are 0 for monologue housekeeping.0:00–8:40 · Matt pushing back 2/10 Panelist Origin Stories and Recruitment Challenges Host Matt Turck opens the panel by framing startup origin questions and recruiting challenges. Guests share founding stories and recruitment realities, with Matt interjecting to press on developer hiring.8:40–13:59 · Matt pushing back 0/10 The Future of Big Data Applications and Intelligent Agents Audience Q&A begins while host moderates. Dennis and Todd explain how intelligent agents will replace static dashboards across consumer intelligence applications.13:59–16:32 · Matt pushing back 0/10 Simplifying Hadoop and Democratizing Big Data Development Todd responds to an audience question about simplifying Hadoop, drawing parallels to kernel development and the Spring framework.16:32–18:49 · Matt pushing back 0/10 Real-Time Relevance, Algorithmic Anomalies, and Infrastructure Scaling Neil Capel addresses audience questions about real-time recommendations, infrastructure scaling, and filtering behavioral anomalies.18:49–22:14 · Matt pushing back 0/10 Open Source Infrastructure vs Proprietary Tech Stack Strategy Todd explains why web-scale tech giants build and open-source infrastructure layers rather than selling proprietary stacks.22:14–24:38 · Matt pushing back 0/10 Predictive Accuracy, Data Quality, and Modeling Realities Daniel clarifies that predictive success stems from high-quality data rather than complex modeling tricks, reframing the questioner's assumptions.24:38–30:38 · Matt pushing back 0/10 Intelligent Agents, Closed Feedback Loops, and Editorial Control Todd and Dennis describe closed feedback loops in editorial systems where algorithms suggest content while respecting human editorial overrides.30:38–34:33 · Matt pushing back 0/10 Real-Time Model Updates versus Batch Retraining Data scientists on the panel break down continuous model updating versus periodic batch retraining depending on concept drift.34:33–38:21 · Matt pushing back 0/10 Data Privacy, Open Business Data, and Healthcare Potential Dennis rejects the premise of selling client data while Todd proposes opening anonymized healthcare data for societal benefit.38:21–41:43 · Matt pushing back 1/10 Limitations of Legacy BI Tools and the Rise of Schema-at-Read Todd corrects an audience member's summary of his remarks before highlighting the industry shift from schema-at-write to schema-at-read.41:43–44:34 · Matt pushing back 0/10 Enterprise Client Data Integration and High-Touch Sales Models Neil and Dennis detail enterprise client integration, balancing automated data collection with high-touch editorial discovery.44:34–44:59 · Matt pushing back 0/10 Event Conclusion and Social Media Presenter Information Host Matt Turck wraps up the panel session and invites attendees to network and drink. Scores are 0 for monologue housekeeping.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 42.2% · guest 57.8%0:00 · Matt 42.2% · guest 57.8%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 11.8% · guest 88.2%6:00 · Matt 11.8% · guest 88.2%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 0% · guest 100%21:00 · Matt 0% · guest 100%24:00 · Matt 0% · guest 100%24:00 · Matt 0% · guest 100%27:00 · Matt 0% · guest 100%27:00 · Matt 0% · guest 100%30:00 · Matt 0% · guest 100%30:00 · Matt 0% · guest 100%33:00 · Matt 0% · guest 100%33:00 · Matt 0% · guest 100%36:00 · Matt 2.2% · guest 97.8%36:00 · Matt 2.2% · guest 97.8%39:00 · Matt 0% · guest 100%39:00 · Matt 0% · guest 100%42:00 · Matt 3.5% · guest 96.5%42:00 · Matt 3.5% · guest 96.5%
Sharpest disagreement ▶ 38:49 Todd Corrects Audience Premise

Todd directly pushes back against an audience member's misquote, stating 'That's not exactly what I said, I don't think' to correct the record.

Hardest push from Matt ▶ 7:48 Host Probes Developer Recruitment

Host Matt Turck humorously challenges Todd's casual summary of raising capital and presses directly on the difficulty of recruiting infrastructure engineers.

Biggest teaching moment ▶ 23:04 Data Quality Beats Complex Modeling

Daniel educates the audience on predictive realities, explaining that simple modeling with great data consistently beats elite modeling with mediocre data.

Matt holds his own ▶ 0:01 Host Sets Panel Agenda

Matt Turck demonstrates solid industry grasp by structuring the discussion around practical startup origins and data recruitment obstacles.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Panelist Origin Stories and Recruitment Challenges 3312 Host Matt Turck opens the panel by framing startup origin questions and recruiting challenges. Guests share founding stories and recruitment realities, with Matt interjecting to press on developer hiring.
The Future of Big Data Applications and Intelligent Agents 1410 Audience Q&A begins while host moderates. Dennis and Todd explain how intelligent agents will replace static dashboards across consumer intelligence applications.
Simplifying Hadoop and Democratizing Big Data Development 1410 Todd responds to an audience question about simplifying Hadoop, drawing parallels to kernel development and the Spring framework.
Real-Time Relevance, Algorithmic Anomalies, and Infrastructure Scaling 1310 Neil Capel addresses audience questions about real-time recommendations, infrastructure scaling, and filtering behavioral anomalies.
Open Source Infrastructure vs Proprietary Tech Stack Strategy 1410 Todd explains why web-scale tech giants build and open-source infrastructure layers rather than selling proprietary stacks.
Predictive Accuracy, Data Quality, and Modeling Realities 0520 Daniel clarifies that predictive success stems from high-quality data rather than complex modeling tricks, reframing the questioner's assumptions.
Intelligent Agents, Closed Feedback Loops, and Editorial Control 0410 Todd and Dennis describe closed feedback loops in editorial systems where algorithms suggest content while respecting human editorial overrides.
Real-Time Model Updates versus Batch Retraining 0410 Data scientists on the panel break down continuous model updating versus periodic batch retraining depending on concept drift.
Data Privacy, Open Business Data, and Healthcare Potential 0420 Dennis rejects the premise of selling client data while Todd proposes opening anonymized healthcare data for societal benefit.
Limitations of Legacy BI Tools and the Rise of Schema-at-Read 1331 Todd corrects an audience member's summary of his remarks before highlighting the industry shift from schema-at-write to schema-at-read.
Enterprise Client Data Integration and High-Touch Sales Models 0410 Neil and Dennis detail enterprise client integration, balancing automated data collection with high-touch editorial discovery.
Event Conclusion and Social Media Presenter Information 0000 Host Matt Turck wraps up the panel session and invites attendees to network and drink. Scores are 0 for monologue housekeeping.

Statements from this episode (12)

Assertion Not checkable as stated
Half of newspaper article views in 2012 originated from homepages
“About half of all the article views on any given day for any one of these properties were driven off the homepage or the section front page.”
Dennis Mortensen Dec 5, 2013 ▶ 1:49
Insight
Launching a startup as a VC EIR beats self-funding
“The best way to launch a company is not to do it on your kitchen table or in your front room with seven guys that you fly over from Europe. It's to have a VC pay you to go and launch a company.”
Todd Papaioannou Dec 5, 2013 ▶ 6:59
Insight
Founders should make their operational mistakes using investor capital
“One of the lessons I learned is that you know, it's always better to make mistakes and learn things on other people's money, right?”
Todd Papaioannou Dec 5, 2013 ▶ 7:12
Assertion Not checkable as stated
Yahoo had 200 petabytes of data when Todd Papaioannou left
“At Yahoo, when I left, we had 200 petabytes of data that we used to, you know, build models and do personalization and scoring and ranking and all of that stuff.”
Todd Papaioannou Dec 5, 2013 ▶ 10:56
Insight
Giving Google Analytics logins to non-analysts is naive
“Editors should be editors, not analysts. I think that would work in pretty much any industry, because the idea of handing out random logins to Google Analytics and believe that your organization is wiser seems Naive tomorrow, perhaps slightly smart today, but …”
Dennis Mortensen Dec 5, 2013 ▶ 11:22
Assertion Not checkable as stated
Papaioannou in 2012: HDFS will scale forever for most companies
“HDFS, which will scale forever for probably like 99.9% of the companies on the planet. They're just never going to run out of space, you know, to put data into it.”
Todd Papaioannou Dec 5, 2013 ▶ 15:08
Assertion Supported
Facebook built Hive to provide a SQL interface on Hadoop
“Facebook built Hive, right, because they needed a tool to sit on top of Hadoop, you know, to allow their business analysts to kind of sequel interface to this big data platform.”
Todd Papaioannou Dec 5, 2013 ▶ 20:08
Disclosure
Yahoo open-sourced infrastructure because its competitive advantage was in applications
“And so this is actually at Yahoo, a very distinctive strategy that we adopted in the global cloud computing group, which to say, we will open source all of our infrastructure because it's not a core competitive advantage to us. The applications and the secret …”
Todd Papaioannou Dec 5, 2013 ▶ 21:03
Assertion Not checkable as stated
Yahoo drove a 300% CTR lift by combining algorithms with human editors
“Once we moved it all to algorithm and had editors just use their, you know, what humans are good at, which is this kind of like, you know, just being able to do correlations that the algorithms maybe didn't, we were able to drive like a 300% improvement in cli…”
Todd Papaioannou Dec 5, 2013 ▶ 27:45
Opinion
Anonymized healthcare data should be made openly available for research
“I would love it in this kind of, like, altruistic world if all Healthcare data, anonymized healthcare data, was made available for researchers and people to actually build applications on top of, because just imagine the efficiencies and benefits you could get…”
Todd Papaioannou Dec 5, 2013 ▶ 36:21
Insight
Data cubes only answer the questions users already know to ask
“Cubes are basically like answers to questions you already knew.”
Todd Papaioannou Dec 5, 2013 ▶ 40:57
Prediction Not checkable as stated
In 2012, Papaioannou predicted easy data tools for non-programmers by 2019
“And I think it's going to take us, like I said, five to seven years until we get to the point where you don't have to be a programmer who understands Python or R or, you know, pick your favorite, you know, language. So actually go and get that insight.”
Todd Papaioannou Dec 5, 2013 ▶ 41:16
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.