May 28, 2015 · 24m · mad

David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)

David Luan · 18m spoken Matt Turck · 1m spoken Andrew / Eric Smith · 19s spoken Richard Blue · 9s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Data Driven NYC, David Luan presents Dextro's computer vision and deep learning platform, explaining how automated feature extraction, timeline analysis, live stream processing, and custom taxonomy mapping convert unstructured video streams into actionable metadata for media enterprises.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 5.6% of the talking time here. How this is scored →

Matt as informed peer 0.7 Guest teaching 3.0 Guest disagreement 0.2 Matt pushing back 0.3
05100:0010:0020:000:13–2:21 · Matt as informed peer 0/10 Dextro's Core Mission and Scale of Video Data Presentation monologue where David Luan introduces Dextro and outlines the vast scale of video uploaded to platforms like YouTube. The host does not speak in this segment.2:21–4:58 · Matt as informed peer 0/10 Developer-Friendly JSON API Infrastructure David continues his presentation on Dextro's JSON API output and real-time live streaming analysis on Periscope. Host is silent during the presentation.4:58–7:18 · Matt as informed peer 0/10 High-Value Applications: Discovery, Curation, and Audience Insights David outlines historical milestones in deep learning, such as Krzyzewski's ImageNet neural network and Karpathy's image captioning. Monologue format with no host interaction.7:18–9:45 · Matt as informed peer 0/10 The Allure of Doing Everything vs. Vertical Focus David explains the strategic trade-off between building general horizontal ML tools and solving vertical customer problems. Uninterrupted presentation.9:45–12:45 · Matt as informed peer 0/10 Real-World Content vs. Iconic Stock Images David contrasts ideal iconic stock photos with noisy real-world video content and explains how Dextro handles partner taxonomies. Monologue format.12:45–24:55 · Matt as informed peer 4/10 Incorporating Motion Cues with Video-Specific Models Matt Turck opens Q&A with conceptual questions regarding predictive video analytics and horizontal AI layers, prompting David and audience members to discuss practical video AI applications.0:13–2:21 · Guest teaching 2/10 Dextro's Core Mission and Scale of Video Data Presentation monologue where David Luan introduces Dextro and outlines the vast scale of video uploaded to platforms like YouTube. The host does not speak in this segment.2:21–4:58 · Guest teaching 2/10 Developer-Friendly JSON API Infrastructure David continues his presentation on Dextro's JSON API output and real-time live streaming analysis on Periscope. Host is silent during the presentation.4:58–7:18 · Guest teaching 3/10 High-Value Applications: Discovery, Curation, and Audience Insights David outlines historical milestones in deep learning, such as Krzyzewski's ImageNet neural network and Karpathy's image captioning. Monologue format with no host interaction.7:18–9:45 · Guest teaching 4/10 The Allure of Doing Everything vs. Vertical Focus David explains the strategic trade-off between building general horizontal ML tools and solving vertical customer problems. Uninterrupted presentation.9:45–12:45 · Guest teaching 3/10 Real-World Content vs. Iconic Stock Images David contrasts ideal iconic stock photos with noisy real-world video content and explains how Dextro handles partner taxonomies. Monologue format.12:45–24:55 · Guest teaching 4/10 Incorporating Motion Cues with Video-Specific Models Matt Turck opens Q&A with conceptual questions regarding predictive video analytics and horizontal AI layers, prompting David and audience members to discuss practical video AI applications.0:13–2:21 · Guest disagreement 0/10 Dextro's Core Mission and Scale of Video Data Presentation monologue where David Luan introduces Dextro and outlines the vast scale of video uploaded to platforms like YouTube. The host does not speak in this segment.2:21–4:58 · Guest disagreement 0/10 Developer-Friendly JSON API Infrastructure David continues his presentation on Dextro's JSON API output and real-time live streaming analysis on Periscope. Host is silent during the presentation.4:58–7:18 · Guest disagreement 0/10 High-Value Applications: Discovery, Curation, and Audience Insights David outlines historical milestones in deep learning, such as Krzyzewski's ImageNet neural network and Karpathy's image captioning. Monologue format with no host interaction.7:18–9:45 · Guest disagreement 0/10 The Allure of Doing Everything vs. Vertical Focus David explains the strategic trade-off between building general horizontal ML tools and solving vertical customer problems. Uninterrupted presentation.9:45–12:45 · Guest disagreement 0/10 Real-World Content vs. Iconic Stock Images David contrasts ideal iconic stock photos with noisy real-world video content and explains how Dextro handles partner taxonomies. Monologue format.12:45–24:55 · Guest disagreement 1/10 Incorporating Motion Cues with Video-Specific Models Matt Turck opens Q&A with conceptual questions regarding predictive video analytics and horizontal AI layers, prompting David and audience members to discuss practical video AI applications.0:13–2:21 · Matt pushing back 0/10 Dextro's Core Mission and Scale of Video Data Presentation monologue where David Luan introduces Dextro and outlines the vast scale of video uploaded to platforms like YouTube. The host does not speak in this segment.2:21–4:58 · Matt pushing back 0/10 Developer-Friendly JSON API Infrastructure David continues his presentation on Dextro's JSON API output and real-time live streaming analysis on Periscope. Host is silent during the presentation.4:58–7:18 · Matt pushing back 0/10 High-Value Applications: Discovery, Curation, and Audience Insights David outlines historical milestones in deep learning, such as Krzyzewski's ImageNet neural network and Karpathy's image captioning. Monologue format with no host interaction.7:18–9:45 · Matt pushing back 0/10 The Allure of Doing Everything vs. Vertical Focus David explains the strategic trade-off between building general horizontal ML tools and solving vertical customer problems. Uninterrupted presentation.9:45–12:45 · Matt pushing back 0/10 Real-World Content vs. Iconic Stock Images David contrasts ideal iconic stock photos with noisy real-world video content and explains how Dextro handles partner taxonomies. Monologue format.12:45–24:55 · Matt pushing back 2/10 Incorporating Motion Cues with Video-Specific Models Matt Turck opens Q&A with conceptual questions regarding predictive video analytics and horizontal AI layers, prompting David and audience members to discuss practical video AI applications.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 14.3% · guest 85.7%12:00 · Matt 14.3% · guest 85.7%15:00 · Matt 20.7% · guest 79.3%15:00 · Matt 20.7% · guest 79.3%18:00 · Matt 1.1% · guest 98.9%18:00 · Matt 1.1% · guest 98.9%21:00 · Matt 9.2% · guest 90.8%21:00 · Matt 9.2% · guest 90.8%24:00 · Matt 3.9% · guest 96.1%24:00 · Matt 3.9% · guest 96.1%
Sharpest disagreement ▶ 18:12 Deflecting Unreleased Feature Details

When asked if audio cues like truck crash sounds are integrated into the model, David humorously deflects the question with 'Let's talk in a couple months', setting a firm boundary on unannounced technical features.

Hardest push from Matt ▶ 15:18 Challenging the Vertical AI Premise

Matt Turck pushes past the guest's thesis on vertical focus by asking whether a universal horizontal AI layer is coming or remains pure science fiction.

Biggest teaching moment ▶ 16:10 Distinguishing Algorithms from Last-Mile Solutions

David educates the host and audience on why raw deep learning algorithms fail commercially without the additional last-mile labor of data collection and fine-tuning for specific customer workflows.

Matt holds his own ▶ 14:31 Proposing Real-Time Predictive Video Analytics

Matt Turck demonstrates keen technical foresight by asking whether real-time computer vision can evolve into spatial-temporal predictive analytics for tracking live movement.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Dextro's Core Mission and Scale of Video Data 0200 Presentation monologue where David Luan introduces Dextro and outlines the vast scale of video uploaded to platforms like YouTube. The host does not speak in this segment.
Developer-Friendly JSON API Infrastructure 0200 David continues his presentation on Dextro's JSON API output and real-time live streaming analysis on Periscope. Host is silent during the presentation.
High-Value Applications: Discovery, Curation, and Audience Insights 0300 David outlines historical milestones in deep learning, such as Krzyzewski's ImageNet neural network and Karpathy's image captioning. Monologue format with no host interaction.
The Allure of Doing Everything vs. Vertical Focus 0400 David explains the strategic trade-off between building general horizontal ML tools and solving vertical customer problems. Uninterrupted presentation.
Real-World Content vs. Iconic Stock Images 0300 David contrasts ideal iconic stock photos with noisy real-world video content and explains how Dextro handles partner taxonomies. Monologue format.
Incorporating Motion Cues with Video-Specific Models 4412 Matt Turck opens Q&A with conceptual questions regarding predictive video analytics and horizontal AI layers, prompting David and audience members to discuss practical video AI applications.

Statements from this episode (14)

Assertion Supported
In 2015, 300 hours of video were uploaded to YouTube every minute
“Around 300 hours of video, it's probably even more now actually, are uploaded to YouTube every minute right now.”
David Luan May 28, 2015 ▶ 0:40
Disclosure
Dextro uses a salience graph to measure video concept prominence
“So what we do is we provide what's, what's also a salience graph, which is a discounted score of how important every concept Or how prominent a particular category is over the video as a whole. So it's not a measure of our confidence, but of actually how impor…”
David Luan May 28, 2015 ▶ 1:44
Assertion Supported
Dextro was the first to offer automated video analysis as a service
“We were the first company to figure out how to get this level of analysis of what's happening in videos as a service.”
David Luan May 28, 2015 ▶ 3:01
Disclosure
Dextro analyzes video strictly through computer vision, not metadata
“This is all just done with computer vision. We don't use any of the metadata whatsoever to identify what's actually happening.”
David Luan May 28, 2015 ▶ 3:55
Assertion Partly supported
AlexNet halved state-of-the-art visual recognition error rates in a single year
“Krzyzewski's work on ILS VRC, which is the ImageNet Large Scale Visual Recognition Challenge, where in one year, essentially they blew away the previous state of the art with a deep convolutional neural network by about, like, half of the final error rate on t…”
David Luan May 28, 2015 ▶ 6:23
Opinion
Academic reviewers are tired of papers blindly applying deep learning
“And I think the reviewers now are pretty tired of that and they're moving on, but.”
David Luan May 28, 2015 ▶ 7:34
Insight
Enterprise customers will not pay for 80% accurate machine learning
“But in most cases, customers aren't willing to pay for a product that only gets them 80% of the way. You have to kind of like specialize and focus on the problem to make sure that you get to being 100%.”
David Luan May 28, 2015 ▶ 8:44
Insight
Models trained on stock images fail to generalize to real-world video
“Classifiers that are and models that are trained in tune on this particular, on iconic type of data don't really generalize as well when you apply it to something that you might see in a real-world video like on YouTube or on Periscope.”
David Luan May 28, 2015 ▶ 10:38
Insight
Frame-level tags lack the high-level semantic context required for video discovery
“Video or frame-level tags, the sort that you might see on, that I've put on the screen right now, didn't actually solve their problem. Because that was, it was too low level in like a, in a, in not in terms of a granularity sense, but in terms of how much addi…”
David Luan May 28, 2015 ▶ 11:05
Disclosure
Dextro built architecture to map custom customer taxonomies without retraining models
“Our core machine learning system needed to be able to easily adapt and generalize to customer and partner taxonomies without restarting training and data collection and everything like that from scratch every single time. And so how we solved that problem was …”
David Luan May 28, 2015 ▶ 12:21
Insight
Frame-by-frame video analysis discards critical temporal motion data
“The naive approach to generalizing the video is to analyze videos as just a frame by frame sequence of photos. But there's so much encoded in the motion information in a video that to just do that actually just throws all of it out.”
David Luan May 28, 2015 ▶ 13:02
Prediction Not checkable as stated
Enterprise deep learning will remain vertical before converging horizontally
“In terms of where we're going to see it In industry, it's likely to be still in a kind of vertical by vertical system for a little while before we start seeing a little more convergence.”
David Luan May 28, 2015 ▶ 16:43
Insight
Public user-generated video tags are too noisy for training vision models
“So with regard to using tags that are already present on the internet, we run into a bunch of different problems, which is that, one, so there is some information to be gained there in general, but it's extremely noisy, and the, another big thing that we see i…”
David Luan May 28, 2015 ▶ 19:20
Disclosure
Dextro created a real-time aggregator for all live public Periscope streams
“Tomorrow we're just ironing out that one issue, but you guys should all check out stream.dextro.co, which aggregates every live public Periscope stream, and you can kind of browse based on what you think is most interesting at any given moment to discover, lik…”
David Luan May 28, 2015 ▶ 24:29
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.