Insight certainty 4/5 debate potential 2/5

Frame-level tags lack the high-level semantic context required for video discovery

David Luan · David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC) · May 28, 2015 · at 11:05

David Luan, co-founder of computer vision startup Dextro, discusses why simple object detection tags fall short for media platforms.

0:00 / 0:15exact quote · 15.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Video or frame-level tags, the sort that you might see on, that I've put on the screen right now, didn't actually solve their problem. Because that was, it was too low level in like a, in a, in not in terms of a granularity sense, but in terms of how much additional information it added about what was happening in the video as a whole.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from David Luan

Assertion Supported
Dextro was the first to offer automated video analysis as a service
“We were the first company to figure out how to get this level of analysis of what's happening in videos as a service.”
David Luan May 28, 2015 ▶ 3:01 David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Opinion
Academic reviewers are tired of papers blindly applying deep learning
“And I think the reviewers now are pretty tired of that and they're moving on, but.”
David Luan May 28, 2015 ▶ 7:34 David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Insight
Enterprise customers will not pay for 80% accurate machine learning
“But in most cases, customers aren't willing to pay for a product that only gets them 80% of the way. You have to kind of like specialize and focus on the problem to make sure that you get to being 100%.”
David Luan May 28, 2015 ▶ 8:44 David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Disclosure
Dextro analyzes video strictly through computer vision, not metadata
“This is all just done with computer vision. We don't use any of the metadata whatsoever to identify what's actually happening.”
David Luan May 28, 2015 ▶ 3:55 David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Insight
Models trained on stock images fail to generalize to real-world video
“Classifiers that are and models that are trained in tune on this particular, on iconic type of data don't really generalize as well when you apply it to something that you might see in a real-world video like on YouTube or on Periscope.”
David Luan May 28, 2015 ▶ 10:38 David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Insight
Frame-by-frame video analysis discards critical temporal motion data
“The naive approach to generalizing the video is to analyze videos as just a frame by frame sequence of photos. But there's so much encoded in the motion information in a video that to just do that actually just throws all of it out.”
David Luan May 28, 2015 ▶ 13:02 David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.