David Luan, co-founder of Dextro, answers an audience question about leveraging existing web video tags like YouTube tags to train computer vision models.
“So with regard to using tags that are already present on the internet, we run into a bunch of different problems, which is that, one, so there is some information to be gained there in general, but it's extremely noisy, and the, another big thing that we see is, is a difference between editorial tags of, like, what's, of, like, how, how would an editor looking at this piece of content Figure out what's most salient and what's actually visually prominent. And so as a result, they're correlated, but they're not so well to the point where we could just directly use that to directly use that to train.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from David Luan
AssertionSupported
Dextro was the first to offer automated video analysis as a service
“We were the first company to figure out how to get this level of analysis of what's happening in videos as a service.”
David LuanMay 28, 2015▶ 3:01David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Opinion
Academic reviewers are tired of papers blindly applying deep learning
“And I think the reviewers now are pretty tired of that and they're moving on, but.”
David LuanMay 28, 2015▶ 7:34David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Insight
Enterprise customers will not pay for 80% accurate machine learning
“But in most cases, customers aren't willing to pay for a product that only gets them 80% of the way. You have to kind of like specialize and focus on the problem to make sure that you get to being 100%.”
David LuanMay 28, 2015▶ 8:44David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Disclosure
Dextro analyzes video strictly through computer vision, not metadata
“This is all just done with computer vision. We don't use any of the metadata whatsoever to identify what's actually happening.”
David LuanMay 28, 2015▶ 3:55David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Insight
Models trained on stock images fail to generalize to real-world video
“Classifiers that are and models that are trained in tune on this particular, on iconic type of data don't really generalize as well when you apply it to something that you might see in a real-world video like on YouTube or on Periscope.”
David LuanMay 28, 2015▶ 10:38David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Insight
Frame-level tags lack the high-level semantic context required for video discovery
“Video or frame-level tags, the sort that you might see on, that I've put on the screen right now, didn't actually solve their problem. Because that was, it was too low level in like a, in a, in not in terms of a granularity sense, but in terms of how much addi…”
David LuanMay 28, 2015▶ 11:05David Luan, Dextro // Real-World Video Understanding (FirstMark / Data Driven NYC)
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 400 conversations transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.