Gilad Lotan, Chief Data Scientist at Betaworks, answers an audience question regarding the database technologies he uses to store and index graph data.
Disclosure
Lotan: Betaworks was an early investor in Tumblr and Kickstarter
“So we were one of the earliest investors in Tumblr and Kickstarter, and a whole bunch of, ah, companies across, ah, the tech scene in New York City.”
Assertion Partly supported
Lotan: TweetDeck was built at Betaworks and sold to Twitter
“TweetDeck was actually built at Betaworks and then sold to Twitter.”
Insight
Lotan: Gephi is ideal for quick exploratory graph analysis
“It's a great tool to pull in graphs and do some exploratory data analysis, so the section where you're sort of trying to explore a data set, you don't want to put too much effort into it and build something for it, you can just easily use this open source tool…”
Insight
Lotan: Network graph analysis effectively isolates spam and off-topic data
“And it's actually also a great way to get rid of spam, things that aren't related not that I think that Python snakes are spammy, because they're pretty awesome, but it's a way to sort of, to identify them as separate from the context that we're trying to unde…”
Assertion Not checkable as stated
Early Giphy tags were manually labeled by humans for cleaner data
“They're manually labeled, so we're getting these labels from, ah, actual humans, ah, so they're really, really clean and great data.”
Assertion Not checkable as stated
Giphy's top tag clusters in 2013 were funny content, art, and movies
“We get three dominant sort of clusters in this data, and it's, I don't think it's surprising There's the lol, right, lots of just funny, funny stuff. There's, like, kind of pretty, beautiful content, like photography, just artistic stuff, and then lots of movi…”