Dec 5, 2013 · 22m · mad

Gilad Lotan, Betaworks // Data Driven NYC 20 // Nov 2013

Gilad Lotan · 17m spoken Matt Turck · 1m spoken AV Technician · 0s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At Data Driven NYC, Betaworks Chief Data Scientist Gilad Lotan demonstrates how graph theory, network visualization tools, and term co-occurrence modeling can be applied across portfolio products like Twitter, Giphy, and Digg Reader to contextualize social data and power recommendation engines.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 5.9% of the talking time here. How this is scored →

Matt as informed peer 0.7 Guest teaching 4.7 Guest disagreement 0.0 Matt pushing back 0.3
05100:0010:0020:002:59–11:22 · Matt as informed peer 0/10 Graph Analysis Fundamentals and Social Networks In this monologue presentation, Gilad demonstrates network graph analysis using Gephi to map Python Twitter users. The host does not speak during this segment, necessitating zero scores for host expertise and pushback. Gilad educates the audience on modularity clustering, revealing distinct groups such as Japanese developers, Python hackers, and Monty Python fans.11:22–16:11 · Matt as informed peer 0/10 Deep-Dive into Twitter Python Community Clusters Gilad continues his presentation uninterrupted, explaining how tag co-occurrence analysis on Giphy separates content into distinct clusters like LOL versus funny. Because the host remains silent throughout the presentation section, all host-side metrics are strictly zero. The dynamic is purely instructive and collaborative.16:11–22:03 · Matt as informed peer 2/10 Digg Reader Feed Clustering and Presentation Conclusion Matt Turck opens the Q&A by asking about Gilad's daily toolkit and inquiring whether Betaworks' shared data science model could scale to large conglomerates like Viacom. The interaction is entirely friendly and constructive, with Gilad detailing how different-sized startups utilize central data resources. Matt asks intelligent organizational questions without pushing back or challenging premises.2:59–11:22 · Guest teaching 6/10 Graph Analysis Fundamentals and Social Networks In this monologue presentation, Gilad demonstrates network graph analysis using Gephi to map Python Twitter users. The host does not speak during this segment, necessitating zero scores for host expertise and pushback. Gilad educates the audience on modularity clustering, revealing distinct groups such as Japanese developers, Python hackers, and Monty Python fans.11:22–16:11 · Guest teaching 5/10 Deep-Dive into Twitter Python Community Clusters Gilad continues his presentation uninterrupted, explaining how tag co-occurrence analysis on Giphy separates content into distinct clusters like LOL versus funny. Because the host remains silent throughout the presentation section, all host-side metrics are strictly zero. The dynamic is purely instructive and collaborative.16:11–22:03 · Guest teaching 3/10 Digg Reader Feed Clustering and Presentation Conclusion Matt Turck opens the Q&A by asking about Gilad's daily toolkit and inquiring whether Betaworks' shared data science model could scale to large conglomerates like Viacom. The interaction is entirely friendly and constructive, with Gilad detailing how different-sized startups utilize central data resources. Matt asks intelligent organizational questions without pushing back or challenging premises.2:59–11:22 · Guest disagreement 0/10 Graph Analysis Fundamentals and Social Networks In this monologue presentation, Gilad demonstrates network graph analysis using Gephi to map Python Twitter users. The host does not speak during this segment, necessitating zero scores for host expertise and pushback. Gilad educates the audience on modularity clustering, revealing distinct groups such as Japanese developers, Python hackers, and Monty Python fans.11:22–16:11 · Guest disagreement 0/10 Deep-Dive into Twitter Python Community Clusters Gilad continues his presentation uninterrupted, explaining how tag co-occurrence analysis on Giphy separates content into distinct clusters like LOL versus funny. Because the host remains silent throughout the presentation section, all host-side metrics are strictly zero. The dynamic is purely instructive and collaborative.16:11–22:03 · Guest disagreement 0/10 Digg Reader Feed Clustering and Presentation Conclusion Matt Turck opens the Q&A by asking about Gilad's daily toolkit and inquiring whether Betaworks' shared data science model could scale to large conglomerates like Viacom. The interaction is entirely friendly and constructive, with Gilad detailing how different-sized startups utilize central data resources. Matt asks intelligent organizational questions without pushing back or challenging premises.2:59–11:22 · Matt pushing back 0/10 Graph Analysis Fundamentals and Social Networks In this monologue presentation, Gilad demonstrates network graph analysis using Gephi to map Python Twitter users. The host does not speak during this segment, necessitating zero scores for host expertise and pushback. Gilad educates the audience on modularity clustering, revealing distinct groups such as Japanese developers, Python hackers, and Monty Python fans.11:22–16:11 · Matt pushing back 0/10 Deep-Dive into Twitter Python Community Clusters Gilad continues his presentation uninterrupted, explaining how tag co-occurrence analysis on Giphy separates content into distinct clusters like LOL versus funny. Because the host remains silent throughout the presentation section, all host-side metrics are strictly zero. The dynamic is purely instructive and collaborative.16:11–22:03 · Matt pushing back 1/10 Digg Reader Feed Clustering and Presentation Conclusion Matt Turck opens the Q&A by asking about Gilad's daily toolkit and inquiring whether Betaworks' shared data science model could scale to large conglomerates like Viacom. The interaction is entirely friendly and constructive, with Gilad detailing how different-sized startups utilize central data resources. Matt asks intelligent organizational questions without pushing back or challenging premises.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 12.1% · guest 87.9%0:00 · Matt 12.1% · guest 87.9%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%12:00 · Matt 0% · guest 100%15:00 · Matt 9.5% · guest 90.5%15:00 · Matt 9.5% · guest 90.5%18:00 · Matt 22% · guest 78%18:00 · Matt 22% · guest 78%21:00 · Matt 3.6% · guest 96.4%21:00 · Matt 3.6% · guest 96.4%
Sharpest disagreement ▶ 21:40 Playful deflection regarding Viacom comparison

In an entirely agreeable session, Gilad playfully sidesteps Matt's comparison to Viacom by noting he doesn't know if the model works there until Betaworks becomes the next Viacom.

Hardest push from Matt ▶ 20:09 Questioning horizontal scalability of shared data units

Matt politely probes the limits of Betaworks' shared data science model, asking whether providing centralized data services is actually replicable for traditional conglomerates.

Biggest teaching moment ▶ 14:45 Explaining the data distinction between 'LOL' and 'funny'

Gilad demonstrates how graph modularity revealed that 'LOL' tagged content predominantly featured animals and bizarre clips, whereas 'funny' tagged content featured human faces laughing.

Matt holds his own ▶ 20:09 Framing the organizational dynamics of startup data infrastructure

Matt frames a thoughtful macro question about whether centralized data science functions as a horizontal capability across diverse business units.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Graph Analysis Fundamentals and Social Networks 0600 In this monologue presentation, Gilad demonstrates network graph analysis using Gephi to map Python Twitter users. The host does not speak during this segment, necessitating zero scores for host expertise and pushback. Gilad educates the audience on modularity clustering, revealing distinct groups such as Japanese developers, Python hackers, and Monty Python fans.
Deep-Dive into Twitter Python Community Clusters 0500 Gilad continues his presentation uninterrupted, explaining how tag co-occurrence analysis on Giphy separates content into distinct clusters like LOL versus funny. Because the host remains silent throughout the presentation section, all host-side metrics are strictly zero. The dynamic is purely instructive and collaborative.
Digg Reader Feed Clustering and Presentation Conclusion 2301 Matt Turck opens the Q&A by asking about Gilad's daily toolkit and inquiring whether Betaworks' shared data science model could scale to large conglomerates like Viacom. The interaction is entirely friendly and constructive, with Gilad detailing how different-sized startups utilize central data resources. Matt asks intelligent organizational questions without pushing back or challenging premises.

Statements from this episode (9)

Disclosure
Lotan: Betaworks was an early investor in Tumblr and Kickstarter
“So we were one of the earliest investors in Tumblr and Kickstarter, and a whole bunch of, ah, companies across, ah, the tech scene in New York City.”
Gilad Lotan Dec 5, 2013 ▶ 0:56
Assertion Partly supported
Lotan: TweetDeck was built at Betaworks and sold to Twitter
“TweetDeck was actually built at Betaworks and then sold to Twitter.”
Gilad Lotan Dec 5, 2013 ▶ 1:13
Insight
Lotan: Gephi is ideal for quick exploratory graph analysis
“It's a great tool to pull in graphs and do some exploratory data analysis, so the section where you're sort of trying to explore a data set, you don't want to put too much effort into it and build something for it, you can just easily use this open source tool…”
Gilad Lotan Dec 5, 2013 ▶ 6:14
Insight
Lotan: Network graph analysis effectively isolates spam and off-topic data
“And it's actually also a great way to get rid of spam, things that aren't related not that I think that Python snakes are spammy, because they're pretty awesome, but it's a way to sort of, to identify them as separate from the context that we're trying to unde…”
Gilad Lotan Dec 5, 2013 ▶ 12:23
Assertion Not checkable as stated
Early Giphy tags were manually labeled by humans for cleaner data
“They're manually labeled, so we're getting these labels from, ah, actual humans, ah, so they're really, really clean and great data.”
Gilad Lotan Dec 5, 2013 ▶ 13:15
Assertion Not checkable as stated
Giphy's top tag clusters in 2013 were funny content, art, and movies
“We get three dominant sort of clusters in this data, and it's, I don't think it's surprising There's the lol, right, lots of just funny, funny stuff. There's, like, kind of pretty, beautiful content, like photography, just artistic stuff, and then lots of movi…”
Gilad Lotan Dec 5, 2013 ▶ 13:47
Assertion Not checkable as stated
Lotan: Giphy GIFs tagged 'funny' feature laughing people, while 'LOL' features animals
“Content that has the tag funny usually consists of sort of people laughing here, like funny in terms of funny, they're, you know, you see their face and you see them laughing. LOL tended to have, for some weird reason, have animals doing weird stuff. Or young …”
Gilad Lotan Dec 5, 2013 ▶ 15:09
Disclosure
Hypertable is highly efficient at writing content for graph storage
“I currently use Hypertable which is just super efficient at writing content.”
Gilad Lotan Dec 5, 2013 ▶ 19:48
Assertion Supported
Digg operated with only 15 employees under Betaworks in 2013
“Digg, for example, is fairly large, ah, in terms of Betaworks companies. It's 15 people.”
Gilad Lotan Dec 5, 2013 ▶ 20:58
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.