Dec 5, 2013 · 22m · mad
Gilad Lotan, Betaworks // Data Driven NYC 20 // Nov 2013
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
At Data Driven NYC, Betaworks Chief Data Scientist Gilad Lotan demonstrates how graph theory, network visualization tools, and term co-occurrence modeling can be applied across portfolio products like Twitter, Giphy, and Digg Reader to contextualize social data and power recommendation engines.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 5.9% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
In an entirely agreeable session, Gilad playfully sidesteps Matt's comparison to Viacom by noting he doesn't know if the model works there until Betaworks becomes the next Viacom.
Hardest push from Matt ▶ 20:09 Questioning horizontal scalability of shared data unitsMatt politely probes the limits of Betaworks' shared data science model, asking whether providing centralized data services is actually replicable for traditional conglomerates.
Biggest teaching moment ▶ 14:45 Explaining the data distinction between 'LOL' and 'funny'Gilad demonstrates how graph modularity revealed that 'LOL' tagged content predominantly featured animals and bizarre clips, whereas 'funny' tagged content featured human faces laughing.
Matt holds his own ▶ 20:09 Framing the organizational dynamics of startup data infrastructureMatt frames a thoughtful macro question about whether centralized data science functions as a horizontal capability across diverse business units.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Graph Analysis Fundamentals and Social Networks | 0 | 6 | 0 | 0 | In this monologue presentation, Gilad demonstrates network graph analysis using Gephi to map Python Twitter users. The host does not speak during this segment, necessitating zero scores for host expertise and pushback. Gilad educates the audience on modularity clustering, revealing distinct groups such as Japanese developers, Python hackers, and Monty Python fans. | |
| Deep-Dive into Twitter Python Community Clusters | 0 | 5 | 0 | 0 | Gilad continues his presentation uninterrupted, explaining how tag co-occurrence analysis on Giphy separates content into distinct clusters like LOL versus funny. Because the host remains silent throughout the presentation section, all host-side metrics are strictly zero. The dynamic is purely instructive and collaborative. | |
| Digg Reader Feed Clustering and Presentation Conclusion | 2 | 3 | 0 | 1 | Matt Turck opens the Q&A by asking about Gilad's daily toolkit and inquiring whether Betaworks' shared data science model could scale to large conglomerates like Viacom. The interaction is entirely friendly and constructive, with Gilad detailing how different-sized startups utilize central data resources. Matt asks intelligent organizational questions without pushing back or challenging premises. |