Mar 19, 2019 · 20m · mad
Building Machines That Can Read and Write // Sean Gourley, Primer (FirstMark's Data Driven NYC)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Data Driven NYC presentation, Primer founder Sean Gourley demonstrates how automated natural language understanding and fact-aware generation enable machines to read, synthesize, and write complex intelligence reports at scale.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.9% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Sean Gourley mildly reframes the host's question regarding dialect support by clarifying the specific regional Arabic variants and Chinese script support Primer handles rather than broad generic language categories.
Hardest push from Matt ▶ 16:48 Probing on language variant limitationsHost Matt Turck presses Gourley on how Primer handles distinct regional dialects like Cantonese vs Mandarin or Iraqi vs Moroccan Arabic.
Biggest teaching moment ▶ 16:07 Explaining graph layer limitations over ElasticSearchSean Gourley educates the host on system design, explaining why standard text search engines like ElasticSearch stall when executing relational graph joins on derived entity data.
Matt holds his own ▶ 15:54 Inquiring about underlying graph database mechanicsMatt Turck demonstrates domain expertise by correctly anticipating and asking whether a structured graph database system is integrated behind the text generation engine.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Sean Gourley Takes Stage and Primer Company Overview | 0 | 0 | 0 | 0 | Sean Gourley introduces Primer and outlines the core challenge of exponential data growth versus linear human capacity. This is a solo presentation monologue with no host presence or interaction. | |
| Select Enterprise and Government Customers | 0 | 0 | 0 | 0 | Gourley outlines key customers across sovereign wealth funds, enterprise, and intelligence agencies, demonstrating multi-document report generation from thousands of native Russian sources. The segment is entirely monologued by the guest. | |
| Fact-Aware Language Generation vs Generic Language Models | 0 | 0 | 0 | 0 | Gourley explains fact-aware language generation in contrast to open-ended models like GPT-2, conducting a live audience comparison test between human and machine headlines. No host participation occurs during this presentation segment. | |
| Building a Self-Writing Wikipedia Knowledge Base | 0 | 0 | 0 | 0 | Gourley demonstrates how Primer automatically generated 40,000 scientist profiles to highlight and correct systemic demographic recall biases in Wikipedia. This segment remains a uninterrupted monologue by the guest. | |
| Automated Knowledge Graph Updates and Model Performance | 0 | 0 | 0 | 0 | Gourley discusses benchmark results against Wikidata and details how Primer detects synthetic generated text by checking for non-factual seeds, such as misidentifying Jim Mattis's role. This is an uninterrupted presentation monologue. | |
| Detecting Synthetic Text and Mapping Information Propagation | 5 | 3 | 1 | 3 | Matt Turck opens the Q&A section with informed questions regarding operational deployments, graph database backends, and dialectical language nuances across Arabic and Chinese. Gourley patiently addresses each technical inquiry and audience question in a highly collaborative manner. |