Mar 19, 2019 · 20m · mad

Building Machines That Can Read and Write // Sean Gourley, Primer (FirstMark's Data Driven NYC)

Sean Gourley · 16m spoken Matt Turck · 43s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC presentation, Primer founder Sean Gourley demonstrates how automated natural language understanding and fact-aware generation enable machines to read, synthesize, and write complex intelligence reports at scale.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.9% of the talking time here. How this is scored →

Matt as informed peer 0.8 Guest teaching 0.5 Guest disagreement 0.2 Matt pushing back 0.5
05100:0010:0020:000:09–2:11 · Matt as informed peer 0/10 Sean Gourley Takes Stage and Primer Company Overview Sean Gourley introduces Primer and outlines the core challenge of exponential data growth versus linear human capacity. This is a solo presentation monologue with no host presence or interaction.2:11–4:35 · Matt as informed peer 0/10 Select Enterprise and Government Customers Gourley outlines key customers across sovereign wealth funds, enterprise, and intelligence agencies, demonstrating multi-document report generation from thousands of native Russian sources. The segment is entirely monologued by the guest.4:35–7:32 · Matt as informed peer 0/10 Fact-Aware Language Generation vs Generic Language Models Gourley explains fact-aware language generation in contrast to open-ended models like GPT-2, conducting a live audience comparison test between human and machine headlines. No host participation occurs during this presentation segment.7:32–9:57 · Matt as informed peer 0/10 Building a Self-Writing Wikipedia Knowledge Base Gourley demonstrates how Primer automatically generated 40,000 scientist profiles to highlight and correct systemic demographic recall biases in Wikipedia. This segment remains a uninterrupted monologue by the guest.9:57–12:03 · Matt as informed peer 0/10 Automated Knowledge Graph Updates and Model Performance Gourley discusses benchmark results against Wikidata and details how Primer detects synthetic generated text by checking for non-factual seeds, such as misidentifying Jim Mattis's role. This is an uninterrupted presentation monologue.12:03–20:30 · Matt as informed peer 5/10 Detecting Synthetic Text and Mapping Information Propagation Matt Turck opens the Q&A section with informed questions regarding operational deployments, graph database backends, and dialectical language nuances across Arabic and Chinese. Gourley patiently addresses each technical inquiry and audience question in a highly collaborative manner.0:09–2:11 · Guest teaching 0/10 Sean Gourley Takes Stage and Primer Company Overview Sean Gourley introduces Primer and outlines the core challenge of exponential data growth versus linear human capacity. This is a solo presentation monologue with no host presence or interaction.2:11–4:35 · Guest teaching 0/10 Select Enterprise and Government Customers Gourley outlines key customers across sovereign wealth funds, enterprise, and intelligence agencies, demonstrating multi-document report generation from thousands of native Russian sources. The segment is entirely monologued by the guest.4:35–7:32 · Guest teaching 0/10 Fact-Aware Language Generation vs Generic Language Models Gourley explains fact-aware language generation in contrast to open-ended models like GPT-2, conducting a live audience comparison test between human and machine headlines. No host participation occurs during this presentation segment.7:32–9:57 · Guest teaching 0/10 Building a Self-Writing Wikipedia Knowledge Base Gourley demonstrates how Primer automatically generated 40,000 scientist profiles to highlight and correct systemic demographic recall biases in Wikipedia. This segment remains a uninterrupted monologue by the guest.9:57–12:03 · Guest teaching 0/10 Automated Knowledge Graph Updates and Model Performance Gourley discusses benchmark results against Wikidata and details how Primer detects synthetic generated text by checking for non-factual seeds, such as misidentifying Jim Mattis's role. This is an uninterrupted presentation monologue.12:03–20:30 · Guest teaching 3/10 Detecting Synthetic Text and Mapping Information Propagation Matt Turck opens the Q&A section with informed questions regarding operational deployments, graph database backends, and dialectical language nuances across Arabic and Chinese. Gourley patiently addresses each technical inquiry and audience question in a highly collaborative manner.0:09–2:11 · Guest disagreement 0/10 Sean Gourley Takes Stage and Primer Company Overview Sean Gourley introduces Primer and outlines the core challenge of exponential data growth versus linear human capacity. This is a solo presentation monologue with no host presence or interaction.2:11–4:35 · Guest disagreement 0/10 Select Enterprise and Government Customers Gourley outlines key customers across sovereign wealth funds, enterprise, and intelligence agencies, demonstrating multi-document report generation from thousands of native Russian sources. The segment is entirely monologued by the guest.4:35–7:32 · Guest disagreement 0/10 Fact-Aware Language Generation vs Generic Language Models Gourley explains fact-aware language generation in contrast to open-ended models like GPT-2, conducting a live audience comparison test between human and machine headlines. No host participation occurs during this presentation segment.7:32–9:57 · Guest disagreement 0/10 Building a Self-Writing Wikipedia Knowledge Base Gourley demonstrates how Primer automatically generated 40,000 scientist profiles to highlight and correct systemic demographic recall biases in Wikipedia. This segment remains a uninterrupted monologue by the guest.9:57–12:03 · Guest disagreement 0/10 Automated Knowledge Graph Updates and Model Performance Gourley discusses benchmark results against Wikidata and details how Primer detects synthetic generated text by checking for non-factual seeds, such as misidentifying Jim Mattis's role. This is an uninterrupted presentation monologue.12:03–20:30 · Guest disagreement 1/10 Detecting Synthetic Text and Mapping Information Propagation Matt Turck opens the Q&A section with informed questions regarding operational deployments, graph database backends, and dialectical language nuances across Arabic and Chinese. Gourley patiently addresses each technical inquiry and audience question in a highly collaborative manner.0:09–2:11 · Matt pushing back 0/10 Sean Gourley Takes Stage and Primer Company Overview Sean Gourley introduces Primer and outlines the core challenge of exponential data growth versus linear human capacity. This is a solo presentation monologue with no host presence or interaction.2:11–4:35 · Matt pushing back 0/10 Select Enterprise and Government Customers Gourley outlines key customers across sovereign wealth funds, enterprise, and intelligence agencies, demonstrating multi-document report generation from thousands of native Russian sources. The segment is entirely monologued by the guest.4:35–7:32 · Matt pushing back 0/10 Fact-Aware Language Generation vs Generic Language Models Gourley explains fact-aware language generation in contrast to open-ended models like GPT-2, conducting a live audience comparison test between human and machine headlines. No host participation occurs during this presentation segment.7:32–9:57 · Matt pushing back 0/10 Building a Self-Writing Wikipedia Knowledge Base Gourley demonstrates how Primer automatically generated 40,000 scientist profiles to highlight and correct systemic demographic recall biases in Wikipedia. This segment remains a uninterrupted monologue by the guest.9:57–12:03 · Matt pushing back 0/10 Automated Knowledge Graph Updates and Model Performance Gourley discusses benchmark results against Wikidata and details how Primer detects synthetic generated text by checking for non-factual seeds, such as misidentifying Jim Mattis's role. This is an uninterrupted presentation monologue.12:03–20:30 · Matt pushing back 3/10 Detecting Synthetic Text and Mapping Information Propagation Matt Turck opens the Q&A section with informed questions regarding operational deployments, graph database backends, and dialectical language nuances across Arabic and Chinese. Gourley patiently addresses each technical inquiry and audience question in a highly collaborative manner.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 9.3% · guest 90.7%12:00 · Matt 9.3% · guest 90.7%15:00 · Matt 16.5% · guest 83.5%15:00 · Matt 16.5% · guest 83.5%18:00 · Matt 1.8% · guest 98.2%18:00 · Matt 1.8% · guest 98.2%
Sharpest disagreement ▶ 16:57 Gentle correction on regional dialect coverage

Sean Gourley mildly reframes the host's question regarding dialect support by clarifying the specific regional Arabic variants and Chinese script support Primer handles rather than broad generic language categories.

Hardest push from Matt ▶ 16:48 Probing on language variant limitations

Host Matt Turck presses Gourley on how Primer handles distinct regional dialects like Cantonese vs Mandarin or Iraqi vs Moroccan Arabic.

Biggest teaching moment ▶ 16:07 Explaining graph layer limitations over ElasticSearch

Sean Gourley educates the host on system design, explaining why standard text search engines like ElasticSearch stall when executing relational graph joins on derived entity data.

Matt holds his own ▶ 15:54 Inquiring about underlying graph database mechanics

Matt Turck demonstrates domain expertise by correctly anticipating and asking whether a structured graph database system is integrated behind the text generation engine.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Sean Gourley Takes Stage and Primer Company Overview 0000 Sean Gourley introduces Primer and outlines the core challenge of exponential data growth versus linear human capacity. This is a solo presentation monologue with no host presence or interaction.
Select Enterprise and Government Customers 0000 Gourley outlines key customers across sovereign wealth funds, enterprise, and intelligence agencies, demonstrating multi-document report generation from thousands of native Russian sources. The segment is entirely monologued by the guest.
Fact-Aware Language Generation vs Generic Language Models 0000 Gourley explains fact-aware language generation in contrast to open-ended models like GPT-2, conducting a live audience comparison test between human and machine headlines. No host participation occurs during this presentation segment.
Building a Self-Writing Wikipedia Knowledge Base 0000 Gourley demonstrates how Primer automatically generated 40,000 scientist profiles to highlight and correct systemic demographic recall biases in Wikipedia. This segment remains a uninterrupted monologue by the guest.
Automated Knowledge Graph Updates and Model Performance 0000 Gourley discusses benchmark results against Wikidata and details how Primer detects synthetic generated text by checking for non-factual seeds, such as misidentifying Jim Mattis's role. This is an uninterrupted presentation monologue.
Detecting Synthetic Text and Mapping Information Propagation 5313 Matt Turck opens the Q&A section with informed questions regarding operational deployments, graph database backends, and dialectical language nuances across Arabic and Chinese. Gourley patiently addresses each technical inquiry and audience question in a highly collaborative manner.

Statements from this episode (15)

Assertion Supported
Gourley: Primer has raised $54M to date, led by Lux Capital
“And you know, we've raised about fifty-four million dollars to date, most notably led with a Series B by Lux based out of here in New York.”
Sean Gourley Mar 19, 2019 ▶ 0:19
Disclosure
Primer processes text natively across English, Chinese, Russian, and Arabic
“So we do this in not just English, but we do it in also Chinese, Russian, and new for us this year, Has been Arabic.”
Sean Gourley Mar 19, 2019 ▶ 1:51
Disclosure
Gourley: Primer counts GIC, Walmart, and US intelligence agencies as customers
“We've got some major customers that we can talk about, some that we can't but GIC is the biggest sovereign wealth fund out of Singapore. Walmart, obviously half a trillion of revenue, and also through our partnerships within Qtel we're deployed into a large nu…”
Sean Gourley Mar 19, 2019 ▶ 2:12
Assertion Not checkable as stated
Gourley: Primer can condense 17,000 Russian documents into a one-page report
“We can take a set of 17,000 Russian language documents talking about Syria, and we can write a one-page report as an output.”
Sean Gourley Mar 19, 2019 ▶ 2:56
Assertion Not checkable as stated
Gourley: OpenAI's GPT-2 was not fact-aware
“Some of the colleagues out of OpenAI released GPT-II. That was language generation, but it wasn't fact-aware.”
Sean Gourley Mar 19, 2019 ▶ 4:43
Assertion Not checkable as stated
Gourley: Primer headline generator achieves 84% okay-or-better rating versus 87% for humans
“So on humans, we do about 87% are okay or better. At Primer we do about 84% are okay or better.”
Sean Gourley Mar 19, 2019 ▶ 6:55
Assertion Not checkable as stated
Gourley: Primer headlines beat human writers 29% of the time in side-by-side tests
“If we do this in a side-by-side comparison we can see around 20, 29 percent of the time primer is better and we can see around 44% of the time on the high-volume model that humans are better.”
Sean Gourley Mar 19, 2019 ▶ 7:05
Assertion Not checkable as stated
Primer generated 40,000 new scientist profiles using RNNs and scientific papers
“So we took 30,000 scientists on Wikipedia and took their profiles as of January 18. And then input into this, we put 650,000 scientific papers and around 1.2 million news articles that talked about the scientists. And then we used an RNN-based generative model…”
Sean Gourley Mar 19, 2019 ▶ 7:56
Assertion Supported
Gourley: Wikipedia had no page for Andrej Karpathy when Primer built his bio
“Well, we could actually compare that to what was on Wikipedia at the time when it was generated, and the answer was nothing. Right? Now, if you go there now, he's taking care of this. We brought it up with Tesla, and their PR department got onto this, but ther…”
Sean Gourley Mar 19, 2019 ▶ 8:46
Assertion Not checkable as stated
Sean Gourley: Wikipedia has a massive recall problem despite high precision
“And this is actually not just endemic to Andre, there's nothing particularly about him, but it's actually, Wikipedia has a massive recall problem. We think a lot about its precision, and we measure its precision, but its recall is actually pretty bad when it l…”
Sean Gourley Mar 19, 2019 ▶ 9:00
Assertion Not checkable as stated
Gourley: Machine-generated content is becoming indistinguishable from true reality
“And it's important to do this because we're living in a world, which I would term, of the generation age, where it is increasingly cheap and easy to generate a reality that is almost indistinguishable from true reality by machines.”
Sean Gourley Mar 19, 2019 ▶ 11:04
Assertion Supported
Sean Gourley: Early deepfakes fail to simulate natural human blinking patterns
“He doesn't quite blink enough, by the way, if you've actually kind of followed this. The blinking is, is not quite right.”
Sean Gourley Mar 19, 2019 ▶ 11:47
Assertion Supported
Gourley: Information propagation in networks is central to Russian military strategy
“And it's important to know the context of where information came from, because as information is seeded into the networks, it starts to propagate, and this becomes key attack vectors that are actually unfolding, and if you listen to some of the major Russian g…”
Sean Gourley Mar 19, 2019 ▶ 12:41
Assertion Not checkable as stated
Gourley: Primer cut bank analyst report summary time, saving $6M
“We're able to reduce that down to about 25% of the time. And that nets out at about a six million dollar saving for that bank.”
Sean Gourley Mar 19, 2019 ▶ 14:29
Insight
Gourley: Graph joins on Elasticsearch slow text analysis to unusable levels
“If you try and do that on top of, like, an elastic search, you know, your joins that are going to emerge on top of the derived information are going to just slow that thing down to being unusable.”
Sean Gourley Mar 19, 2019 ▶ 16:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.