May 13, 2019 · 23m · mad

AI for Sensitive Information // Apoorv Agarwal, Text IQ (FirstMark's Data Driven NYC)

Apoorv Agarwal · 17m spoken Matt Turck · 36s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Apoorv Agarwal, CEO of Text IQ, presents his company's contextual AI platform designed to automatically detect sensitive risks, legal liabilities, and compliance threats across unstructured enterprise data.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 2.8% of the talking time here. How this is scored →

Matt as informed peer 0.7 Guest teaching 3.2 Guest disagreement 0.7 Matt pushing back 0.7
05100:0010:0020:000:38–2:40 · Matt as informed peer 0/10 Company Origin & Columbia PhD Research Apoorv presents a monologue covering his Columbia PhD research on relationship extraction from unstructured text and how government grants launched Text IQ. Host Matt Turck does not participate in this section, resulting in zero host scores.2:40–5:39 · Matt as informed peer 0/10 High-Stakes Disasters & The Status Quo Problem Apoorv details the shortcomings of legacy search term methods and 13,000-person review teams using the Ruben sandwich coded email example. As this is a presentation monologue, host participation scores are strictly zero.5:39–7:48 · Matt as informed peer 0/10 The Text IQ Platform Architecture & Applications Apoorv explains the Text IQ platform architecture and presents a case study saving a healthcare client $3M while running 10x faster. Host metrics remain zero during the continuous monologue.7:48–10:11 · Matt as informed peer 0/10 Long-Term Vision: Unstructured to Structured Data Apoorv outlines the company's long-term product vision of converting unstructured to structured data using unsupervised machine learning, drawing a comparison to Splunk. Host is silent.10:11–13:06 · Matt as informed peer 0/10 Core Technical Challenges in Sensitive AI Apoorv reviews key technical challenges including multilingual semantics, short text, model interpretability, and continuous human-in-the-loop learning. Host is not present on mic.13:06–23:12 · Matt as informed peer 4/10 Company Expansion & Recruitment Pitch Matt Turck steps in for Q&A, pressing Apoorv on how AI handles high-stakes sensitive data given typical 80-90% accuracy ceilings. Apoorv counters gently that the enterprise benchmark isn't perfection, but beating legacy human keyword search.0:38–2:40 · Guest teaching 2/10 Company Origin & Columbia PhD Research Apoorv presents a monologue covering his Columbia PhD research on relationship extraction from unstructured text and how government grants launched Text IQ. Host Matt Turck does not participate in this section, resulting in zero host scores.2:40–5:39 · Guest teaching 3/10 High-Stakes Disasters & The Status Quo Problem Apoorv details the shortcomings of legacy search term methods and 13,000-person review teams using the Ruben sandwich coded email example. As this is a presentation monologue, host participation scores are strictly zero.5:39–7:48 · Guest teaching 3/10 The Text IQ Platform Architecture & Applications Apoorv explains the Text IQ platform architecture and presents a case study saving a healthcare client $3M while running 10x faster. Host metrics remain zero during the continuous monologue.7:48–10:11 · Guest teaching 3/10 Long-Term Vision: Unstructured to Structured Data Apoorv outlines the company's long-term product vision of converting unstructured to structured data using unsupervised machine learning, drawing a comparison to Splunk. Host is silent.10:11–13:06 · Guest teaching 3/10 Core Technical Challenges in Sensitive AI Apoorv reviews key technical challenges including multilingual semantics, short text, model interpretability, and continuous human-in-the-loop learning. Host is not present on mic.13:06–23:12 · Guest teaching 5/10 Company Expansion & Recruitment Pitch Matt Turck steps in for Q&A, pressing Apoorv on how AI handles high-stakes sensitive data given typical 80-90% accuracy ceilings. Apoorv counters gently that the enterprise benchmark isn't perfection, but beating legacy human keyword search.0:38–2:40 · Guest disagreement 0/10 Company Origin & Columbia PhD Research Apoorv presents a monologue covering his Columbia PhD research on relationship extraction from unstructured text and how government grants launched Text IQ. Host Matt Turck does not participate in this section, resulting in zero host scores.2:40–5:39 · Guest disagreement 1/10 High-Stakes Disasters & The Status Quo Problem Apoorv details the shortcomings of legacy search term methods and 13,000-person review teams using the Ruben sandwich coded email example. As this is a presentation monologue, host participation scores are strictly zero.5:39–7:48 · Guest disagreement 0/10 The Text IQ Platform Architecture & Applications Apoorv explains the Text IQ platform architecture and presents a case study saving a healthcare client $3M while running 10x faster. Host metrics remain zero during the continuous monologue.7:48–10:11 · Guest disagreement 0/10 Long-Term Vision: Unstructured to Structured Data Apoorv outlines the company's long-term product vision of converting unstructured to structured data using unsupervised machine learning, drawing a comparison to Splunk. Host is silent.10:11–13:06 · Guest disagreement 0/10 Core Technical Challenges in Sensitive AI Apoorv reviews key technical challenges including multilingual semantics, short text, model interpretability, and continuous human-in-the-loop learning. Host is not present on mic.13:06–23:12 · Guest disagreement 3/10 Company Expansion & Recruitment Pitch Matt Turck steps in for Q&A, pressing Apoorv on how AI handles high-stakes sensitive data given typical 80-90% accuracy ceilings. Apoorv counters gently that the enterprise benchmark isn't perfection, but beating legacy human keyword search.0:38–2:40 · Matt pushing back 0/10 Company Origin & Columbia PhD Research Apoorv presents a monologue covering his Columbia PhD research on relationship extraction from unstructured text and how government grants launched Text IQ. Host Matt Turck does not participate in this section, resulting in zero host scores.2:40–5:39 · Matt pushing back 0/10 High-Stakes Disasters & The Status Quo Problem Apoorv details the shortcomings of legacy search term methods and 13,000-person review teams using the Ruben sandwich coded email example. As this is a presentation monologue, host participation scores are strictly zero.5:39–7:48 · Matt pushing back 0/10 The Text IQ Platform Architecture & Applications Apoorv explains the Text IQ platform architecture and presents a case study saving a healthcare client $3M while running 10x faster. Host metrics remain zero during the continuous monologue.7:48–10:11 · Matt pushing back 0/10 Long-Term Vision: Unstructured to Structured Data Apoorv outlines the company's long-term product vision of converting unstructured to structured data using unsupervised machine learning, drawing a comparison to Splunk. Host is silent.10:11–13:06 · Matt pushing back 0/10 Core Technical Challenges in Sensitive AI Apoorv reviews key technical challenges including multilingual semantics, short text, model interpretability, and continuous human-in-the-loop learning. Host is not present on mic.13:06–23:12 · Matt pushing back 4/10 Company Expansion & Recruitment Pitch Matt Turck steps in for Q&A, pressing Apoorv on how AI handles high-stakes sensitive data given typical 80-90% accuracy ceilings. Apoorv counters gently that the enterprise benchmark isn't perfection, but beating legacy human keyword search.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 18.1% · guest 81.9%12:00 · Matt 18.1% · guest 81.9%15:00 · Matt 0% · guest 100%15:00 · Matt 0% · guest 100%18:00 · Matt 0.6% · guest 99.4%18:00 · Matt 0.6% · guest 99.4%21:00 · Matt 4.3% · guest 95.7%21:00 · Matt 4.3% · guest 95.7%
Sharpest disagreement ▶ 14:37 Reframing the AI accuracy benchmark

Apoorv rejects the premise that AI must achieve 100% accuracy in high-stakes environments, pushing back that the real threshold to win deals is merely outperforming flawed manual search terms.

Hardest push from Matt ▶ 14:05 Host challenges guest on playing with fire regarding missing sensitive data

Matt Turck pushes Apoorv on the inherent risk of AI missing sensitive needles when models usually max out at 80-90% accuracy.

Biggest teaching moment ▶ 14:37 Educating on real-world enterprise procurement baselines

Apoorv clarifies to the host that enterprise legal teams buy based on relative improvement over 500 manual reviewers rather than absolute model perfection.

Matt holds his own ▶ 14:05 Host highlights AI performance ceiling in high-risk contexts

Matt demonstrates technical awareness of machine learning accuracy ceilings (80-90%) to press the guest on workflow vulnerabilities.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Company Origin & Columbia PhD Research 0200 Apoorv presents a monologue covering his Columbia PhD research on relationship extraction from unstructured text and how government grants launched Text IQ. Host Matt Turck does not participate in this section, resulting in zero host scores.
High-Stakes Disasters & The Status Quo Problem 0310 Apoorv details the shortcomings of legacy search term methods and 13,000-person review teams using the Ruben sandwich coded email example. As this is a presentation monologue, host participation scores are strictly zero.
The Text IQ Platform Architecture & Applications 0300 Apoorv explains the Text IQ platform architecture and presents a case study saving a healthcare client $3M while running 10x faster. Host metrics remain zero during the continuous monologue.
Long-Term Vision: Unstructured to Structured Data 0300 Apoorv outlines the company's long-term product vision of converting unstructured to structured data using unsupervised machine learning, drawing a comparison to Splunk. Host is silent.
Core Technical Challenges in Sensitive AI 0300 Apoorv reviews key technical challenges including multilingual semantics, short text, model interpretability, and continuous human-in-the-loop learning. Host is not present on mic.
Company Expansion & Recruitment Pitch 4534 Matt Turck steps in for Q&A, pressing Apoorv on how AI handles high-stakes sensitive data given typical 80-90% accuracy ceilings. Apoorv counters gently that the enterprise benchmark isn't perfection, but beating legacy human keyword search.

Statements from this episode (10)

Assertion Supported
Text IQ was launched via a government research grant in 2014
“Around the time in 214, ah, when I was graduating, ah, the government, ah, who had funded a lot of my research, saw a lot of commercial potential in what we had built, ah, they gave us a grant, ah, and that's how we got started.”
Apoorv Agarwal May 13, 2019 ▶ 1:33
Assertion Not checkable as stated
Text IQ achieved profitability by 2017 with a 100% pilot conversion rate
“We've been profitable since 2017 but really the metric that we are proud of is that we, we've had a hundred percent pilot to customer conversion rate.”
Apoorv Agarwal May 13, 2019 ▶ 2:11
Assertion Not checkable as stated
Enterprise compliance still relies mostly on keyword search and manual review
“The status quo method, the most popular method of finding these needles in a haystack remains to be based on search terms, ah, and manual review.”
Apoorv Agarwal May 13, 2019 ▶ 3:24
Assertion Not checkable as stated
A major New York bank hired 13,000 compliance reviewers since 2012
“A major bank in New York alone, they've hired over 13,000 people since 2012 looking for sensitive needles in this haystack.”
Apoorv Agarwal May 13, 2019 ▶ 3:37
Assertion Supported
Text IQ infers employee roles and relationships to flag sensitive information
“We can analyze human communication, we can infer things about roles of people within an organization, their relationship, and that allows us to zoom into or find many different kinds of sensitive information.”
Apoorv Agarwal May 13, 2019 ▶ 5:21
What-if
Legacy compliance review for a healthcare client would have cost $4 million
“Had they used the status quo method, which is very much based on search terms and manual review, it would have taken them about 20 weeks To review all these documents. Ah, we later showed them that their search would have missed about 10% of the sensitive info…”
Apoorv Agarwal May 13, 2019 ▶ 6:43
Insight
Structuring enterprise unstructured data requires unsupervised machine learning
“There's just so many different types of documents and types of data that there's just no way we can, you know, tackle this problem using supervised machine learning. So a lot of the machine learning we use in-house you know, is, is mostly unsupervised and that…”
Apoorv Agarwal May 13, 2019 ▶ 9:37
Insight
Agarwal: Analyzing short text requires different ML techniques due to context limits
“Dealing with short text, it's a totally different you know, animal compared to longer text. There's very little context.”
Apoorv Agarwal May 13, 2019 ▶ 10:58
Assertion Not checkable as stated
Agarwal: Text IQ grew 5x over 6-7 months
“We've grown like five times over the past six, seven months.”
Apoorv Agarwal May 13, 2019 ▶ 13:27
Insight
Enterprise AI only needs to outperform keyword searches and human reviewers
“The bar is not to be hundred percent accurate. The bar is search tomes and 500 humans, right? And as it turns out that bar is not that hard to beat with AI.”
Apoorv Agarwal May 13, 2019 ▶ 14:39
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.