Mar 2, 2018 · 23m · mad

Text Analytics for Finance // Amanda Stent, Bloomberg (FirstMark's Data Driven)

Amanda Stent · 17m spoken Matt Turck · 49s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

At a Data Driven NYC event, Bloomberg NLP Researcher Amanda Stent explores how text analytics and natural language processing are applied to unstructured financial data to generate actionable market insights. She outlines core application areas, real-time speed requirements, high-precision entity extraction, and the importance of hybrid human-in-the-loop workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 3.9% of the talking time here. How this is scored →

Matt as informed peer 0.9 Guest teaching 1.0 Guest disagreement 0.6 Matt pushing back 0.1
05100:0010:0020:001:45–5:09 · Matt as informed peer 0/10 Four Key NLP Applications at Bloomberg Monologue presentation segment delivered by Amanda Stent introducing Bloomberg's core mission and NLP enrichment applications. The host is not active during this segment, requiring host-side scores to be zero.5:09–7:30 · Matt as informed peer 0/10 Three Core Challenges in Financial Text Analytics Monologue presentation detailing idiomatic financial language and document understanding challenges. The host is inactive throughout this portion.7:30–9:38 · Matt as informed peer 0/10 The Need for Speed and High-Volume Social Processing Monologue segment focusing on latency, processing speed, and tweet funnels. The host remains silent throughout the talk.9:38–11:54 · Matt as informed peer 0/10 Precision vs. Recall and Ethical Data Science Monologue presentation on precision versus recall, ethical data science, and human-in-the-loop workflows. Host activity is zero.11:54–15:36 · Matt as informed peer 4/10 Presentation Wrap-Up and Transition to Q&A Matt Turck steps in, citing his four-year background at Bloomberg and asking about human-in-the-loop mechanics and data pipelines. Stent gently reframes his question about a single data pipeline, clarifying that Bloomberg operates between 15,000 and 40,000 distinct functions.15:36–20:29 · Matt as informed peer 1/10 Audience Q&A: Social Media Manipulation, Speed Limits, and Market Truth Audience members ask detailed questions on bot manipulation, speed calculus, and market expectations. The host acts only as a moderator, while Stent provides detailed technical and philosophical answers.20:29–21:28 · Matt as informed peer 1/10 Audience Q&A: Competitor Usage of Automated Content Audience questions cover competitor usage of automated text and sentiment extraction from earnings calls. Stent shares informative responses while the host maintains a minimal moderating role.1:45–5:09 · Guest teaching 0/10 Four Key NLP Applications at Bloomberg Monologue presentation segment delivered by Amanda Stent introducing Bloomberg's core mission and NLP enrichment applications. The host is not active during this segment, requiring host-side scores to be zero.5:09–7:30 · Guest teaching 0/10 Three Core Challenges in Financial Text Analytics Monologue presentation detailing idiomatic financial language and document understanding challenges. The host is inactive throughout this portion.7:30–9:38 · Guest teaching 0/10 The Need for Speed and High-Volume Social Processing Monologue segment focusing on latency, processing speed, and tweet funnels. The host remains silent throughout the talk.9:38–11:54 · Guest teaching 0/10 Precision vs. Recall and Ethical Data Science Monologue presentation on precision versus recall, ethical data science, and human-in-the-loop workflows. Host activity is zero.11:54–15:36 · Guest teaching 3/10 Presentation Wrap-Up and Transition to Q&A Matt Turck steps in, citing his four-year background at Bloomberg and asking about human-in-the-loop mechanics and data pipelines. Stent gently reframes his question about a single data pipeline, clarifying that Bloomberg operates between 15,000 and 40,000 distinct functions.15:36–20:29 · Guest teaching 2/10 Audience Q&A: Social Media Manipulation, Speed Limits, and Market Truth Audience members ask detailed questions on bot manipulation, speed calculus, and market expectations. The host acts only as a moderator, while Stent provides detailed technical and philosophical answers.20:29–21:28 · Guest teaching 2/10 Audience Q&A: Competitor Usage of Automated Content Audience questions cover competitor usage of automated text and sentiment extraction from earnings calls. Stent shares informative responses while the host maintains a minimal moderating role.1:45–5:09 · Guest disagreement 0/10 Four Key NLP Applications at Bloomberg Monologue presentation segment delivered by Amanda Stent introducing Bloomberg's core mission and NLP enrichment applications. The host is not active during this segment, requiring host-side scores to be zero.5:09–7:30 · Guest disagreement 0/10 Three Core Challenges in Financial Text Analytics Monologue presentation detailing idiomatic financial language and document understanding challenges. The host is inactive throughout this portion.7:30–9:38 · Guest disagreement 0/10 The Need for Speed and High-Volume Social Processing Monologue segment focusing on latency, processing speed, and tweet funnels. The host remains silent throughout the talk.9:38–11:54 · Guest disagreement 0/10 Precision vs. Recall and Ethical Data Science Monologue presentation on precision versus recall, ethical data science, and human-in-the-loop workflows. Host activity is zero.11:54–15:36 · Guest disagreement 2/10 Presentation Wrap-Up and Transition to Q&A Matt Turck steps in, citing his four-year background at Bloomberg and asking about human-in-the-loop mechanics and data pipelines. Stent gently reframes his question about a single data pipeline, clarifying that Bloomberg operates between 15,000 and 40,000 distinct functions.15:36–20:29 · Guest disagreement 1/10 Audience Q&A: Social Media Manipulation, Speed Limits, and Market Truth Audience members ask detailed questions on bot manipulation, speed calculus, and market expectations. The host acts only as a moderator, while Stent provides detailed technical and philosophical answers.20:29–21:28 · Guest disagreement 1/10 Audience Q&A: Competitor Usage of Automated Content Audience questions cover competitor usage of automated text and sentiment extraction from earnings calls. Stent shares informative responses while the host maintains a minimal moderating role.1:45–5:09 · Matt pushing back 0/10 Four Key NLP Applications at Bloomberg Monologue presentation segment delivered by Amanda Stent introducing Bloomberg's core mission and NLP enrichment applications. The host is not active during this segment, requiring host-side scores to be zero.5:09–7:30 · Matt pushing back 0/10 Three Core Challenges in Financial Text Analytics Monologue presentation detailing idiomatic financial language and document understanding challenges. The host is inactive throughout this portion.7:30–9:38 · Matt pushing back 0/10 The Need for Speed and High-Volume Social Processing Monologue segment focusing on latency, processing speed, and tweet funnels. The host remains silent throughout the talk.9:38–11:54 · Matt pushing back 0/10 Precision vs. Recall and Ethical Data Science Monologue presentation on precision versus recall, ethical data science, and human-in-the-loop workflows. Host activity is zero.11:54–15:36 · Matt pushing back 1/10 Presentation Wrap-Up and Transition to Q&A Matt Turck steps in, citing his four-year background at Bloomberg and asking about human-in-the-loop mechanics and data pipelines. Stent gently reframes his question about a single data pipeline, clarifying that Bloomberg operates between 15,000 and 40,000 distinct functions.15:36–20:29 · Matt pushing back 0/10 Audience Q&A: Social Media Manipulation, Speed Limits, and Market Truth Audience members ask detailed questions on bot manipulation, speed calculus, and market expectations. The host acts only as a moderator, while Stent provides detailed technical and philosophical answers.20:29–21:28 · Matt pushing back 0/10 Audience Q&A: Competitor Usage of Automated Content Audience questions cover competitor usage of automated text and sentiment extraction from earnings calls. Stent shares informative responses while the host maintains a minimal moderating role.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0% · guest 100%0:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 27.3% · guest 72.7%12:00 · Matt 27.3% · guest 72.7%15:00 · Matt 1.8% · guest 98.2%15:00 · Matt 1.8% · guest 98.2%18:00 · Matt 0.7% · guest 99.3%18:00 · Matt 0.7% · guest 99.3%21:00 · Matt 2% · guest 98%21:00 · Matt 2% · guest 98%
Sharpest disagreement ▶ 14:09 Clarifying Bloomberg's infrastructure architecture

Stent politely rejects the host's premise of a single data pipeline, explaining that Bloomberg runs 15,000 to 40,000 separate functional pipelines.

Hardest push from Matt ▶ 12:52 Host pushes on practical human-in-the-loop execution

Matt Turck anchors his question in his own four-year experience at Bloomberg and presses for practical details on how human-machine handoffs actually function.

Biggest teaching moment ▶ 14:09 Educating on multi-function pipeline complexity

Stent corrects the host's impression of backend architecture by breaking down how thousands of feeds feed tens of thousands of individual functions.

Matt holds his own ▶ 12:26 Host establishes domain authority and Bloomberg history

Matt Turck highlights Bloomberg's foundational tech heritage and mentions his four-year tenure at the company to frame his line of questioning.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Four Key NLP Applications at Bloomberg 0000 Monologue presentation segment delivered by Amanda Stent introducing Bloomberg's core mission and NLP enrichment applications. The host is not active during this segment, requiring host-side scores to be zero.
Three Core Challenges in Financial Text Analytics 0000 Monologue presentation detailing idiomatic financial language and document understanding challenges. The host is inactive throughout this portion.
The Need for Speed and High-Volume Social Processing 0000 Monologue segment focusing on latency, processing speed, and tweet funnels. The host remains silent throughout the talk.
Precision vs. Recall and Ethical Data Science 0000 Monologue presentation on precision versus recall, ethical data science, and human-in-the-loop workflows. Host activity is zero.
Presentation Wrap-Up and Transition to Q&A 4321 Matt Turck steps in, citing his four-year background at Bloomberg and asking about human-in-the-loop mechanics and data pipelines. Stent gently reframes his question about a single data pipeline, clarifying that Bloomberg operates between 15,000 and 40,000 distinct functions.
Audience Q&A: Social Media Manipulation, Speed Limits, and Market Truth 1210 Audience members ask detailed questions on bot manipulation, speed calculus, and market expectations. The host acts only as a moderator, while Stent provides detailed technical and philosophical answers.
Audience Q&A: Competitor Usage of Automated Content 1210 Audience questions cover competitor usage of automated text and sentiment extraction from earnings calls. Stent shares informative responses while the host maintains a minimal moderating role.

Statements from this episode (14)

Assertion Supported
Stent: One-minute news reporting delay corresponded to 3% stock price move
“The distance between 10 38 and 10 39 Was about three percent in the stock price. The distance between the SEC headline and 1055, 20 minutes later, was a 12% absolute difference in the stock price.”
Amanda Stent Mar 2, 2018 ▶ 3:13
Assertion Supported
Stent: SEC declared Twitter valid for official corporate disclosures
“SEC has declared that Twitter, in particular, can be used to provide company, official company announcements.”
Amanda Stent Mar 2, 2018 ▶ 3:38
Disclosure
Stent: Bloomberg NLP aims to inform trading decisions, not track sentiment
“At Bloomberg, we're not trying to identify net promoter score. We're trying to really understand what's going on so that someone else can decide whether they should buy or sell this company.”
Amanda Stent Mar 2, 2018 ▶ 5:31
Insight
Stent: Vertical NLP faces key hurdles in domain language, speed, and precision
“Here are three key challenges that we face that would be present to a greater or lesser extent in any other highly highly vertical industry. The first is financial language, idiomatic language. The second is the need for speed. Remember, I said a minute. And t…”
Amanda Stent Mar 2, 2018 ▶ 5:41
Assertion Supported
Stent: No generic text analytics platform can handle financial language
“There is no on the market generic text analytics platform that can handle this type of language. No way.”
Amanda Stent Mar 2, 2018 ▶ 6:51
Disclosure
Bloomberg can automatically generate news articles without humans in two minutes
“And this kind of article we can now automatically generate. So you can get this with no human intervention within two minutes, and then within five minutes a human journalist can come along in art color.”
Amanda Stent Mar 2, 2018 ▶ 8:29
Assertion Not checkable as stated
Bloomberg NLP models feed over 100 Bloomberg Terminal screens
“And they feed over a hundred different screens within the Bloomberg terminal, which provide company trend information, company tweet information, sentiment information, sentiment alerts, tweets on mergers, tweets by the president of the United States, et ceter…”
Amanda Stent Mar 2, 2018 ▶ 9:17
Assertion Partly supported
Generic NLP mistakenly linked Bill O'Reilly news to O'Reilly Auto Parts
“Okay, so Bill O'Reilly said something, and there's a small and innocent company called O'Reilly Auto Parts somewhere in the Midwest, and a bunch of generic text analytics engines linked Bill O'Reilly to O'Reilly Auto Parts, and the O'Reilly Auto Parts stock dr…”
Amanda Stent Mar 2, 2018 ▶ 9:48
Disclosure
Stent: Bloomberg shifted data labeling to deep learning and decision trees
“And that's something that historically has been done mostly with rule-based systems, but today we do it with deep learning and a lot of decision trees.”
Amanda Stent Mar 2, 2018 ▶ 13:49
Disclosure
Stent: Bloomberg Terminal contains between 15,000 and 40,000 functions
“There are, nobody has told me the actual number. I've heard numbers from between 15 and 40,000 individual functions inside the Bloomberg. Each of them is fed by at least one data feed, and in many cases, several.”
Amanda Stent Mar 2, 2018 ▶ 14:12
Disclosure
Bloomberg filters social media manipulation using account metadata and supply chain data
“We do a lot of historical monitoring, so we can know how old a social media account is. We also have a very, very large collection of pieces of information about entities in the world, so we can know how closely associated a particular handle is. With a decisi…”
Amanda Stent Mar 2, 2018 ▶ 16:15
Opinion
Stent: Bloomberg's human-written stories are much more informative than automated ones
“Our manually generated stories, which are frankly much more informative”
Amanda Stent Mar 2, 2018 ▶ 21:12
Insight
Stent: Earnings call cues cannot reliably predict financial performance
“So we can't always predict based on the non-linguistic or even the linguistic features of a, of an earnings call or an earnings release.”
Amanda Stent Mar 2, 2018 ▶ 23:13
Assertion Contradicted
Stent: Bloomberg sentiment feed outputs probabilistic buy or sell signals
“So the Bloomberg sentiment feed is not just positive or negative. I like it right. It's buy or sell, and it's with probability so that you can use those in your own models.”
Amanda Stent Mar 2, 2018 ▶ 23:30
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.