May 24, 2017 · 23m · mad

How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)

Parth Vasa · 15m spoken Bloomberg BI Employee · 49s spoken Matt Turck · 19s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Data Driven NYC presentation, Parth Vasa, Head of Data Science at Bloomberg Engineering, explains how Bloomberg leverages open-source Apache Solr and advanced machine learning models like LambdaMART to build high-precision federated search systems across complex financial datasets.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1.6% of the talking time here. How this is scored →

Matt as informed peer 0.8 Guest teaching 5.5 Guest disagreement 0.5 Matt pushing back 0.1
05100:0010:0020:000:33–2:55 · Matt as informed peer 0/10 Overview of Bloomberg Data Diversity Parth gives a presentation monologue introducing Bloomberg's diverse financial market data types and why high-precision search is critical when real money is at stake. The host does not speak during this segment, so host scores are set to zero.2:55–6:01 · Matt as informed peer 0/10 Federated Search on the Bloomberg Terminal Parth details unified natural language search across 60 data sets and Bloomberg's contributions to open-source Solr. Host remains silent during the monologue.6:01–8:07 · Matt as informed peer 0/10 Learning to Rank and Problem Formulation (NDCG) Parth educates the audience on machine learning search metrics, specifically defining Normalized Discounted Cumulative Gain (NDCG) and its mathematical rationale. Host does not participate.8:07–10:15 · Matt as informed peer 0/10 LambdaMART and Tree-Based Ranking Models Parth explains gradient boosted regression trees (LambdaMART) and how tree-based ranking allows complex feature interactions like age versus asset type. Host is non-talker in this monologue segment.10:15–13:02 · Matt as informed peer 0/10 Integrating ML-Based Reranking directly into Solr Parth covers pushing ranking models directly into Solr via ZooKeeper and contrasts A/B testing with interleaving for small, specialized user bases. Monologue segment requires zero host scores.13:02–16:09 · Matt as informed peer 4/10 Audience Q&A - Solr vs. Elasticsearch and Model Maintenance Host Matt Turck opens Q&A by asking a knowledgeable question comparing Solr to Elasticsearch. Parth explains open-source governance differences using a Beatles vs. Stones analogy, followed by an audience question on model maintenance.16:09–20:12 · Matt as informed peer 1/10 Audience Q&A - Low Volume Query Strategy and Search Entry Points Audience members ask about low daily query volume and question if Bloomberg is overcomplicating search when users know specific functions. Parth constructively defends unified natural language search as a cross-asset discoverability layer.20:12–23:09 · Matt as informed peer 1/10 Audience Q&A - Exogenous Events, Unsupervised Learning, and Conclusion Parth answers questions regarding market volatility spikes using online learning algorithms and explains why unsupervised learning cannot be directly relied upon due to strict explainability and precision requirements. Matt Turck concludes the session.0:33–2:55 · Guest teaching 5/10 Overview of Bloomberg Data Diversity Parth gives a presentation monologue introducing Bloomberg's diverse financial market data types and why high-precision search is critical when real money is at stake. The host does not speak during this segment, so host scores are set to zero.2:55–6:01 · Guest teaching 5/10 Federated Search on the Bloomberg Terminal Parth details unified natural language search across 60 data sets and Bloomberg's contributions to open-source Solr. Host remains silent during the monologue.6:01–8:07 · Guest teaching 6/10 Learning to Rank and Problem Formulation (NDCG) Parth educates the audience on machine learning search metrics, specifically defining Normalized Discounted Cumulative Gain (NDCG) and its mathematical rationale. Host does not participate.8:07–10:15 · Guest teaching 6/10 LambdaMART and Tree-Based Ranking Models Parth explains gradient boosted regression trees (LambdaMART) and how tree-based ranking allows complex feature interactions like age versus asset type. Host is non-talker in this monologue segment.10:15–13:02 · Guest teaching 6/10 Integrating ML-Based Reranking directly into Solr Parth covers pushing ranking models directly into Solr via ZooKeeper and contrasts A/B testing with interleaving for small, specialized user bases. Monologue segment requires zero host scores.13:02–16:09 · Guest teaching 5/10 Audience Q&A - Solr vs. Elasticsearch and Model Maintenance Host Matt Turck opens Q&A by asking a knowledgeable question comparing Solr to Elasticsearch. Parth explains open-source governance differences using a Beatles vs. Stones analogy, followed by an audience question on model maintenance.16:09–20:12 · Guest teaching 5/10 Audience Q&A - Low Volume Query Strategy and Search Entry Points Audience members ask about low daily query volume and question if Bloomberg is overcomplicating search when users know specific functions. Parth constructively defends unified natural language search as a cross-asset discoverability layer.20:12–23:09 · Guest teaching 6/10 Audience Q&A - Exogenous Events, Unsupervised Learning, and Conclusion Parth answers questions regarding market volatility spikes using online learning algorithms and explains why unsupervised learning cannot be directly relied upon due to strict explainability and precision requirements. Matt Turck concludes the session.0:33–2:55 · Guest disagreement 1/10 Overview of Bloomberg Data Diversity Parth gives a presentation monologue introducing Bloomberg's diverse financial market data types and why high-precision search is critical when real money is at stake. The host does not speak during this segment, so host scores are set to zero.2:55–6:01 · Guest disagreement 0/10 Federated Search on the Bloomberg Terminal Parth details unified natural language search across 60 data sets and Bloomberg's contributions to open-source Solr. Host remains silent during the monologue.6:01–8:07 · Guest disagreement 0/10 Learning to Rank and Problem Formulation (NDCG) Parth educates the audience on machine learning search metrics, specifically defining Normalized Discounted Cumulative Gain (NDCG) and its mathematical rationale. Host does not participate.8:07–10:15 · Guest disagreement 0/10 LambdaMART and Tree-Based Ranking Models Parth explains gradient boosted regression trees (LambdaMART) and how tree-based ranking allows complex feature interactions like age versus asset type. Host is non-talker in this monologue segment.10:15–13:02 · Guest disagreement 0/10 Integrating ML-Based Reranking directly into Solr Parth covers pushing ranking models directly into Solr via ZooKeeper and contrasts A/B testing with interleaving for small, specialized user bases. Monologue segment requires zero host scores.13:02–16:09 · Guest disagreement 1/10 Audience Q&A - Solr vs. Elasticsearch and Model Maintenance Host Matt Turck opens Q&A by asking a knowledgeable question comparing Solr to Elasticsearch. Parth explains open-source governance differences using a Beatles vs. Stones analogy, followed by an audience question on model maintenance.16:09–20:12 · Guest disagreement 2/10 Audience Q&A - Low Volume Query Strategy and Search Entry Points Audience members ask about low daily query volume and question if Bloomberg is overcomplicating search when users know specific functions. Parth constructively defends unified natural language search as a cross-asset discoverability layer.20:12–23:09 · Guest disagreement 0/10 Audience Q&A - Exogenous Events, Unsupervised Learning, and Conclusion Parth answers questions regarding market volatility spikes using online learning algorithms and explains why unsupervised learning cannot be directly relied upon due to strict explainability and precision requirements. Matt Turck concludes the session.0:33–2:55 · Matt pushing back 0/10 Overview of Bloomberg Data Diversity Parth gives a presentation monologue introducing Bloomberg's diverse financial market data types and why high-precision search is critical when real money is at stake. The host does not speak during this segment, so host scores are set to zero.2:55–6:01 · Matt pushing back 0/10 Federated Search on the Bloomberg Terminal Parth details unified natural language search across 60 data sets and Bloomberg's contributions to open-source Solr. Host remains silent during the monologue.6:01–8:07 · Matt pushing back 0/10 Learning to Rank and Problem Formulation (NDCG) Parth educates the audience on machine learning search metrics, specifically defining Normalized Discounted Cumulative Gain (NDCG) and its mathematical rationale. Host does not participate.8:07–10:15 · Matt pushing back 0/10 LambdaMART and Tree-Based Ranking Models Parth explains gradient boosted regression trees (LambdaMART) and how tree-based ranking allows complex feature interactions like age versus asset type. Host is non-talker in this monologue segment.10:15–13:02 · Matt pushing back 0/10 Integrating ML-Based Reranking directly into Solr Parth covers pushing ranking models directly into Solr via ZooKeeper and contrasts A/B testing with interleaving for small, specialized user bases. Monologue segment requires zero host scores.13:02–16:09 · Matt pushing back 1/10 Audience Q&A - Solr vs. Elasticsearch and Model Maintenance Host Matt Turck opens Q&A by asking a knowledgeable question comparing Solr to Elasticsearch. Parth explains open-source governance differences using a Beatles vs. Stones analogy, followed by an audience question on model maintenance.16:09–20:12 · Matt pushing back 0/10 Audience Q&A - Low Volume Query Strategy and Search Entry Points Audience members ask about low daily query volume and question if Bloomberg is overcomplicating search when users know specific functions. Parth constructively defends unified natural language search as a cross-asset discoverability layer.20:12–23:09 · Matt pushing back 0/10 Audience Q&A - Exogenous Events, Unsupervised Learning, and Conclusion Parth answers questions regarding market volatility spikes using online learning algorithms and explains why unsupervised learning cannot be directly relied upon due to strict explainability and precision requirements. Matt Turck concludes the session.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 0.2% · guest 99.8%0:00 · Matt 0.2% · guest 99.8%3:00 · Matt 0% · guest 100%3:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%6:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 8.4% · guest 91.6%12:00 · Matt 8.4% · guest 91.6%15:00 · Matt 1.5% · guest 98.5%15:00 · Matt 1.5% · guest 98.5%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 4.3% · guest 95.7%21:00 · Matt 4.3% · guest 95.7%
Sharpest disagreement ▶ 18:16 Defending product necessity against audience challenge

An audience member challenges Parth on whether Bloomberg is making search more complex than necessary, prompting Parth to firmly defend the need for cross-asset discoverability over manual function navigation.

Hardest push from Matt ▶ 13:05 Host questions architectural stack choice

Matt Turck directly probes the core technical architecture decision, asking why Bloomberg chose Solr over the widely popular Elasticsearch alternative.

Biggest teaching moment ▶ 11:50 Why A/B testing fails for expert niche domains

Parth educates the audience on why conventional consumer A/B testing assumptions fall apart on small, expert user populations and why interleaving ranking evaluation is required.

Matt holds his own ▶ 13:05 Host demonstrates enterprise search stack familiarity

Matt Turck shows solid domain awareness by pinpointing Solr versus Elasticsearch as the primary technology decision for enterprise search pipelines.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Overview of Bloomberg Data Diversity 0510 Parth gives a presentation monologue introducing Bloomberg's diverse financial market data types and why high-precision search is critical when real money is at stake. The host does not speak during this segment, so host scores are set to zero.
Federated Search on the Bloomberg Terminal 0500 Parth details unified natural language search across 60 data sets and Bloomberg's contributions to open-source Solr. Host remains silent during the monologue.
Learning to Rank and Problem Formulation (NDCG) 0600 Parth educates the audience on machine learning search metrics, specifically defining Normalized Discounted Cumulative Gain (NDCG) and its mathematical rationale. Host does not participate.
LambdaMART and Tree-Based Ranking Models 0600 Parth explains gradient boosted regression trees (LambdaMART) and how tree-based ranking allows complex feature interactions like age versus asset type. Host is non-talker in this monologue segment.
Integrating ML-Based Reranking directly into Solr 0600 Parth covers pushing ranking models directly into Solr via ZooKeeper and contrasts A/B testing with interleaving for small, specialized user bases. Monologue segment requires zero host scores.
Audience Q&A - Solr vs. Elasticsearch and Model Maintenance 4511 Host Matt Turck opens Q&A by asking a knowledgeable question comparing Solr to Elasticsearch. Parth explains open-source governance differences using a Beatles vs. Stones analogy, followed by an audience question on model maintenance.
Audience Q&A - Low Volume Query Strategy and Search Entry Points 1520 Audience members ask about low daily query volume and question if Bloomberg is overcomplicating search when users know specific functions. Parth constructively defends unified natural language search as a cross-asset discoverability layer.
Audience Q&A - Exogenous Events, Unsupervised Learning, and Conclusion 1600 Parth answers questions regarding market volatility spikes using online learning algorithms and explains why unsupervised learning cannot be directly relied upon due to strict explainability and precision requirements. Matt Turck concludes the session.

Statements from this episode (10)

Assertion Not checkable as stated
Bloomberg cannot rely on large-scale A/B testing or usage data
“So a lot of the luxury that other companies have, like doing large scale machine learning based on massive amount of usage data, or doing large scale A-B testing to decide which model works better, doesn't work for us.”
Parth Vasa May 24, 2017 ▶ 2:12
Disclosure
Bloomberg plans to consolidate terminal search into a single ranked list
“And we want to get to a place where you can search all that in one screen, and we will give you that information in one single ranked list, so you don't have to learn more and more different places to search for data.”
Parth Vasa May 24, 2017 ▶ 3:53
Assertion Supported
Bloomberg contributed code to the last fifteen Apache Solr releases
“I think about the last 15, 15 or 16 versions of Solr have had some sort of our code in it”
Parth Vasa May 24, 2017 ▶ 4:55
Disclosure
Bloomberg uses LambdaMART decision tree algorithms for terminal search ranking
“The one we use is called Lambda Mart, which is based on decision trees, or gradient, regression trees, actually.”
Parth Vasa May 24, 2017 ▶ 8:09
Disclosure
Learning-to-rank algorithms continuously improve search models without manual human tweaking
“We had a huge improvement once we started using it, but more than that improvement, the good part about learning to rank is it always continues to improve your model without constant need of people tweaking it.”
Parth Vasa May 24, 2017 ▶ 9:59
Insight
Embedding machine learning reranking inside Apache Solr improves search performance
“And also, it reduces the hop between two services, so your performance gets much better.”
Parth Vasa May 24, 2017 ▶ 10:59
Disclosure
Bloomberg uses interleaving instead of A/B testing to deploy search models
“So that's what we use for, ah, we use to decide which model to push out.”
Parth Vasa May 24, 2017 ▶ 12:36
Disclosure
Bloomberg chose Apache Solr over Elasticsearch to maintain open-source control
“One was, we were looking for something that was truly open source, where we had a lot more control over open source and, sorry, source code, and a lot more saying how the direction goes.”
Parth Vasa May 24, 2017 ▶ 13:16
Assertion Not checkable as stated
Bloomberg Terminal search volume reaches the low hundreds of thousands daily
“I think on a good day we are talking about low 100,000.”
Parth Vasa May 24, 2017 ▶ 16:43
Disclosure
Bloomberg limits unsupervised learning to data bootstrapping due to precision needs
“We use unsupervised learning a lot to Sort of bootstrap our data. For example, word to whack, right? It's a perfect unsupervised learning algorithm. We use that a lot to, for query reformulation and all that, but, ah, it's a little risky for something as high …”
Parth Vasa May 24, 2017 ▶ 22:10
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.