May 24, 2017 · 23m · mad
How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this Data Driven NYC presentation, Parth Vasa, Head of Data Science at Bloomberg Engineering, explains how Bloomberg leverages open-source Apache Solr and advanced machine learning models like LambdaMART to build high-precision federated search systems across complex financial datasets.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 1.6% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
An audience member challenges Parth on whether Bloomberg is making search more complex than necessary, prompting Parth to firmly defend the need for cross-asset discoverability over manual function navigation.
Hardest push from Matt ▶ 13:05 Host questions architectural stack choiceMatt Turck directly probes the core technical architecture decision, asking why Bloomberg chose Solr over the widely popular Elasticsearch alternative.
Biggest teaching moment ▶ 11:50 Why A/B testing fails for expert niche domainsParth educates the audience on why conventional consumer A/B testing assumptions fall apart on small, expert user populations and why interleaving ranking evaluation is required.
Matt holds his own ▶ 13:05 Host demonstrates enterprise search stack familiarityMatt Turck shows solid domain awareness by pinpointing Solr versus Elasticsearch as the primary technology decision for enterprise search pipelines.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Overview of Bloomberg Data Diversity | 0 | 5 | 1 | 0 | Parth gives a presentation monologue introducing Bloomberg's diverse financial market data types and why high-precision search is critical when real money is at stake. The host does not speak during this segment, so host scores are set to zero. | |
| Federated Search on the Bloomberg Terminal | 0 | 5 | 0 | 0 | Parth details unified natural language search across 60 data sets and Bloomberg's contributions to open-source Solr. Host remains silent during the monologue. | |
| Learning to Rank and Problem Formulation (NDCG) | 0 | 6 | 0 | 0 | Parth educates the audience on machine learning search metrics, specifically defining Normalized Discounted Cumulative Gain (NDCG) and its mathematical rationale. Host does not participate. | |
| LambdaMART and Tree-Based Ranking Models | 0 | 6 | 0 | 0 | Parth explains gradient boosted regression trees (LambdaMART) and how tree-based ranking allows complex feature interactions like age versus asset type. Host is non-talker in this monologue segment. | |
| Integrating ML-Based Reranking directly into Solr | 0 | 6 | 0 | 0 | Parth covers pushing ranking models directly into Solr via ZooKeeper and contrasts A/B testing with interleaving for small, specialized user bases. Monologue segment requires zero host scores. | |
| Audience Q&A - Solr vs. Elasticsearch and Model Maintenance | 4 | 5 | 1 | 1 | Host Matt Turck opens Q&A by asking a knowledgeable question comparing Solr to Elasticsearch. Parth explains open-source governance differences using a Beatles vs. Stones analogy, followed by an audience question on model maintenance. | |
| Audience Q&A - Low Volume Query Strategy and Search Entry Points | 1 | 5 | 2 | 0 | Audience members ask about low daily query volume and question if Bloomberg is overcomplicating search when users know specific functions. Parth constructively defends unified natural language search as a cross-asset discoverability layer. | |
| Audience Q&A - Exogenous Events, Unsupervised Learning, and Conclusion | 1 | 6 | 0 | 0 | Parth answers questions regarding market volatility spikes using online learning algorithms and explains why unsupervised learning cannot be directly relied upon due to strict explainability and precision requirements. Matt Turck concludes the session. |