The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 10 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Disclosure
Bloomberg chose Apache Solr over Elasticsearch to maintain open-source control
“One was, we were looking for something that was truly open source, where we had a lot more control over open source and, sorry, source code, and a lot more saying how the direction goes.”
Parth Vasa May 24, 2017 ▶ 13:16 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Assertion Not checkable as stated
Bloomberg cannot rely on large-scale A/B testing or usage data
“So a lot of the luxury that other companies have, like doing large scale machine learning based on massive amount of usage data, or doing large scale A-B testing to decide which model works better, doesn't work for us.”
Parth Vasa May 24, 2017 ▶ 2:12 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Disclosure
Learning-to-rank algorithms continuously improve search models without manual human tweaking
“We had a huge improvement once we started using it, but more than that improvement, the good part about learning to rank is it always continues to improve your model without constant need of people tweaking it.”
Parth Vasa May 24, 2017 ▶ 9:59 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Insight
Embedding machine learning reranking inside Apache Solr improves search performance
“And also, it reduces the hop between two services, so your performance gets much better.”
Parth Vasa May 24, 2017 ▶ 10:59 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Disclosure
Bloomberg limits unsupervised learning to data bootstrapping due to precision needs
“We use unsupervised learning a lot to Sort of bootstrap our data. For example, word to whack, right? It's a perfect unsupervised learning algorithm. We use that a lot to, for query reformulation and all that, but, ah, it's a little risky for something as high …”
Parth Vasa May 24, 2017 ▶ 22:10 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Disclosure
Bloomberg plans to consolidate terminal search into a single ranked list
“And we want to get to a place where you can search all that in one screen, and we will give you that information in one single ranked list, so you don't have to learn more and more different places to search for data.”
Parth Vasa May 24, 2017 ▶ 3:53 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Assertion Supported
Bloomberg contributed code to the last fifteen Apache Solr releases
“I think about the last 15, 15 or 16 versions of Solr have had some sort of our code in it”
Parth Vasa May 24, 2017 ▶ 4:55 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Disclosure
Bloomberg uses LambdaMART decision tree algorithms for terminal search ranking
“The one we use is called Lambda Mart, which is based on decision trees, or gradient, regression trees, actually.”
Parth Vasa May 24, 2017 ▶ 8:09 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Disclosure
Bloomberg uses interleaving instead of A/B testing to deploy search models
“So that's what we use for, ah, we use to decide which model to push out.”
Parth Vasa May 24, 2017 ▶ 12:36 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Assertion Not checkable as stated
Bloomberg Terminal search volume reaches the low hundreds of thousands daily
“I think on a good day we are talking about low 100,000.”
Parth Vasa May 24, 2017 ▶ 16:43 How to Train Your Search Engine // Parth Vasa, Bloomberg (FirstMark's Data Driven)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.