Mayur Thakur, Managing Director at Goldman Sachs, discusses framing data architectures without relying on trendy machine learning terminology.
Disclosure
Goldman Sachs targets 95% recall and 20% precision for compliance surveillance
“And we generally hold ourselves to, you know, It's a very recall-driven, if people know what that means, but I'll describe quickly what it means. It's like, the, we want to catch a lot of things that we should be catching, like, 95% of that, right? The more th…”
Insight
Detecting market manipulation is essentially a database joining problem
“A problem about manipulating the market comes down to a problem about kind of joining databases, really, right?”
Assertion Not checkable as stated
Every single alert from Goldman Sachs' compliance surveillance is reviewed manually
“Every single alert that we produce is looked by some human.”
Assertion Not checkable as stated
Goldman Sachs' internal surveillance graph has 100M nodes and 1B edges
“Or in our case, slightly smaller than the Facebook graph, but still about a hundred million nodes and about a billion edges.”
Assertion Not checkable as stated
Goldman Sachs indexes up to 3 billion internal communications for surveillance
“Think of all the emails, all the chats, all the Bloomberg messages that come or go out of Goldman as being the, ah, the data behind it, right? So like two to three billion documents, ah, behind, billion documents behind it.”
Assertion Not checkable as stated
Goldman's internal surveillance search queries billions of documents in sub-second time
“All these things you can enter the query, and within sub-second, get results, ah, where it's actually going through, ah, well, billions of documents, and finding the relevant results.”