The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Richard Craib no published score: no usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 14 produced feed exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
14exchanges match
0on raw tape
1redirected or not addressed
Answered produced feed D 5 · C 5 · P 5 · Cm 5 5.00

Q Why are these people engaging in this data science contest?

A Well, for sure, they find data science fun. They could have a job as a data scientist somewhere. We have someone who's, uh, who is, uh, working at NASA Jet Propulsion Lab, and then on the side, he would train a numeri model. And so there's a lot of different types of people in these fields where they have to pick up data science skills for their job, but then they want to practice. So that's a big motivator. And another big motivator is just competition, becoming a master. At something is very compelling. And finally, money. Numeri has created the highest paying data science tournament on the internet. We pay out more than Kaggle does, and we've paid out over thirty million dollars since we started. So it's 30 Netflix prizes worth of pet prizes. So that's definitely a piece. But we don't know. I mean, sometimes our users are anonymous and we don't even know why they're there or what their, what their motivations are.

AI assessment note: “Well, for sure, they find data science fun. They could have a job”

Answered produced feed D 5 · C 5 · P 5 · Cm 5 5.00

Q What goes into developing a new, say, good feature?

A Yeah. Well, the first part is it shouldn't be correlated with existing features. That's one thing, but there are many other things besides. One thing we don't like our features to do is move fast, like have high turnover themselves. If we stuck a seven day momentum feature into Numerai's dataset, the Numerai users would like it and their models would pick up on it. But coming to trade execution, we wouldn't want to trade that fast. And then we wouldn't actually make any money from that. So there's a sensitivity to churn. And then we also like to see features that don't have long periods of not working. Like they should have something to say about the variance of stocks all the time.

AI assessment note: “the first part is it shouldn't be correlated with existing features.”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q Great. Well, let's walk through what this is. And so as a strategy, Why don't we begin breaking this down with who are these people that are getting involved in creating, you know, what becomes the inputs for the model?

A The data scientists are typically some of the best data scientists in the world, frankly. I mean, one of the nice things about starting Numeri was that Kaggle already existed and already had Kaggle grandmasters, people who'd won multiple competitions in different fields. And so they, Kaggle had kind of networked the data science community already. And so when Numeri started, they all talked to each other about it. They all joined and I would see people on Numeri's leaderboard and be like, oh, I recognize that guy's name. He's a, he's a Kaggle user that I competed against or something. So, but that's the thing. Data science is not quant. None of these people were quants. None of them had any background in quant. They're just the types of people who can solve abstract problems to do with data. And that's something people miss about Numeri. Like, People are not coming to Numeri and working on making a trading strategy or something like that. They are modeling data. Just in a similar way, the people at Google Translate, who are working on Google Translate, do not know all the languages. They just model the data, and that's what, what's happening at Numeri, and just like Google Translate, they can be a lot better than a professional Translator on certain languages, and Numeri can be better than a professional investor.

AI assessment note: “Kaggle grandmasters, people who'd won multiple competitions in different fields”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q So you mentioned paying out thirty million dollars of compensation. What's the mechanism and the process that you pay that out?

A Yeah, so we, we actually used to pay people out in dollars, but we would find that we would have data scientists around the world that said, oh, we don't have PayPal. Well, you can't send us dollars. Can you please just send us Bitcoin? And that's how we got into paying crypto in the first place. It was just a basic need to pay our data scientists. But since then, we started moving to paying in our own cryptocurrency, which we made in 2017. And so that thirty million dollars, a lot of it's coming from our cryptocurrency, uh, almost all of it, but, uh, it's quite liquid. In fact, our cryptocurrency is more liquid than some stocks, and so people will earn some of it, maybe they'll earn 10,000 dollars by staking on Numurai, and then they'll go to Coinbase and sell it into dollars, or they'll keep staking it on their model.

AI assessment note: “we started moving to paying in our own cryptocurrency... by staking on Numurai”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q What are the frequency of the predictions that come in for you to build into the model and portfolio?

A We get the predictions sent to us now every day, but we only trade quite slowly. So we'll only turn over the portfolio quite slowly, minimizing our market impact. And we end up holding positions for three or four months. And that's what people find quite unusual. Like they look at Numurai, like this is a quant fund. Oh, so you're going to be sub one week holding periods. But Numurai, we somehow doing things on three to four months and we still very, very high sharp. So I think that's a nice part about it. Like, even though we're playing in this competitive space, we've kind of done things a little bit differently. And that's what's paying off because I don't think people are working in the same region of space as we are.

AI assessment note: “We get the predictions sent to us now every day”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q At some point, you can imagine that the models that come in could overlap with each other, have high correlation with each other. How do you integrate that into your process?

A That's a really good point. So one of the reasons Millennium works is because they can get strategies that are uncorrelated, and that's what boosts the sharp. They have so many different strategies that are all uncorrelated, it's starts to become very hard to have a down year. Numurai, it's the exact same. We figured out, maybe about a year ago, how, you know, we were seeing people who'd say submit a model with 99% correlation to some basic model that we'd released for free. So it's like pretty clear that that data scientist didn't really even try to be different to the example script, but we started a new way of paying our users aside from just their IC, their correlation with subsequent returns. We said, we're going to pay you for being uncorrelated effectively. So if you make a mediocre model that's very uncorrelated, you can do particularly well. Because of its lack of correlation is actually helping Numeri more than a good model that's correlated with the one we already have. And so that's a key piece of what's going on right now is people are working on strategies that are, wow, I never thought of that or I never thought of doing that. And they end up being like 40, 30% correlated with the average model, which is really good.

AI assessment note: “we started a new way of paying our users... for being uncorrelated effectively.”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q Since you launched six years ago, what are some of the adoptions that you made in how this approach has worked?

A The biggest one was cracking staking. I mean, we had a period where, yes, we were getting a lot of people submitting models, but we couldn't trust them. We couldn't trust that they would keep working out of sample. We didn't even know who the people were, but having staking was huge because suddenly, oh wow, I'm going to stop submitting my bad models and only submit my best model because that's the one I'm putting money on. And that just cleared up everything. And made Numeri start working. Before then, we really weren't working. Like, first couple of years, things were not working. We were growing our user base, but the rest of it was, had a long way to go. The other piece is optimization. Being factor neutral, but also neutralizing all kinds of other risks, I think has been very important for Numeri. We have a very, very good drawdown characteristics, like the numerized fund hasn't historically gone down very much in periods of market stress, and that's what people want from a market neutral quant fund.

AI assessment note: “The biggest one was cracking staking... The other piece is optimization.”

Answered produced feed D 5 · C 5 · P 5 · Cm 4 4.85

Q So you've done this first run in your career at building these models. What was the path from that to starting to think about Numeri?

A It was quite early on in this job that I started competing in data science competitions myself. There's a data science competition website called Kaggle. And they listed all these problems, and people could go and try to solve the problems with machine learning. And so for anyone interested in machine learning at that time, this was definitely the place to go to try to test your skills on new data sets. What was really remarkable was that every single time a company put up a data set, and they said, try to solve this problem. Like famously, the Netflix prize, they said, try to solve this movie recommendation problem. And as whenever it was done, the data scientists who approached the problem demolished the benchmark. I mean, the benchmark was at some level, and the data scientists got way better. And that's an interesting thing that, like, even a great company like Netflix could have an internal model. And have people on the outside quite quickly do a much better job. And that's, uh, what got me thinking about Numeri. Could you have a hedge fund that worked this way? Could you have a hedge fund where the data of the hedge fund was completely open and free? And then anybody could join and do data science on the data. And then you'd have the best hedge fund potentially, because you would have the most talent being applied to your problem. You'd have more than any other hedge fund…

AI assessment note: “what got me thinking about Numeri. Could you have a hedge fund that worked this way?”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q How have you thought about the ability of this system to continue to scale and deliver with larger asset size?

A Yeah, that is the trick with quant funds. I mean, most of them just stopped working very quickly because they focused on a small asset amount. Numariah, we were just super strict on this, and this was from my experience of working at a fund where I was trying to build a high capacity strategy. I basically told the team, like, we are never allowed to do a back test of fifty million dollars, because even though we only had 10 at the time, we had to assume we're doing way more. And that discipline forces you into these higher spaces of capacity. The horizon is very, very long. And so that helps. But the other piece is also that more and more models helps too. You know, we can always grow the data set size. We can always grow our data science community and machine learning as a whole right now is taking off like crazy. Like the number of new research papers being released that apply to new Mariah. Is very large. So I think capacity is kind of like in our DNA. We never wanted to build a small hedge fund. So we've never touched the small opportunities. We've always gone for the larger stocks, larger horizon.

AI assessment note: “we had to assume we're doing way more. And that discipline forces you”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q Why don't you take me back to your path to this particularly interesting approach that you're taking to investing?

A Well, to go way back, uh, when I was eight years old, my dad, who worked as a portfolio manager, he gave me my first stocks, and I, I would follow them in the newspaper, and I remember rushing to my grandmother's room to try to see how the stocks were doing, and I was just fascinated by it. I was fascinated by the fact that these kind of numbers represented real companies in the world doing real things, and I was fascinated by Things like, well, why can a stock move 20% in a single day? Why isn't it two percent? That's a lot instead, you know, why is it 20? And then also got fascinated with the idea that it's very hard to beat the market portfolio and that nearly all investors lose to it over a long period. And then, you know, when I was a teenager, I started trading stocks on the U.S. markets. Our high school actually ended at 3:15 p.m. South African time. Which was 15 minutes before US Open. So I could run to the library and trade stocks from the library computers in Cape Town. And this is just what I was interested in. You could, if you looked at my library card, you would see that nearly every book I took out was about business or finance. So it's just been a lifelong thing for me.

AI assessment note: “when I was eight years old, my dad, who worked as a portfolio manager”

Answered produced feed D 4 · C 5 · P 4 · Cm 4 4.30

Q How do you decide how much leverage to run on the strategy?

A It's a good question. That is in some ways the magic of quant. Jim Simons would make absolutely no money if it weren't for leverage. That's just how, how it works. But you want to leverage things that are safe. So what we try to do is make sure that even on very high leverage, our volatility is way lower than the, say the S&P. When you're building these portfolios and you're so hedged to so many things, It's very hard for things to affect you. For example, you know, in January, momentum, the factor crashed, and a lot of quants did badly. Numeri is neutral to momentum, so we basically didn't even notice this crash because we're hedged to that risk, and we do that with as many things as possible to make our track record more and more idiosyncratic and focused on alpha.

AI assessment note: “make sure that even on very high leverage, our volatility is way lower”

Answered produced feed D 5 · C 4 · P 4 · Cm 3 4.15

Q How have you gone about engaging with this community of data scientists?

A We have a discord channel. We have, uh, what we call fireside chats, which is once a quarter where we have a chat with all of our community and answer the tough questions that they have. And then we have conferences. We had a big conference in San Francisco called Numicon. Where we had probably the largest gathering of Kaggle grandmasters in real life ever. The number of bright, amazing people who came was staggering. And so that's the type of thing we do, but it's strange. I mean, it's a bit like the, who, who's a hedge fund customer, right? Or LPs. But then in some ways you want to be like, well, wait a second. So are our data scientists because they're kind of a customer of our websites experience and data systems. But then they also kind of work for us, because they're helping to build the strategy. That's why I think it's quite a cool company, because it's like a modern thing. Like, how is it that you have a quant hedge fund that doesn't build any of their own models, and has people that they don't even know building the models, and they are getting paid in cryptocurrency? There's a lot of interesting things to think about.

AI assessment note: “We have a discord channel. We have, uh, what we call fireside chats”

Partly produced feed D 3 · C 4 · P 3 · Cm 3 3.30

Q So in that role, you're building a machine learning model from your elementary roots in self-taught finance. The inputs can come both from the fundamental side of the companies and then trading. I'm kind of curious, what were you putting into the machine and what was it trying to learn?

A Well, that's a really good point. So there were many funds that were trading with very short-term signals, a fund like Renaissance would, would classify, which I looked up to. But what was interesting was if you do short-term trading, your capacity is very low. So if you trade very frequently in very small stocks, you're moving the market a lot, and it's very hard to have a strategy that scales. And working with my boss, he would say, well, I could see that working, but only on fifty million dollars. I want you to make something that works on a billion dollars because the firm had fifteen billion dollars and they wanted to put money to work. So he kept forcing me to say, no, make it work on a longer horizon. Make it work with fundamental data, not technical data. And so I started to develop something quite strange about, there's a, there's a very different approach to tackling that problem. Versus tackling the short-term trading problem. But I think it's more valuable because it's higher capacity.

AI assessment note: “Make it work with fundamental data, not technical data.”

Not addressed produced feed D 1 · C 2 · P 2 · Cm 3 1.85

Q And how do you drive that forward in the incentive system for this group of developers you may or may not know?

A Well, for that project, it might involve some internal work, uh, with our own internal data scientists, because we would be the ones building the feature to add to our data set as a new feature, but absolutely the case that other data scientists could do similar things and probably are already. So it's a very interesting time for machine learning, and I think people are going to be very surprised Like the difference between having an AI hedge fund, a good one in your portfolio versus not. I know it sounds self-serving, but the difference will be very stark in the long run because there is going to be basically a wave, a third generation of hedge funds that do things differently to others that I, I'm quite excited to see.

AI assessment note: “it might involve some internal work, uh, with our own internal data scientists”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 700 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.