May 11, 2023 · 49m · capital-allocators

Richard Craib – Crowdsourcing Data Science for Returns at Numerai (Capital Allocators, EP.314)

Richard Craib · 35m spoken Ted Seides · 6m spoken
0:00 / 0:00

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Capital Allocators, host Ted Seides interviews Richard Craib, founder and CEO of Numerai, exploring how the firm crowdsources quantitative models from thousands of global data scientists using obfuscated data, machine learning, and cryptocurrency staking. Craib outlines the mechanics of synthesizing distributed forecasts into a scalable, factor-neutral hedge fund portfolio and shares his vision for the next generation of AI-driven asset management.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Ted holds 15.8% of the talking time here. How this is scored →

Ted as informed peer 2.6 Guest teaching 3.1 Guest disagreement 0.6 Ted pushing back 0.5
05100:0015:0030:0045:000:32–5:27 · Ted as informed peer 1/10 Legal Disclaimer and Advisory Notice Ted opens with a legal disclaimer and introductory biographical framing before Richard shares his early fascination with stock trading and math education. The dynamic is purely conversational and exploratory.5:29–7:54 · Ted as informed peer 3/10 Early Quantitative Experience and High-Capacity Machine Learning Ted asks how fundamental and technical data were integrated into Richard's early ML models. Richard explains the distinction between low-capacity short-term trading and high-capacity fundamental horizons.7:55–10:01 · Ted as informed peer 3/10 Theory-Free Machine Learning vs Traditional Hypothesis Testing When Ted asks about formulating investment hypotheses, Richard gently corrects him by explaining that machine learning inverts traditional finance by being explicitly theory-free and purely data-driven.10:02–12:36 · Ted as informed peer 2/10 The Genesis and Early Venture Backing of Numerai Richard explains how participating in Kaggle competitions inspired him to crowdsource quantitative modeling on open, obfuscated datasets, leading to backing from Howard Morgan.12:38–15:13 · Ted as informed peer 2/10 Crowdsourcing Global Data Science Talent Over Traditional Quants Richard highlights that data scientists modeling abstract obfuscated data outperform traditional quants, comparing it to Google Translate outperforming human linguists.15:14–17:41 · Ted as informed peer 2/10 The Numerai Dataset, Data Science Tournaments, and Payouts Richard details the structure of the million-row obfuscated dataset and explains why participants are motivated by both intellectual challenge and the substantial prize pool.17:42–20:31 · Ted as informed peer 3/10 The Staking Mechanism and Alignment via Cryptocurrency Richard details Numerai's staking mechanism where contributors put their own cryptocurrency at risk to ensure skin in the game, comparing the model to Millennium's multi-manager alignment.20:33–23:38 · Ted as informed peer 3/10 Model Turnover, Strategy Decay, and Crypto Market Resilience Ted probes whether crypto market volatility disrupted contributor incentives. Richard explains that user engagement doubled even when token prices dropped over 80%.23:39–25:50 · Ted as informed peer 3/10 Data Architecture and Feature Engineering Standards Richard outlines the engineering criteria for features, emphasizing low churn and multi-decade applicability across global equities rather than short-term momentum signals.25:51–28:01 · Ted as informed peer 3/10 From Daily Predictions to the Meta-Model and Factor Neutral Portfolio Richard explains the synthesis of user signals into the stake-weighted metamodel, which is subsequently passed through an optimizer to achieve country, sector, and factor neutrality.28:01–31:23 · Ted as informed peer 4/10 Portfolio Risk Hedging and Combating Machine Learning Overfitting Ted questions whether anonymous machine learning models simply overfit or data-mine noise. Richard pushes back by distinguishing modern ML generalization from naive data mining and explains how they incentivize uncorrelated models.31:24–33:55 · Ted as informed peer 3/10 Balance Sheet Construction, Leverage, and Idiosyncratic Alpha Richard breaks down portfolio leverage and risk constraints, highlighting how strict factor neutralization protects the strategy from systemic factor crashes like momentum shocks.33:57–38:06 · Ted as informed peer 4/10 Core Breakthroughs: Staking Mechanism and Robust Drawdown Control Ted asks how allocators overcome the black-box perception of Numerai. Richard explains that sophisticated quantitative allocators look past process opacity to statistically significant alpha attribution.38:08–42:44 · Ted as informed peer 3/10 Drawdown Alignment and Scalable Portfolio Architecture Richard discusses capacity discipline and drawdown mechanics, noting that contributors suffer token burns during model degradation, driving rapid corrective iteration.42:46–48:59 · Ted as informed peer 2/10 Future R&D: Large Language Models and the Next Generation of Hedge Funds Richard discusses the potential of LLMs for feature generation and answers Ted's concluding personal and career questions in a collaborative, reflective tone.49:01–49:23 · Ted as informed peer 0/10 Podcast Conclusion and Sponsored Episode Announcement Ted provides an outro monologue detailing how managers can sponsor insights on the Capital Allocators podcast.0:32–5:27 · Guest teaching 1/10 Legal Disclaimer and Advisory Notice Ted opens with a legal disclaimer and introductory biographical framing before Richard shares his early fascination with stock trading and math education. The dynamic is purely conversational and exploratory.5:29–7:54 · Guest teaching 3/10 Early Quantitative Experience and High-Capacity Machine Learning Ted asks how fundamental and technical data were integrated into Richard's early ML models. Richard explains the distinction between low-capacity short-term trading and high-capacity fundamental horizons.7:55–10:01 · Guest teaching 5/10 Theory-Free Machine Learning vs Traditional Hypothesis Testing When Ted asks about formulating investment hypotheses, Richard gently corrects him by explaining that machine learning inverts traditional finance by being explicitly theory-free and purely data-driven.10:02–12:36 · Guest teaching 3/10 The Genesis and Early Venture Backing of Numerai Richard explains how participating in Kaggle competitions inspired him to crowdsource quantitative modeling on open, obfuscated datasets, leading to backing from Howard Morgan.12:38–15:13 · Guest teaching 4/10 Crowdsourcing Global Data Science Talent Over Traditional Quants Richard highlights that data scientists modeling abstract obfuscated data outperform traditional quants, comparing it to Google Translate outperforming human linguists.15:14–17:41 · Guest teaching 3/10 The Numerai Dataset, Data Science Tournaments, and Payouts Richard details the structure of the million-row obfuscated dataset and explains why participants are motivated by both intellectual challenge and the substantial prize pool.17:42–20:31 · Guest teaching 4/10 The Staking Mechanism and Alignment via Cryptocurrency Richard details Numerai's staking mechanism where contributors put their own cryptocurrency at risk to ensure skin in the game, comparing the model to Millennium's multi-manager alignment.20:33–23:38 · Guest teaching 3/10 Model Turnover, Strategy Decay, and Crypto Market Resilience Ted probes whether crypto market volatility disrupted contributor incentives. Richard explains that user engagement doubled even when token prices dropped over 80%.23:39–25:50 · Guest teaching 4/10 Data Architecture and Feature Engineering Standards Richard outlines the engineering criteria for features, emphasizing low churn and multi-decade applicability across global equities rather than short-term momentum signals.25:51–28:01 · Guest teaching 4/10 From Daily Predictions to the Meta-Model and Factor Neutral Portfolio Richard explains the synthesis of user signals into the stake-weighted metamodel, which is subsequently passed through an optimizer to achieve country, sector, and factor neutrality.28:01–31:23 · Guest teaching 4/10 Portfolio Risk Hedging and Combating Machine Learning Overfitting Ted questions whether anonymous machine learning models simply overfit or data-mine noise. Richard pushes back by distinguishing modern ML generalization from naive data mining and explains how they incentivize uncorrelated models.31:24–33:55 · Guest teaching 3/10 Balance Sheet Construction, Leverage, and Idiosyncratic Alpha Richard breaks down portfolio leverage and risk constraints, highlighting how strict factor neutralization protects the strategy from systemic factor crashes like momentum shocks.33:57–38:06 · Guest teaching 4/10 Core Breakthroughs: Staking Mechanism and Robust Drawdown Control Ted asks how allocators overcome the black-box perception of Numerai. Richard explains that sophisticated quantitative allocators look past process opacity to statistically significant alpha attribution.38:08–42:44 · Guest teaching 3/10 Drawdown Alignment and Scalable Portfolio Architecture Richard discusses capacity discipline and drawdown mechanics, noting that contributors suffer token burns during model degradation, driving rapid corrective iteration.42:46–48:59 · Guest teaching 2/10 Future R&D: Large Language Models and the Next Generation of Hedge Funds Richard discusses the potential of LLMs for feature generation and answers Ted's concluding personal and career questions in a collaborative, reflective tone.49:01–49:23 · Guest teaching 0/10 Podcast Conclusion and Sponsored Episode Announcement Ted provides an outro monologue detailing how managers can sponsor insights on the Capital Allocators podcast.0:32–5:27 · Guest disagreement 0/10 Legal Disclaimer and Advisory Notice Ted opens with a legal disclaimer and introductory biographical framing before Richard shares his early fascination with stock trading and math education. The dynamic is purely conversational and exploratory.5:29–7:54 · Guest disagreement 1/10 Early Quantitative Experience and High-Capacity Machine Learning Ted asks how fundamental and technical data were integrated into Richard's early ML models. Richard explains the distinction between low-capacity short-term trading and high-capacity fundamental horizons.7:55–10:01 · Guest disagreement 2/10 Theory-Free Machine Learning vs Traditional Hypothesis Testing When Ted asks about formulating investment hypotheses, Richard gently corrects him by explaining that machine learning inverts traditional finance by being explicitly theory-free and purely data-driven.10:02–12:36 · Guest disagreement 0/10 The Genesis and Early Venture Backing of Numerai Richard explains how participating in Kaggle competitions inspired him to crowdsource quantitative modeling on open, obfuscated datasets, leading to backing from Howard Morgan.12:38–15:13 · Guest disagreement 1/10 Crowdsourcing Global Data Science Talent Over Traditional Quants Richard highlights that data scientists modeling abstract obfuscated data outperform traditional quants, comparing it to Google Translate outperforming human linguists.15:14–17:41 · Guest disagreement 0/10 The Numerai Dataset, Data Science Tournaments, and Payouts Richard details the structure of the million-row obfuscated dataset and explains why participants are motivated by both intellectual challenge and the substantial prize pool.17:42–20:31 · Guest disagreement 1/10 The Staking Mechanism and Alignment via Cryptocurrency Richard details Numerai's staking mechanism where contributors put their own cryptocurrency at risk to ensure skin in the game, comparing the model to Millennium's multi-manager alignment.20:33–23:38 · Guest disagreement 1/10 Model Turnover, Strategy Decay, and Crypto Market Resilience Ted probes whether crypto market volatility disrupted contributor incentives. Richard explains that user engagement doubled even when token prices dropped over 80%.23:39–25:50 · Guest disagreement 0/10 Data Architecture and Feature Engineering Standards Richard outlines the engineering criteria for features, emphasizing low churn and multi-decade applicability across global equities rather than short-term momentum signals.25:51–28:01 · Guest disagreement 0/10 From Daily Predictions to the Meta-Model and Factor Neutral Portfolio Richard explains the synthesis of user signals into the stake-weighted metamodel, which is subsequently passed through an optimizer to achieve country, sector, and factor neutrality.28:01–31:23 · Guest disagreement 2/10 Portfolio Risk Hedging and Combating Machine Learning Overfitting Ted questions whether anonymous machine learning models simply overfit or data-mine noise. Richard pushes back by distinguishing modern ML generalization from naive data mining and explains how they incentivize uncorrelated models.31:24–33:55 · Guest disagreement 0/10 Balance Sheet Construction, Leverage, and Idiosyncratic Alpha Richard breaks down portfolio leverage and risk constraints, highlighting how strict factor neutralization protects the strategy from systemic factor crashes like momentum shocks.33:57–38:06 · Guest disagreement 1/10 Core Breakthroughs: Staking Mechanism and Robust Drawdown Control Ted asks how allocators overcome the black-box perception of Numerai. Richard explains that sophisticated quantitative allocators look past process opacity to statistically significant alpha attribution.38:08–42:44 · Guest disagreement 1/10 Drawdown Alignment and Scalable Portfolio Architecture Richard discusses capacity discipline and drawdown mechanics, noting that contributors suffer token burns during model degradation, driving rapid corrective iteration.42:46–48:59 · Guest disagreement 0/10 Future R&D: Large Language Models and the Next Generation of Hedge Funds Richard discusses the potential of LLMs for feature generation and answers Ted's concluding personal and career questions in a collaborative, reflective tone.49:01–49:23 · Guest disagreement 0/10 Podcast Conclusion and Sponsored Episode Announcement Ted provides an outro monologue detailing how managers can sponsor insights on the Capital Allocators podcast.0:32–5:27 · Ted pushing back 0/10 Legal Disclaimer and Advisory Notice Ted opens with a legal disclaimer and introductory biographical framing before Richard shares his early fascination with stock trading and math education. The dynamic is purely conversational and exploratory.5:29–7:54 · Ted pushing back 1/10 Early Quantitative Experience and High-Capacity Machine Learning Ted asks how fundamental and technical data were integrated into Richard's early ML models. Richard explains the distinction between low-capacity short-term trading and high-capacity fundamental horizons.7:55–10:01 · Ted pushing back 1/10 Theory-Free Machine Learning vs Traditional Hypothesis Testing When Ted asks about formulating investment hypotheses, Richard gently corrects him by explaining that machine learning inverts traditional finance by being explicitly theory-free and purely data-driven.10:02–12:36 · Ted pushing back 0/10 The Genesis and Early Venture Backing of Numerai Richard explains how participating in Kaggle competitions inspired him to crowdsource quantitative modeling on open, obfuscated datasets, leading to backing from Howard Morgan.12:38–15:13 · Ted pushing back 0/10 Crowdsourcing Global Data Science Talent Over Traditional Quants Richard highlights that data scientists modeling abstract obfuscated data outperform traditional quants, comparing it to Google Translate outperforming human linguists.15:14–17:41 · Ted pushing back 0/10 The Numerai Dataset, Data Science Tournaments, and Payouts Richard details the structure of the million-row obfuscated dataset and explains why participants are motivated by both intellectual challenge and the substantial prize pool.17:42–20:31 · Ted pushing back 0/10 The Staking Mechanism and Alignment via Cryptocurrency Richard details Numerai's staking mechanism where contributors put their own cryptocurrency at risk to ensure skin in the game, comparing the model to Millennium's multi-manager alignment.20:33–23:38 · Ted pushing back 1/10 Model Turnover, Strategy Decay, and Crypto Market Resilience Ted probes whether crypto market volatility disrupted contributor incentives. Richard explains that user engagement doubled even when token prices dropped over 80%.23:39–25:50 · Ted pushing back 0/10 Data Architecture and Feature Engineering Standards Richard outlines the engineering criteria for features, emphasizing low churn and multi-decade applicability across global equities rather than short-term momentum signals.25:51–28:01 · Ted pushing back 0/10 From Daily Predictions to the Meta-Model and Factor Neutral Portfolio Richard explains the synthesis of user signals into the stake-weighted metamodel, which is subsequently passed through an optimizer to achieve country, sector, and factor neutrality.28:01–31:23 · Ted pushing back 2/10 Portfolio Risk Hedging and Combating Machine Learning Overfitting Ted questions whether anonymous machine learning models simply overfit or data-mine noise. Richard pushes back by distinguishing modern ML generalization from naive data mining and explains how they incentivize uncorrelated models.31:24–33:55 · Ted pushing back 1/10 Balance Sheet Construction, Leverage, and Idiosyncratic Alpha Richard breaks down portfolio leverage and risk constraints, highlighting how strict factor neutralization protects the strategy from systemic factor crashes like momentum shocks.33:57–38:06 · Ted pushing back 1/10 Core Breakthroughs: Staking Mechanism and Robust Drawdown Control Ted asks how allocators overcome the black-box perception of Numerai. Richard explains that sophisticated quantitative allocators look past process opacity to statistically significant alpha attribution.38:08–42:44 · Ted pushing back 1/10 Drawdown Alignment and Scalable Portfolio Architecture Richard discusses capacity discipline and drawdown mechanics, noting that contributors suffer token burns during model degradation, driving rapid corrective iteration.42:46–48:59 · Ted pushing back 0/10 Future R&D: Large Language Models and the Next Generation of Hedge Funds Richard discusses the potential of LLMs for feature generation and answers Ted's concluding personal and career questions in a collaborative, reflective tone.49:01–49:23 · Ted pushing back 0/10 Podcast Conclusion and Sponsored Episode Announcement Ted provides an outro monologue detailing how managers can sponsor insights on the Capital Allocators podcast.

speaking balance: gold is Ted, purple is the guest (3 minute bins)

0:00 · Ted 62.4% · guest 37.6%0:00 · Ted 62.4% · guest 37.6%3:00 · Ted 2.3% · guest 97.7%3:00 · Ted 2.3% · guest 97.7%6:00 · Ted 17% · guest 83%6:00 · Ted 17% · guest 83%9:00 · Ted 11.6% · guest 88.4%9:00 · Ted 11.6% · guest 88.4%12:00 · Ted 12.1% · guest 87.9%12:00 · Ted 12.1% · guest 87.9%15:00 · Ted 8.4% · guest 91.6%15:00 · Ted 8.4% · guest 91.6%18:00 · Ted 16.8% · guest 83.2%18:00 · Ted 16.8% · guest 83.2%21:00 · Ted 16.1% · guest 83.9%21:00 · Ted 16.1% · guest 83.9%24:00 · Ted 13.5% · guest 86.5%24:00 · Ted 13.5% · guest 86.5%27:00 · Ted 19.2% · guest 80.8%27:00 · Ted 19.2% · guest 80.8%30:00 · Ted 8.5% · guest 91.5%30:00 · Ted 8.5% · guest 91.5%33:00 · Ted 11.5% · guest 88.5%33:00 · Ted 11.5% · guest 88.5%36:00 · Ted 17.8% · guest 82.2%36:00 · Ted 17.8% · guest 82.2%39:00 · Ted 10.5% · guest 89.5%39:00 · Ted 10.5% · guest 89.5%42:00 · Ted 13.6% · guest 86.4%42:00 · Ted 13.6% · guest 86.4%45:00 · Ted 5.8% · guest 94.2%45:00 · Ted 5.8% · guest 94.2%48:00 · Ted 46.5% · guest 53.5%48:00 · Ted 46.5% · guest 53.5%
Sharpest disagreement ▶ 7:55 Rejecting traditional hypothesis-driven investing

Richard immediately rejects Ted's premise about forming investment hypotheses, emphasizing that machine learning is defined by operating without predefined theories.

Hardest push from Ted ▶ 28:33 Challenging machine learning as data mining

Ted presses Richard on whether crowdsourced, anonymized machine learning models are merely overfitting noise and executing data mining rather than finding real signal.

Biggest teaching moment ▶ 14:05 Data scientists vs traditional quants

Richard reframes the talent model of quantitative finance, explaining that domain-blind data scientists outmodel professional quants in the same way language-agnostic models outperform human translators.

Ted holds their own ▶ 35:54 Questioning allocator comfort with black-box architectures

Ted leverages allocator psychology to challenge Richard on how institutional investors can bridge the gap between necessary due diligence and an opaque, double-anonymized black box.

the scores for every segment, with the reasoning behind each
ChapterTopicTed as informed peerGuest teachingGuest disagreementTed pushing backWhy
Legal Disclaimer and Advisory Notice 1100 Ted opens with a legal disclaimer and introductory biographical framing before Richard shares his early fascination with stock trading and math education. The dynamic is purely conversational and exploratory.
Early Quantitative Experience and High-Capacity Machine Learning 3311 Ted asks how fundamental and technical data were integrated into Richard's early ML models. Richard explains the distinction between low-capacity short-term trading and high-capacity fundamental horizons.
Theory-Free Machine Learning vs Traditional Hypothesis Testing 3521 When Ted asks about formulating investment hypotheses, Richard gently corrects him by explaining that machine learning inverts traditional finance by being explicitly theory-free and purely data-driven.
The Genesis and Early Venture Backing of Numerai 2300 Richard explains how participating in Kaggle competitions inspired him to crowdsource quantitative modeling on open, obfuscated datasets, leading to backing from Howard Morgan.
Crowdsourcing Global Data Science Talent Over Traditional Quants 2410 Richard highlights that data scientists modeling abstract obfuscated data outperform traditional quants, comparing it to Google Translate outperforming human linguists.
The Numerai Dataset, Data Science Tournaments, and Payouts 2300 Richard details the structure of the million-row obfuscated dataset and explains why participants are motivated by both intellectual challenge and the substantial prize pool.
The Staking Mechanism and Alignment via Cryptocurrency 3410 Richard details Numerai's staking mechanism where contributors put their own cryptocurrency at risk to ensure skin in the game, comparing the model to Millennium's multi-manager alignment.
Model Turnover, Strategy Decay, and Crypto Market Resilience 3311 Ted probes whether crypto market volatility disrupted contributor incentives. Richard explains that user engagement doubled even when token prices dropped over 80%.
Data Architecture and Feature Engineering Standards 3400 Richard outlines the engineering criteria for features, emphasizing low churn and multi-decade applicability across global equities rather than short-term momentum signals.
From Daily Predictions to the Meta-Model and Factor Neutral Portfolio 3400 Richard explains the synthesis of user signals into the stake-weighted metamodel, which is subsequently passed through an optimizer to achieve country, sector, and factor neutrality.
Portfolio Risk Hedging and Combating Machine Learning Overfitting 4422 Ted questions whether anonymous machine learning models simply overfit or data-mine noise. Richard pushes back by distinguishing modern ML generalization from naive data mining and explains how they incentivize uncorrelated models.
Balance Sheet Construction, Leverage, and Idiosyncratic Alpha 3301 Richard breaks down portfolio leverage and risk constraints, highlighting how strict factor neutralization protects the strategy from systemic factor crashes like momentum shocks.
Core Breakthroughs: Staking Mechanism and Robust Drawdown Control 4411 Ted asks how allocators overcome the black-box perception of Numerai. Richard explains that sophisticated quantitative allocators look past process opacity to statistically significant alpha attribution.
Drawdown Alignment and Scalable Portfolio Architecture 3311 Richard discusses capacity discipline and drawdown mechanics, noting that contributors suffer token burns during model degradation, driving rapid corrective iteration.
Future R&D: Large Language Models and the Next Generation of Hedge Funds 2200 Richard discusses the potential of LLMs for feature generation and answers Ted's concluding personal and career questions in a collaborative, reflective tone.
Podcast Conclusion and Sponsored Episode Announcement 0000 Ted provides an outro monologue detailing how managers can sponsor insights on the Capital Allocators podcast.

Statements from this episode (29)

Insight
Craib: Short-term trading strategies lack the capacity to scale
“If you do short-term trading, your capacity is very low. So if you trade very frequently in very small stocks, you're moving the market a lot, and it's very hard to have a strategy that scales.”
Richard Craib May 11, 2023 ▶ 7:06
Insight
Craib: Machine learning quant investing relies entirely on theory-free models
“That's, in some ways, the opposite of machine learning, which is a path of saying, we have no idea what patterns are Real or we don't have any theory. We don't have any hypothesis, but we do have a large data set. What can we learn in that data set that will g…”
Richard Craib May 11, 2023 ▶ 8:22
Disclosure
Craib: Numerai closed seed round the day personal savings ran out
“Turns out I spent that 250,000 dollars in about three months and ran out of money pretty much the day of first round seeding the company as with a venture investment from Howard Morgan and Bill Trenchard.”
Richard Craib May 11, 2023 ▶ 12:08
Assertion Not checkable as stated
Craib: Numerai contributors lack traditional quantitative finance backgrounds
“Data science is not quant. None of these people were quants. None of them had any background in quant. They're just the types of people who can solve abstract problems to do with data.”
Richard Craib May 11, 2023 ▶ 14:26
Insight
Craib: Numerai outperforms pros similarly to Google Translate outperforming human linguists
“They just model the data, and that's what, what's happening at Numeri, and just like Google Translate, they can be a lot better than a professional Translator on certain languages, and Numeri can be better than a professional investor.”
Richard Craib May 11, 2023 ▶ 14:58
Disclosure
Craib: Numerai's public dataset contains over one million rows
“What they're doing is downloading a dataset, and the dataset is about, I think it's over a million rows long. So it's just millions of numbers between zero and one. You can download it, by the way, anyone can. You can just go to a website and download it. But …”
Richard Craib May 11, 2023 ▶ 15:22
Assertion Supported
Craib: Numerai has paid contributors over $30 million, surpassing Kaggle payouts
“Numeri has created the highest paying data science tournament on the internet. We pay out more than Kaggle does, and we've paid out over thirty million dollars since we started.”
Richard Craib May 11, 2023 ▶ 17:12
Assertion Supported
Craib: Numerai pioneered cryptocurrency staking back in 2017
“And that mechanism is something we pioneered back in 2017, long before most people had heard about staking.”
Richard Craib May 11, 2023 ▶ 18:46
Assertion Supported
Craib: Numerai's native cryptocurrency is more liquid than some stocks
“In fact, our cryptocurrency is more liquid than some stocks, and so people will earn some of it, maybe they'll earn 10,000 dollars by staking on Numurai, and then they'll go to Coinbase and sell it into dollars, or they'll keep staking it on their model.”
Richard Craib May 11, 2023 ▶ 20:15
Opinion
Craib: Numerai operates as an AI-driven version of Millennium Management
“I talk about Numerize being a kind of AI version of Millennium. Millennium, they do not have all the same people they had 20 years ago working at Millennium. They have a totally dynamic system of PMs coming in and out, and that's how year over year they've pro…”
Richard Craib May 11, 2023 ▶ 21:23
Assertion Partly supported
Craib: Numerai user base doubled in 2018 despite token crashing 90%
“Well, there was a time in 2018 when NMR, our cryptocurrency, was down maybe, yeah, more than 80%, maybe 90%. And in that year, our user base doubled.”
Richard Craib May 11, 2023 ▶ 22:01
Assertion Not checkable as stated
Craib: Numerai hosted the largest in-person gathering of Kaggle Grandmasters
“We had a big conference in San Francisco called Numicon. Where we had probably the largest gathering of Kaggle grandmasters in real life ever.”
Richard Craib May 11, 2023 ▶ 22:45
Assertion Partly supported
Craib: Numerai scaled to 2,000 data features and 5,500 active modelers
“There are now almost 2000. Three years ago, it was 40. We only had 40 features, and now it's 2000, and over that same period, the number of data scientists submitting models has grown from a hundred to five and a half thousand.”
Richard Craib May 11, 2023 ▶ 24:38
Insight
Craib: Numerai avoids high-turnover trading signals due to execution limits
“If we stuck a seven day momentum feature into Numerai's dataset, the Numerai users would like it and their models would pick up on it. But coming to trade execution, we wouldn't want to trade that fast. And then we wouldn't actually make any money from that.”
Richard Craib May 11, 2023 ▶ 25:20
Disclosure
Craib: Numerai holds portfolio positions for three to four months
“We get the predictions sent to us now every day, but we only trade quite slowly. So we'll only turn over the portfolio quite slowly, minimizing our market impact. And we end up holding positions for three or four months.”
Richard Craib May 11, 2023 ▶ 26:00
Disclosure
Craib: Numerai creates meta-model via stake-weighted average of 5,000 predictions
“They're giving us a signal, which is just a long vector of predictions, 5000 predictions on every stock in the universe. And we take those predictions. We compute the stake weighted average.”
Richard Craib May 11, 2023 ▶ 27:02
Disclosure
Craib: Numerai's portfolio optimizer eliminates country, sector, and factor risk
“And Numeri is factor neutral. So we will have the optimizer take out country risk, sector risk, factor risk.”
Richard Craib May 11, 2023 ▶ 27:53
Insight
Craib: Millennium Management thrives by combining numerous uncorrelated strategies
“One of the reasons Millennium works is because they can get strategies that are uncorrelated, and that's what boosts the sharp. They have so many different strategies that are all uncorrelated, it's starts to become very hard to have a down year.”
Richard Craib May 11, 2023 ▶ 29:38
Insight
Craib: Uncorrelated mediocre models help Numerai more than correlated good models
“If you make a mediocre model that's very uncorrelated, you can do particularly well. Because of its lack of correlation is actually helping Numeri more than a good model that's correlated with the one we already have.”
Richard Craib May 11, 2023 ▶ 30:28
Disclosure
Craib: Numerai holds 500 long and short positions under 3% cap
“It's about 500 stocks long and 500 stocks short, and a lot of the position sizes are kind of equal weight. Some that are higher than others, but it's basically, we never have more than three percent of the fund in one name.”
Richard Craib May 11, 2023 ▶ 31:32
What-if
Craib: Jim Simons would make no money without leverage
“Jim Simons would make absolutely no money if it weren't for leverage. That's just how, how it works.”
Richard Craib May 11, 2023 ▶ 31:59
Assertion Not checkable as stated
Craib: Numerai was unaffected by the January momentum crash
“For example, you know, in January, momentum, the factor crashed, and a lot of quants did badly. Numeri is neutral to momentum, so we basically didn't even notice this crash because we're hedged to that risk, and we do that with as many things as possible to ma…”
Richard Craib May 11, 2023 ▶ 32:23
Insight
Craib: Requiring financial stakes forced contributors to submit their best models
“The biggest one was cracking staking. I mean, we had a period where, yes, we were getting a lot of people submitting models, but we couldn't trust them. We couldn't trust that they would keep working out of sample. We didn't even know who the people were, but …”
Richard Craib May 11, 2023 ▶ 34:07
Opinion
Craib: Numerai accesses unhireable global talent through crowdsourcing
“So I think it's very clear to me that for right now, Numeri has a very big edge on talent. We have people working on our data that aren't even hireable.”
Richard Craib May 11, 2023 ▶ 35:44
Disclosure
Craib: Numerai does not build or know its models' underlying architectures
“We don't even build the models. Our data scientists do, and we don't even know what they built. Like, they might have used a neural network, or they might have done something else. We don't know.”
Richard Craib May 11, 2023 ▶ 36:23
Insight
Craib: Markets require continuous crowdsourcing because they are never fully solved
“With a single model, you can get very high accuracy with that type of image detection. That's sort of like a solved problem. And therefore, why would you need to do crowdsourcing? That's a solved problem. But stock market is never solved. It's a permanent race…”
Richard Craib May 11, 2023 ▶ 41:53
Insight
Craib: A 0.5% edge increase produces massive Sharpe improvements
“To go from 52 to 52 and a half is massive in terms of your Sharpe ratio returns, volatility. And so the fact is, it's just one of these industries where a tiny bit helps.”
Richard Craib May 11, 2023 ▶ 42:27
Disclosure
Craib: Numerai plans to build dataset features using LLMs
“Large language models, which is probably the most hyped thing in the world, is actually something I want to have Numeri become the best at. And I think we have a shot at it because Machine learning is so in our DNA. So I think that's something I'm excited abou…”
Richard Craib May 11, 2023 ▶ 43:07
Opinion
Craib: High PE returns require less skill than market-neutral quant investing
“If a PE fund makes 40% in a single year because they were actually holding positions that were two times market beta and the market went up, That doesn't have anything close to the investment skill of a quant fund that's making money kind of out of nothing by …”
Richard Craib May 11, 2023 ▶ 45:14
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 700 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.