Aug 29, 2025 · 1h 18m · latent-space

Better Data is All You Need — Ari Morcos, Datology

Ari Morcos · 1h 3m spoken Shawn Wang · 6m spoken Alessio Fanelli · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, Datology CEO Ari Morcos discusses why automated data curation is AI's most critical frontier, demonstrating how algorithmic filtering, synthetic rephrasing, and curriculum sequencing enable faster training, superior accuracy, and significantly smaller enterprise models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 14% of the talking time here. How this is scored →

The hosts as informed peer 4.9 Guest teaching 5.4 Guest disagreement 2.6 The hosts pushing back 2.2
05100:0020:0040:001:00:002:04–7:34 · The hosts as informed peer 4/10 Ari Morcos's Background: Neuroscience, Inductive Biases, and the Bitter Lesson Swyx demonstrates technical familiarity with vision transformers and weight mapping between CNNs and ViTs. Ari walks through his transition from neuroscience to inductive biases and his realization that architecture matters far less than data scale.7:34–10:15 · The hosts as informed peer 3/10 The Strategic Shift: Why Data Is AI's Most Under-Invested Frontier Ari strongly criticizes foundational scaling papers by Kaplan and Chinchilla for assuming IID data distributions, calling the premise insane given standard computer science adages. The hosts listen attentively with minimal interjections.10:15–12:31 · The hosts as informed peer 5/10 Historical Prejudices and Misaligned Incentives in Data Research Alessio asks why data research incentives are misaligned given high-profile data companies like Scale AI. Ari differentiates academic research culture, which treated data as fixed grunt work, from industrial priorities and supervised dataset limitations.12:31–14:49 · The hosts as informed peer 4/10 The Self-Supervised Paradigm Shift: Moving to the Underfitting Regime Ari presents a contrarian view on transformer architectures, arguing they are just one of many viable architectures while self-supervised learning on unlabeled data was the true revolutionary breakthrough.14:49–19:34 · The hosts as informed peer 5/10 Data Engineering at Frontier Labs and the Failure of Human Intuition Swyx notes his belief in automated end-to-end learning. Ari backs this up with findings from DCLM showing domain-expert grad students could not predict data filtering classifiers better than chance, proving human intuition fails at scale.19:34–22:46 · The hosts as informed peer 4/10 Concept Complexity, Data Redundancy, and Zero-Shot Generalization Alessio presses Ari on how to empirically define boundaries for concepts. Ari uses an analogy comparing elephants and dog breeds to explain why redundancy requirements vary with semantic complexity.22:46–26:24 · The hosts as informed peer 6/10 The Evolution of Public Datasets, Code Predictors, and Copyright Litigation Swyx rapid-fires standard pre-training datasets including GitHub, Arxiv, and Books3, citing litigation against Anthropic and Meta. Ari notes counterintuitive findings from StarCoder showing GitHub stars do not predict code quality.26:24–28:57 · The hosts as informed peer 4/10 Overcoming the Limits of Power-Law Scaling via Data Efficiency Ari explains the mathematical connection between decaying marginal information gain and power-law scaling, citing his NeurIPS paper on bending scaling curves through active data pruning.28:57–32:28 · The hosts as informed peer 6/10 Quantifying Data Curation Gains Across Speed, Quality, and Size Swyx challenges whether Datology's performance claims are simply overfitting or training to the test on public benchmarks. Ari details their strict protocol using held-out evaluations and unreleased eval suites.32:28–39:12 · The hosts as informed peer 5/10 Why Specialized Data Companies Outperform Internal AI Lab Teams Alessio questions whether private curation layers create friction with open-source dataset creators. Ari details Datology's moat balance between open scientific intuition and proprietary engineering know-how.39:12–45:58 · The hosts as informed peer 6/10 Curation Methodologies and the Power of Synthetic Data Rephrasing Swyx explores synthetic data, drawing distinctions between distillation and steganography. Ari outlines Datology's core approach of rephrasing existing data to bypass mode collapse and teacher model limitations.45:58–49:03 · The hosts as informed peer 5/10 The Supremacy of Data Diversity and the Reality of Epoching Swyx brings up the 'Textbooks Are All You Need' hypothesis. Ari refutes narrow distribution claims, asserting data diversity is paramount and epoching high-quality tokens consistently outperforms low-grade new data.49:03–52:56 · The hosts as informed peer 4/10 The Resurgence of Curriculum Learning in Modern LLM Training Ari explains the conceptual graph theory behind curriculum learning, clarifying that discrete curricula now succeed in underfitted LLM regimes where they previously showed minimal value in saturated supervised settings.52:56–1:00:19 · The hosts as informed peer 6/10 Optimizing Pre-Training Data for Downstream Post-Training and Alignment Swyx challenges the idea of pre-training dependencies given the consensus that post-training is merely capability elicitation. Ari counters that pre-training determines the slope of test-time compute scaling and alignment robustness.1:00:19–1:03:02 · The hosts as informed peer 7/10 Model Pruning Bottlenecks and Complementary Optimization Strategies Swyx brings deep context regarding Jonathan Frankle's lottery ticket hypothesis and parameter pruning. Ari explains how unstructured pruning fell out of favor due to GPU sparse matrix multiply inefficiencies.1:03:03–1:06:53 · The hosts as informed peer 6/10 Future Architecture: Small Models, Test-Time Compute, and Cognitive Cores Swyx references Andre Karpathy's 'cognitive core' concept. Ari agrees that storing raw knowledge inside network parameters is inefficient and smaller models will dominate test-time compute workloads.1:06:53–1:09:46 · The hosts as informed peer 5/10 Case Study: The Arcee Foundation Model and Compounding Curation Gains Alessio asks for concrete metrics from the Arcee training run. Ari details how condensing 25T tokens into 7T allowed a 4.5B model to beat larger baselines before reaching one trillion tokens.1:09:46–1:14:24 · The hosts as informed peer 4/10 Data Valuation: AI's NP-Complete Problem and the Datology Profile Swyx asks recruiting and technical frontier questions. Ari defines target-specific data valuation as AI's central NP-complete problem and outlines Datology's mission to automate zero-shot domain adaptation.2:04–7:34 · Guest teaching 6/10 Ari Morcos's Background: Neuroscience, Inductive Biases, and the Bitter Lesson Swyx demonstrates technical familiarity with vision transformers and weight mapping between CNNs and ViTs. Ari walks through his transition from neuroscience to inductive biases and his realization that architecture matters far less than data scale.7:34–10:15 · Guest teaching 6/10 The Strategic Shift: Why Data Is AI's Most Under-Invested Frontier Ari strongly criticizes foundational scaling papers by Kaplan and Chinchilla for assuming IID data distributions, calling the premise insane given standard computer science adages. The hosts listen attentively with minimal interjections.10:15–12:31 · Guest teaching 5/10 Historical Prejudices and Misaligned Incentives in Data Research Alessio asks why data research incentives are misaligned given high-profile data companies like Scale AI. Ari differentiates academic research culture, which treated data as fixed grunt work, from industrial priorities and supervised dataset limitations.12:31–14:49 · Guest teaching 6/10 The Self-Supervised Paradigm Shift: Moving to the Underfitting Regime Ari presents a contrarian view on transformer architectures, arguing they are just one of many viable architectures while self-supervised learning on unlabeled data was the true revolutionary breakthrough.14:49–19:34 · Guest teaching 7/10 Data Engineering at Frontier Labs and the Failure of Human Intuition Swyx notes his belief in automated end-to-end learning. Ari backs this up with findings from DCLM showing domain-expert grad students could not predict data filtering classifiers better than chance, proving human intuition fails at scale.19:34–22:46 · Guest teaching 6/10 Concept Complexity, Data Redundancy, and Zero-Shot Generalization Alessio presses Ari on how to empirically define boundaries for concepts. Ari uses an analogy comparing elephants and dog breeds to explain why redundancy requirements vary with semantic complexity.22:46–26:24 · Guest teaching 5/10 The Evolution of Public Datasets, Code Predictors, and Copyright Litigation Swyx rapid-fires standard pre-training datasets including GitHub, Arxiv, and Books3, citing litigation against Anthropic and Meta. Ari notes counterintuitive findings from StarCoder showing GitHub stars do not predict code quality.26:24–28:57 · Guest teaching 6/10 Overcoming the Limits of Power-Law Scaling via Data Efficiency Ari explains the mathematical connection between decaying marginal information gain and power-law scaling, citing his NeurIPS paper on bending scaling curves through active data pruning.28:57–32:28 · Guest teaching 6/10 Quantifying Data Curation Gains Across Speed, Quality, and Size Swyx challenges whether Datology's performance claims are simply overfitting or training to the test on public benchmarks. Ari details their strict protocol using held-out evaluations and unreleased eval suites.32:28–39:12 · Guest teaching 4/10 Why Specialized Data Companies Outperform Internal AI Lab Teams Alessio questions whether private curation layers create friction with open-source dataset creators. Ari details Datology's moat balance between open scientific intuition and proprietary engineering know-how.39:12–45:58 · Guest teaching 5/10 Curation Methodologies and the Power of Synthetic Data Rephrasing Swyx explores synthetic data, drawing distinctions between distillation and steganography. Ari outlines Datology's core approach of rephrasing existing data to bypass mode collapse and teacher model limitations.45:58–49:03 · Guest teaching 5/10 The Supremacy of Data Diversity and the Reality of Epoching Swyx brings up the 'Textbooks Are All You Need' hypothesis. Ari refutes narrow distribution claims, asserting data diversity is paramount and epoching high-quality tokens consistently outperforms low-grade new data.49:03–52:56 · Guest teaching 6/10 The Resurgence of Curriculum Learning in Modern LLM Training Ari explains the conceptual graph theory behind curriculum learning, clarifying that discrete curricula now succeed in underfitted LLM regimes where they previously showed minimal value in saturated supervised settings.52:56–1:00:19 · Guest teaching 6/10 Optimizing Pre-Training Data for Downstream Post-Training and Alignment Swyx challenges the idea of pre-training dependencies given the consensus that post-training is merely capability elicitation. Ari counters that pre-training determines the slope of test-time compute scaling and alignment robustness.1:00:19–1:03:02 · Guest teaching 4/10 Model Pruning Bottlenecks and Complementary Optimization Strategies Swyx brings deep context regarding Jonathan Frankle's lottery ticket hypothesis and parameter pruning. Ari explains how unstructured pruning fell out of favor due to GPU sparse matrix multiply inefficiencies.1:03:03–1:06:53 · Guest teaching 5/10 Future Architecture: Small Models, Test-Time Compute, and Cognitive Cores Swyx references Andre Karpathy's 'cognitive core' concept. Ari agrees that storing raw knowledge inside network parameters is inefficient and smaller models will dominate test-time compute workloads.1:06:53–1:09:46 · Guest teaching 5/10 Case Study: The Arcee Foundation Model and Compounding Curation Gains Alessio asks for concrete metrics from the Arcee training run. Ari details how condensing 25T tokens into 7T allowed a 4.5B model to beat larger baselines before reaching one trillion tokens.1:09:46–1:14:24 · Guest teaching 5/10 Data Valuation: AI's NP-Complete Problem and the Datology Profile Swyx asks recruiting and technical frontier questions. Ari defines target-specific data valuation as AI's central NP-complete problem and outlines Datology's mission to automate zero-shot domain adaptation.2:04–7:34 · Guest disagreement 2/10 Ari Morcos's Background: Neuroscience, Inductive Biases, and the Bitter Lesson Swyx demonstrates technical familiarity with vision transformers and weight mapping between CNNs and ViTs. Ari walks through his transition from neuroscience to inductive biases and his realization that architecture matters far less than data scale.7:34–10:15 · Guest disagreement 4/10 The Strategic Shift: Why Data Is AI's Most Under-Invested Frontier Ari strongly criticizes foundational scaling papers by Kaplan and Chinchilla for assuming IID data distributions, calling the premise insane given standard computer science adages. The hosts listen attentively with minimal interjections.10:15–12:31 · Guest disagreement 3/10 Historical Prejudices and Misaligned Incentives in Data Research Alessio asks why data research incentives are misaligned given high-profile data companies like Scale AI. Ari differentiates academic research culture, which treated data as fixed grunt work, from industrial priorities and supervised dataset limitations.12:31–14:49 · Guest disagreement 3/10 The Self-Supervised Paradigm Shift: Moving to the Underfitting Regime Ari presents a contrarian view on transformer architectures, arguing they are just one of many viable architectures while self-supervised learning on unlabeled data was the true revolutionary breakthrough.14:49–19:34 · Guest disagreement 4/10 Data Engineering at Frontier Labs and the Failure of Human Intuition Swyx notes his belief in automated end-to-end learning. Ari backs this up with findings from DCLM showing domain-expert grad students could not predict data filtering classifiers better than chance, proving human intuition fails at scale.19:34–22:46 · Guest disagreement 2/10 Concept Complexity, Data Redundancy, and Zero-Shot Generalization Alessio presses Ari on how to empirically define boundaries for concepts. Ari uses an analogy comparing elephants and dog breeds to explain why redundancy requirements vary with semantic complexity.22:46–26:24 · Guest disagreement 2/10 The Evolution of Public Datasets, Code Predictors, and Copyright Litigation Swyx rapid-fires standard pre-training datasets including GitHub, Arxiv, and Books3, citing litigation against Anthropic and Meta. Ari notes counterintuitive findings from StarCoder showing GitHub stars do not predict code quality.26:24–28:57 · Guest disagreement 3/10 Overcoming the Limits of Power-Law Scaling via Data Efficiency Ari explains the mathematical connection between decaying marginal information gain and power-law scaling, citing his NeurIPS paper on bending scaling curves through active data pruning.28:57–32:28 · Guest disagreement 3/10 Quantifying Data Curation Gains Across Speed, Quality, and Size Swyx challenges whether Datology's performance claims are simply overfitting or training to the test on public benchmarks. Ari details their strict protocol using held-out evaluations and unreleased eval suites.32:28–39:12 · Guest disagreement 2/10 Why Specialized Data Companies Outperform Internal AI Lab Teams Alessio questions whether private curation layers create friction with open-source dataset creators. Ari details Datology's moat balance between open scientific intuition and proprietary engineering know-how.39:12–45:58 · Guest disagreement 2/10 Curation Methodologies and the Power of Synthetic Data Rephrasing Swyx explores synthetic data, drawing distinctions between distillation and steganography. Ari outlines Datology's core approach of rephrasing existing data to bypass mode collapse and teacher model limitations.45:58–49:03 · Guest disagreement 3/10 The Supremacy of Data Diversity and the Reality of Epoching Swyx brings up the 'Textbooks Are All You Need' hypothesis. Ari refutes narrow distribution claims, asserting data diversity is paramount and epoching high-quality tokens consistently outperforms low-grade new data.49:03–52:56 · Guest disagreement 2/10 The Resurgence of Curriculum Learning in Modern LLM Training Ari explains the conceptual graph theory behind curriculum learning, clarifying that discrete curricula now succeed in underfitted LLM regimes where they previously showed minimal value in saturated supervised settings.52:56–1:00:19 · Guest disagreement 4/10 Optimizing Pre-Training Data for Downstream Post-Training and Alignment Swyx challenges the idea of pre-training dependencies given the consensus that post-training is merely capability elicitation. Ari counters that pre-training determines the slope of test-time compute scaling and alignment robustness.1:00:19–1:03:02 · Guest disagreement 2/10 Model Pruning Bottlenecks and Complementary Optimization Strategies Swyx brings deep context regarding Jonathan Frankle's lottery ticket hypothesis and parameter pruning. Ari explains how unstructured pruning fell out of favor due to GPU sparse matrix multiply inefficiencies.1:03:03–1:06:53 · Guest disagreement 2/10 Future Architecture: Small Models, Test-Time Compute, and Cognitive Cores Swyx references Andre Karpathy's 'cognitive core' concept. Ari agrees that storing raw knowledge inside network parameters is inefficient and smaller models will dominate test-time compute workloads.1:06:53–1:09:46 · Guest disagreement 1/10 Case Study: The Arcee Foundation Model and Compounding Curation Gains Alessio asks for concrete metrics from the Arcee training run. Ari details how condensing 25T tokens into 7T allowed a 4.5B model to beat larger baselines before reaching one trillion tokens.1:09:46–1:14:24 · Guest disagreement 2/10 Data Valuation: AI's NP-Complete Problem and the Datology Profile Swyx asks recruiting and technical frontier questions. Ari defines target-specific data valuation as AI's central NP-complete problem and outlines Datology's mission to automate zero-shot domain adaptation.2:04–7:34 · The hosts pushing back 1/10 Ari Morcos's Background: Neuroscience, Inductive Biases, and the Bitter Lesson Swyx demonstrates technical familiarity with vision transformers and weight mapping between CNNs and ViTs. Ari walks through his transition from neuroscience to inductive biases and his realization that architecture matters far less than data scale.7:34–10:15 · The hosts pushing back 1/10 The Strategic Shift: Why Data Is AI's Most Under-Invested Frontier Ari strongly criticizes foundational scaling papers by Kaplan and Chinchilla for assuming IID data distributions, calling the premise insane given standard computer science adages. The hosts listen attentively with minimal interjections.10:15–12:31 · The hosts pushing back 2/10 Historical Prejudices and Misaligned Incentives in Data Research Alessio asks why data research incentives are misaligned given high-profile data companies like Scale AI. Ari differentiates academic research culture, which treated data as fixed grunt work, from industrial priorities and supervised dataset limitations.12:31–14:49 · The hosts pushing back 1/10 The Self-Supervised Paradigm Shift: Moving to the Underfitting Regime Ari presents a contrarian view on transformer architectures, arguing they are just one of many viable architectures while self-supervised learning on unlabeled data was the true revolutionary breakthrough.14:49–19:34 · The hosts pushing back 2/10 Data Engineering at Frontier Labs and the Failure of Human Intuition Swyx notes his belief in automated end-to-end learning. Ari backs this up with findings from DCLM showing domain-expert grad students could not predict data filtering classifiers better than chance, proving human intuition fails at scale.19:34–22:46 · The hosts pushing back 3/10 Concept Complexity, Data Redundancy, and Zero-Shot Generalization Alessio presses Ari on how to empirically define boundaries for concepts. Ari uses an analogy comparing elephants and dog breeds to explain why redundancy requirements vary with semantic complexity.22:46–26:24 · The hosts pushing back 2/10 The Evolution of Public Datasets, Code Predictors, and Copyright Litigation Swyx rapid-fires standard pre-training datasets including GitHub, Arxiv, and Books3, citing litigation against Anthropic and Meta. Ari notes counterintuitive findings from StarCoder showing GitHub stars do not predict code quality.26:24–28:57 · The hosts pushing back 2/10 Overcoming the Limits of Power-Law Scaling via Data Efficiency Ari explains the mathematical connection between decaying marginal information gain and power-law scaling, citing his NeurIPS paper on bending scaling curves through active data pruning.28:57–32:28 · The hosts pushing back 5/10 Quantifying Data Curation Gains Across Speed, Quality, and Size Swyx challenges whether Datology's performance claims are simply overfitting or training to the test on public benchmarks. Ari details their strict protocol using held-out evaluations and unreleased eval suites.32:28–39:12 · The hosts pushing back 3/10 Why Specialized Data Companies Outperform Internal AI Lab Teams Alessio questions whether private curation layers create friction with open-source dataset creators. Ari details Datology's moat balance between open scientific intuition and proprietary engineering know-how.39:12–45:58 · The hosts pushing back 2/10 Curation Methodologies and the Power of Synthetic Data Rephrasing Swyx explores synthetic data, drawing distinctions between distillation and steganography. Ari outlines Datology's core approach of rephrasing existing data to bypass mode collapse and teacher model limitations.45:58–49:03 · The hosts pushing back 2/10 The Supremacy of Data Diversity and the Reality of Epoching Swyx brings up the 'Textbooks Are All You Need' hypothesis. Ari refutes narrow distribution claims, asserting data diversity is paramount and epoching high-quality tokens consistently outperforms low-grade new data.49:03–52:56 · The hosts pushing back 2/10 The Resurgence of Curriculum Learning in Modern LLM Training Ari explains the conceptual graph theory behind curriculum learning, clarifying that discrete curricula now succeed in underfitted LLM regimes where they previously showed minimal value in saturated supervised settings.52:56–1:00:19 · The hosts pushing back 5/10 Optimizing Pre-Training Data for Downstream Post-Training and Alignment Swyx challenges the idea of pre-training dependencies given the consensus that post-training is merely capability elicitation. Ari counters that pre-training determines the slope of test-time compute scaling and alignment robustness.1:00:19–1:03:02 · The hosts pushing back 2/10 Model Pruning Bottlenecks and Complementary Optimization Strategies Swyx brings deep context regarding Jonathan Frankle's lottery ticket hypothesis and parameter pruning. Ari explains how unstructured pruning fell out of favor due to GPU sparse matrix multiply inefficiencies.1:03:03–1:06:53 · The hosts pushing back 2/10 Future Architecture: Small Models, Test-Time Compute, and Cognitive Cores Swyx references Andre Karpathy's 'cognitive core' concept. Ari agrees that storing raw knowledge inside network parameters is inefficient and smaller models will dominate test-time compute workloads.1:06:53–1:09:46 · The hosts pushing back 1/10 Case Study: The Arcee Foundation Model and Compounding Curation Gains Alessio asks for concrete metrics from the Arcee training run. Ari details how condensing 25T tokens into 7T allowed a 4.5B model to beat larger baselines before reaching one trillion tokens.1:09:46–1:14:24 · The hosts pushing back 1/10 Data Valuation: AI's NP-Complete Problem and the Datology Profile Swyx asks recruiting and technical frontier questions. Ari defines target-specific data valuation as AI's central NP-complete problem and outlines Datology's mission to automate zero-shot domain adaptation.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 29.8% · guest 70.2%0:00 · the hosts 29.8% · guest 70.2%3:00 · the hosts 4.7% · guest 95.3%3:00 · the hosts 4.7% · guest 95.3%6:00 · the hosts 0.5% · guest 99.5%6:00 · the hosts 0.5% · guest 99.5%9:00 · the hosts 10.2% · guest 89.8%9:00 · the hosts 10.2% · guest 89.8%12:00 · the hosts 7.1% · guest 92.9%12:00 · the hosts 7.1% · guest 92.9%15:00 · the hosts 22.5% · guest 77.5%15:00 · the hosts 22.5% · guest 77.5%18:00 · the hosts 1.5% · guest 98.5%18:00 · the hosts 1.5% · guest 98.5%21:00 · the hosts 27.4% · guest 72.6%21:00 · the hosts 27.4% · guest 72.6%24:00 · the hosts 33.2% · guest 66.8%24:00 · the hosts 33.2% · guest 66.8%27:00 · the hosts 7.4% · guest 92.6%27:00 · the hosts 7.4% · guest 92.6%30:00 · the hosts 2.7% · guest 97.3%30:00 · the hosts 2.7% · guest 97.3%33:00 · the hosts 19.7% · guest 80.3%33:00 · the hosts 19.7% · guest 80.3%36:00 · the hosts 15.6% · guest 84.4%36:00 · the hosts 15.6% · guest 84.4%39:00 · the hosts 17.9% · guest 82.1%39:00 · the hosts 17.9% · guest 82.1%42:00 · the hosts 7.9% · guest 92.1%42:00 · the hosts 7.9% · guest 92.1%45:00 · the hosts 13.8% · guest 86.2%45:00 · the hosts 13.8% · guest 86.2%48:00 · the hosts 12.7% · guest 87.3%48:00 · the hosts 12.7% · guest 87.3%51:00 · the hosts 6.9% · guest 93.1%51:00 · the hosts 6.9% · guest 93.1%54:00 · the hosts 17.1% · guest 82.9%54:00 · the hosts 17.1% · guest 82.9%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 17.8% · guest 82.2%1:00:00 · the hosts 17.8% · guest 82.2%1:03:00 · the hosts 26.3% · guest 73.7%1:03:00 · the hosts 26.3% · guest 73.7%1:06:00 · the hosts 16.2% · guest 83.8%1:06:00 · the hosts 16.2% · guest 83.8%1:09:00 · the hosts 8.4% · guest 91.6%1:09:00 · the hosts 8.4% · guest 91.6%1:12:00 · the hosts 19.8% · guest 80.2%1:12:00 · the hosts 19.8% · guest 80.2%1:15:00 · the hosts 15.1% · guest 84.9%1:15:00 · the hosts 15.1% · guest 84.9%1:18:00 · the hosts 38% · guest 62%1:18:00 · the hosts 38% · guest 62%
Sharpest disagreement ▶ 7:55 Calling Out Flawed Scaling Law Assumptions

Ari rejects standard scaling law literature from Kaplan and Chinchilla, calling the underlying assumption of IID data insane.

Hardest push from the hosts ▶ 52:56 Swyx Challenges Pre-Training Feedback Loop

Swyx directly presses Ari on how pre-training data can be optimized for downstream tuning if post-training is merely capability elicitation.

Biggest teaching moment ▶ 17:36 Human Intuition Failure in DCLM Study

Ari educates the hosts on the DCLM study showing expert NLP researchers performed no better than chance at guessing data filtering decisions.

The host holds their own ▶ 1:00:45 Swyx on Lottery Ticket Pruning and Hardware Bottlenecks

Swyx brings up specific historical research on parameter pruning and Frankle's lottery ticket hypothesis, demonstrating detailed technical context.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Ari Morcos's Background: Neuroscience, Inductive Biases, and the Bitter Lesson 4621 Swyx demonstrates technical familiarity with vision transformers and weight mapping between CNNs and ViTs. Ari walks through his transition from neuroscience to inductive biases and his realization that architecture matters far less than data scale.
The Strategic Shift: Why Data Is AI's Most Under-Invested Frontier 3641 Ari strongly criticizes foundational scaling papers by Kaplan and Chinchilla for assuming IID data distributions, calling the premise insane given standard computer science adages. The hosts listen attentively with minimal interjections.
Historical Prejudices and Misaligned Incentives in Data Research 5532 Alessio asks why data research incentives are misaligned given high-profile data companies like Scale AI. Ari differentiates academic research culture, which treated data as fixed grunt work, from industrial priorities and supervised dataset limitations.
The Self-Supervised Paradigm Shift: Moving to the Underfitting Regime 4631 Ari presents a contrarian view on transformer architectures, arguing they are just one of many viable architectures while self-supervised learning on unlabeled data was the true revolutionary breakthrough.
Data Engineering at Frontier Labs and the Failure of Human Intuition 5742 Swyx notes his belief in automated end-to-end learning. Ari backs this up with findings from DCLM showing domain-expert grad students could not predict data filtering classifiers better than chance, proving human intuition fails at scale.
Concept Complexity, Data Redundancy, and Zero-Shot Generalization 4623 Alessio presses Ari on how to empirically define boundaries for concepts. Ari uses an analogy comparing elephants and dog breeds to explain why redundancy requirements vary with semantic complexity.
The Evolution of Public Datasets, Code Predictors, and Copyright Litigation 6522 Swyx rapid-fires standard pre-training datasets including GitHub, Arxiv, and Books3, citing litigation against Anthropic and Meta. Ari notes counterintuitive findings from StarCoder showing GitHub stars do not predict code quality.
Overcoming the Limits of Power-Law Scaling via Data Efficiency 4632 Ari explains the mathematical connection between decaying marginal information gain and power-law scaling, citing his NeurIPS paper on bending scaling curves through active data pruning.
Quantifying Data Curation Gains Across Speed, Quality, and Size 6635 Swyx challenges whether Datology's performance claims are simply overfitting or training to the test on public benchmarks. Ari details their strict protocol using held-out evaluations and unreleased eval suites.
Why Specialized Data Companies Outperform Internal AI Lab Teams 5423 Alessio questions whether private curation layers create friction with open-source dataset creators. Ari details Datology's moat balance between open scientific intuition and proprietary engineering know-how.
Curation Methodologies and the Power of Synthetic Data Rephrasing 6522 Swyx explores synthetic data, drawing distinctions between distillation and steganography. Ari outlines Datology's core approach of rephrasing existing data to bypass mode collapse and teacher model limitations.
The Supremacy of Data Diversity and the Reality of Epoching 5532 Swyx brings up the 'Textbooks Are All You Need' hypothesis. Ari refutes narrow distribution claims, asserting data diversity is paramount and epoching high-quality tokens consistently outperforms low-grade new data.
The Resurgence of Curriculum Learning in Modern LLM Training 4622 Ari explains the conceptual graph theory behind curriculum learning, clarifying that discrete curricula now succeed in underfitted LLM regimes where they previously showed minimal value in saturated supervised settings.
Optimizing Pre-Training Data for Downstream Post-Training and Alignment 6645 Swyx challenges the idea of pre-training dependencies given the consensus that post-training is merely capability elicitation. Ari counters that pre-training determines the slope of test-time compute scaling and alignment robustness.
Model Pruning Bottlenecks and Complementary Optimization Strategies 7422 Swyx brings deep context regarding Jonathan Frankle's lottery ticket hypothesis and parameter pruning. Ari explains how unstructured pruning fell out of favor due to GPU sparse matrix multiply inefficiencies.
Future Architecture: Small Models, Test-Time Compute, and Cognitive Cores 6522 Swyx references Andre Karpathy's 'cognitive core' concept. Ari agrees that storing raw knowledge inside network parameters is inefficient and smaller models will dominate test-time compute workloads.
Case Study: The Arcee Foundation Model and Compounding Curation Gains 5511 Alessio asks for concrete metrics from the Arcee training run. Ari details how condensing 25T tokens into 7T allowed a 4.5B model to beat larger baselines before reaching one trillion tokens.
Data Valuation: AI's NP-Complete Problem and the Datology Profile 4521 Swyx asks recruiting and technical frontier questions. Ari defines target-specific data valuation as AI's central NP-complete problem and outlines Datology's mission to automate zero-shot domain adaptation.

Statements from this episode (50)

Insight
Morcos: Data curation choices fundamentally determine machine learning model performance
“There are a ton of choices you would make in that process, ranging from how you're going to filter the data, how you're going to sequence the data, what synthetic data you're going to generate, if any how you're going to batch the data, all of those things. An…”
Ari Morcos Aug 29, 2025 ▶ 0:58
Assertion Not checkable as stated
Morcos: Proper data curation enables smaller models with equal or better performance
“Help the folks we work with to train models much faster to much better performance and to also help them train much smaller models to the same or better performance, which I actually think is some of the most exciting stuff going forward. But fundamentally, th…”
Ari Morcos Aug 29, 2025 ▶ 1:43
Assertion Supported
Morcos: Soft inductive biases become harmful past 1M data points in vision
“Turns out in the small data regime, and when I say small data here, I mean, say less than 500,000 data points. And this was in the context of image self-supervised learning. So in that small data regime, this is super helpful. And where this paper's actually b…”
Ari Morcos Aug 29, 2025 ▶ 6:04
Insight
Morcos: Inductive biases matter not at all at scale compared to data
“Basically, you know, when you get to enough scale, inductive biases matter not at all. All that really matters is the learned posterior from the data distribution, and that's really what defines everything.”
Ari Morcos Aug 29, 2025 ▶ 6:51
Opinion
Morcos: Data is AI's most under-invested research area relative to impact
“Something I've said before and I'll say again is, is that data is the most under-invested in area of research relative to its impact, and I don't think it's even close.”
Ari Morcos Aug 29, 2025 ▶ 8:06
Assertion Supported
Morcos: Kaplan and Chinchilla scaling laws incorrectly assume all data is equal
“And even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data which is insane. We know that all data are not created equal, that garbage in garbage out is like the oldest adage in compu…”
Ari Morcos Aug 29, 2025 ▶ 8:25
Opinion
Morcos: Industry values data work far more than academic AI research
“In general data work has been far more valued in industry consistently than it had been in the research community.”
Ari Morcos Aug 29, 2025 ▶ 10:34
Insight
Morcos: Top AI researchers' secret to success is looking at the data
“If you talk to the most talented AI researchers and you ask them, what's the secret to your success, they'll largely tell you that they look at the data.”
Ari Morcos Aug 29, 2025 ▶ 11:07
Opinion
Morcos: The Transformer is just one of many equivalently good architectures
“And one of my like more controversial viewpoints, I think, is that I think the transformer is a great advance to be sure, but I think it's one of a very large Set of equivalently good architectures that we could have found. And there are many, many ways we cou…”
Ari Morcos Aug 29, 2025 ▶ 12:35
Opinion
Morcos: Modern AI capabilities depended entirely on self-supervised learning
“But I do not think there's any way we could get to where we are today without self-supervised learning and the ability to train on unlabeled data. That was the real advance to my mind that enabled us to get these incredible increases in capabilities.”
Ari Morcos Aug 29, 2025 ▶ 12:52
Insight
Morcos: Better data improves AI performance per dollar by orders of magnitude
“Making the data better can be a massive compute multiplier. It can change the performance per dollar by orders of magnitude.”
Ari Morcos Aug 29, 2025 ▶ 14:37
Insight
Morcos: Data curation requires compounding dozens of individually modest, conflicting techniques
“Data creation also is a hard problem to solve quote unquote, because it's not one where there's a single silver bullet. There's not just do this one trick and all of a sudden things work. It's rather here are these 50 different things that you can do, each of …”
Ari Morcos Aug 29, 2025 ▶ 16:10
Assertion Contradicted
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Ari Morcos Aug 29, 2025 ▶ 18:06
Insight
Morcos: A data point's value depends on its relationship to the full dataset
“The easiest way to think about this is that the value of a data point is not just a function of that data point itself. It's rather a function of how that data point relates to every other data point in the training set, right?”
Ari Morcos Aug 29, 2025 ▶ 18:47
Insight
Morcos: Complex, high-variance concepts require much more data redundancy than simple ones
“The amount of data that I need in order to properly understand dogs is going to be a lot higher than the amount of data I need to understand elephants.”
Ari Morcos Aug 29, 2025 ▶ 20:43
Insight
Morcos: GitHub stars do not predict code quality for model training
“Stars are not a good predictor of whether data is useful for models or not. Like, I think that's, like, the most popular repos are not necessarily higher quality, at least with respect to do they improve a model's coding capabilities.”
Ari Morcos Aug 29, 2025 ▶ 23:36
Disclosure
Morcos: Zuckerberg personally approved high-risk AI training datasets at Meta
“When I was at Meta, certainly legal stuff around data sets was very challenging and becoming increasingly challenging, and there are a number of situations where, you know, the only person that could approve things was Zuck because of the scale of the risk, I …”
Ari Morcos Aug 29, 2025 ▶ 25:45
Insight
Morcos: Power-law scaling yields diminishing returns for every 10x data increase
“Power law scaling is terrible. It means that every time you 10 X your data, you get a diminishing marginal return on performance.”
Ari Morcos Aug 29, 2025 ▶ 26:51
Opinion
Morcos: GPT-4.5 and Llama 4 show limits of naive mega-model scaling
“And I think that's what we've seen to some extent with the failure of the mega models, right? With 4.5 and Lama four and others. I think that there is a challenge of just continuing to do that naively and you have to figure out how to break it.”
Ari Morcos Aug 29, 2025 ▶ 27:07
Assertion Supported
Morcos: Proper data curation can bend neural scaling laws
“And what that paper showed was that if you use your data correctly, you can actually bend the scaling laws themselves.”
Ari Morcos Aug 29, 2025 ▶ 27:35
Assertion Supported
Morcos: Nemotron dataset quality is similar to DCLM despite token gains
“Nematron is actually pretty similar in quality to DCLM. It's, it came out about six months later. It has more unique tokens. They made a really big deal about it having more unique tokens, but on average, the quality is, is pretty straightforward.”
Ari Morcos Aug 29, 2025 ▶ 29:12
Assertion Open · timeframe Aug 2028
Morcos: Datology matches DCLM performance 12x faster with under 10% tokens
“We're able to now get to the same performance as DCLM about 12 x faster. So, you know, in fewer than 10% of the tokens we can match What you get from training to convergence.”
Ari Morcos Aug 29, 2025 ▶ 29:43
Prediction Not checkable as stated
Morcos: Data curation still has at least 100x in performance gains ahead
“You know, we've already been able to get 10 X gains. I think there's at least another hundred X behind this that are still to be done.”
Ari Morcos Aug 29, 2025 ▶ 32:23
Opinion
Morcos: Frontier AI lab data teams are systematically under-resourced
“I think you, what you see in all the frontier labs is that they have data teams. And if you talk to the folks that work on those data teams, what you'll kind of systematically hear is that typically they're under resourced relative to the gains that they're de…”
Ari Morcos Aug 29, 2025 ▶ 33:20
Disclosure
Morcos: Datology publishes intuition in blogs without enabling reproducibility
“What we've tried to do, and I think we've done a good job of, and I'm generally happy with the balance we've struck is try to, in the blog posts that we put out, give a lot of intuition as to kind of what we're doing and how it works without necessarily gettin…”
Ari Morcos Aug 29, 2025 ▶ 37:01
Assertion Partly supported
Morcos: Gemini tech report names data quality as single most important factor
“If you look at like the data section of like the Gemini tech report, it basically says like data quality was the single most important thing for making great model.”
Ari Morcos Aug 29, 2025 ▶ 37:22
Insight
Morcos: Post-training techniques are better applied in pre- and mid-training
“Most of what we do in post-training is better were done in pre and mid training and earlier on in training in general.”
Ari Morcos Aug 29, 2025 ▶ 42:16
Prediction Held up
Morcos: Training a specialized frontier model will cost under $1M very soon
“I believe that getting to a frontier model should cost a million dollars or less for most organizations, at least in a specialized domain, right? And when you think about what enterprises need, that's generally what they need. They don't need a model that can …”
Ari Morcos Aug 29, 2025 ▶ 42:49
Insight
Morcos: Filtering synthetic data between generation cycles prevents model collapse
“If you filter the data after each point, that's now information injection, and that can break all of this and I think can prevent model collapse.”
Ari Morcos Aug 29, 2025 ▶ 43:47
Insight
Morcos: Weak models can rephrase data to train superior models
“Because the model that's doing the rephrasing just needs to know how to rephrase. It doesn't need to know anything about the content itself. It doesn't need to understand it. It means you can use a pretty weak model. To do the rephrasing and have it generalize…”
Ari Morcos Aug 29, 2025 ▶ 44:23
Insight
Morcos: Data diversity is the single most important factor in AI data quality
“If there's only one thing that you should take away from this entire interview about what is good for data quality, it's diversity.”
Ari Morcos Aug 29, 2025 ▶ 46:16
Insight
Morcos: Epoching high-quality data beats training on new average data
“Repeating higher quality tokens is almost always better than seeing net new lower quality tokens. So, like, epoching over higher quality data, almost always better than getting the same amount of new data of an unknown quality, or of average quality, average i…”
Ari Morcos Aug 29, 2025 ▶ 48:15
Prediction Not checkable as stated
Morcos: Proper training curricula could reduce model training costs by 10x
“And getting a curriculum right could literally make the difference between, you know, spending 10 times as much on a model training, you know, hundreds of millions of dollars potentially.”
Ari Morcos Aug 29, 2025 ▶ 51:36
Assertion Not checkable as stated
Morcos: Major AI labs fail to holistically integrate pre-, mid-, and post-training
“Something that you don't see happen even at the big labs because they have entirely separate teams, right? There's a free training team. There's a mid training team. There's a post training team. And like the mid training team is a customer of the free trainin…”
Ari Morcos Aug 29, 2025 ▶ 52:35
Insight
Morcos: Post-training alignment is ineffective long-term compared to pre-training alignment
“Like fundamentally, I think alignment and post training doesn't really make sense as a long-term solution. If you can easily align a model through post training, you can easily misalign a model through post training. If it's easy to put it in, it's easy to tak…”
Ari Morcos Aug 29, 2025 ▶ 53:37
Assertion Not checkable as stated
Morcos: Qwen is much easier to align than Llama due to pre-training
“It's much easier to RL Quen than it is to do Lama. Likely that has to do with the fact that Quen put a lot of synthetic reasoning traces into their training data.”
Ari Morcos Aug 29, 2025 ▶ 54:07
Insight
Morcos: AI total cost of ownership is dominated by inference
“When you think about the total cost of ownership of these models, it's gonna be very, very heavily weighted towards inference. It's all inference.”
Ari Morcos Aug 29, 2025 ▶ 58:04
Prediction Not checkable as stated
Morcos: AI inference costs will skyrocket, penalizing oversized models
“The inference costs are going to skyrocket with these models. And if you use a general purpose model, then you constrain to say, hey, this model knows about everything, but now only do this one thing. That model is going to have a ton of parameters that do not…”
Ari Morcos Aug 29, 2025 ▶ 59:00
Opinion
Morcos: AI training is commoditized while data curation remains hard
“Mosaic was the first one to really recognize that there was a huge opportunity in making this easy. And now this has largely been commoditized by things like SageMaker and Together and lots of different folks that help you on the training side. But on the data…”
Ari Morcos Aug 29, 2025 ▶ 59:48
Insight
Morcos: Lottery ticket initializations fail because they are data dependent
“We actually found out that the problem was that the lottery ticket was actually data dependent. And that was where the fundamental problem came. That as soon as you change the data distribution a little bit, like the winning tickets changed in a really big way…”
Ari Morcos Aug 29, 2025 ▶ 1:01:14
Prediction Not checkable as stated
Morcos: Most AI models used in three years will be under 10B parameters
“Most of the models that the vast majority of people will be using in say three years will be single digit B or smaller.”
Ari Morcos Aug 29, 2025 ▶ 1:03:42
Insight
Morcos: Test-time compute fundamentally favors smaller models to cut multi-step inference costs
“Test time compute as a paradigm really pushes you towards smaller models, right? Because if your cost of solving a problem is cost of inference times number of thinking steps, and you have to do a lot of thinking steps. Well, now this is like a really like min…”
Ari Morcos Aug 29, 2025 ▶ 1:04:35
Opinion
Morcos: Current AI models waste massive capacity memorizing unnecessary knowledge
“We're wasting a ton of capacity in these models on knowledge that is just totally unnecessary for them to have.”
Ari Morcos Aug 29, 2025 ▶ 1:06:42
Assertion Not publicly verifiable
Morcos: Arcee 4.5B beat Gemma before reaching one trillion tokens
“It was beating Gemma pretty consistently before the one trillion mark, which was pretty cool to see.”
Ari Morcos Aug 29, 2025 ▶ 1:07:33
Insight
Morcos: Combining disparate data curation techniques generally fails without difficult tuning
“When you take these different techniques and you try to make them work together, they don't, generally. You can make them work together, but it's quite hard to do so.”
Ari Morcos Aug 29, 2025 ▶ 1:08:12
Insight
Morcos: Curation gains stack multiplicatively and preserve relative dataset advantages
“If we apply our curation on top of say DCLM, and then we apply it on top of FineWeb, the gap between FineWeb and DCLM is maintained in the gap between kind of Datology curated DCLM and Datology curated FineWeb. They both get a lot better, but Datology DCLM is …”
Ari Morcos Aug 29, 2025 ▶ 1:09:07
Insight
Morcos: No universal 'golden' curation exists for AI training data
“There's no golden curation. A curation is only optimal with respect to a given set of downstream use cases or tasks, right?”
Ari Morcos Aug 29, 2025 ▶ 1:12:22
Insight
Morcos: Valuing data for downstream use cases is AI's NP-complete problem
“In many ways, I think that's kind of the NP-complete problem of AI. If you can do that, you can kind of do anything”
Ari Morcos Aug 29, 2025 ▶ 1:13:29
Assertion Not checkable as stated
Morcos: Yann LeCun was never defining Meta's AI strategy
“I don't think he was ever you know, or at least not since the beginning in a role where he was defining AI strategy for Meta. I don't think that's the role he wanted at any point. You know, I think he really wanted to be doing that research, and I think, so I …”
Ari Morcos Aug 29, 2025 ▶ 1:16:13
Prediction Not checkable as stated
Morcos: Meta's metaverse bet will pay off in the long run
“I think the one that's still really up in the air is a metaverse, but I would actually argue that I think that's going to end up paying off in the long run. I think the Ray-Ban glasses pretty darn cool. And a lot of the foundations of what was in reality labs …”
Ari Morcos Aug 29, 2025 ▶ 1:17:20
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.