DCLM
4 statements across 1 episodes · 2 bullish · 0 bearish · 1 people on the record · first statement Aug 29, 2025 by Ari Morcos · said 2 times in 1 episodes since 2024 · across every show →
Mentions by year
brought up most by Loubna Ben Allal (2)
tap a year for its mentions
Everything said about DCLM, oldest first
Aug 29, 2025
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Aug 29, 2025 positive
Morcos: Curation gains stack multiplicatively and preserve relative dataset advantages
“If we apply our curation on top of say DCLM, and then we apply it on top of FineWeb, the gap between FineWeb and DCLM is maintained in the gap between kind of Datology curated DCLM and Datology curated FineWeb. They both get a lot better, but Datology DCLM is …”
Aug 29, 2025 neutral
Morcos: Nemotron dataset quality is similar to DCLM despite token gains
“Nematron is actually pretty similar in quality to DCLM. It's, it came out about six months later. It has more unique tokens. They made a really big deal about it having more unique tokens, but on average, the quality is, is pretty straightforward.”