Datology
company on 1 show · 10 statements across 2 episodes · said 37 times in 2 episodes since 2025
Mentions by year, every show
tap a year for its mentions
Latent Space 37
2026 13 mentions in 1 episode
2025 24 mentions in 1 episode
every mention on every show, scene by scene, with the transcript →
10 statements about Datology, every show
Datology BeyondWeb 3B matches NVIDIA Nemotron 8B in 2.7x less training time
“As you can see that we achieved the same performance as the NVIDIA model in almost, like, 2.7 X, like, lesser time. And then much faster than anything that hugging face or pajama does. Very interestingly, our three B model is pretty much the same performance a…”
Morcos: No universal 'golden' curation exists for AI training data
“There's no golden curation. A curation is only optimal with respect to a given set of downstream use cases or tasks, right?”
Morcos: Curation gains stack multiplicatively and preserve relative dataset advantages
“If we apply our curation on top of say DCLM, and then we apply it on top of FineWeb, the gap between FineWeb and DCLM is maintained in the gap between kind of Datology curated DCLM and Datology curated FineWeb. They both get a lot better, but Datology DCLM is …”
Morcos: Arcee 4.5B beat Gemma before reaching one trillion tokens
“It was beating Gemma pretty consistently before the one trillion mark, which was pretty cool to see.”
Morcos: Datology publishes intuition in blogs without enabling reproducibility
“What we've tried to do, and I think we've done a good job of, and I'm generally happy with the balance we've struck is try to, in the blog posts that we put out, give a lot of intuition as to kind of what we're doing and how it works without necessarily gettin…”
Morcos: Datology matches DCLM performance 12x faster with under 10% tokens
“We're able to now get to the same performance as DCLM about 12 x faster. So, you know, in fewer than 10% of the tokens we can match What you get from training to convergence.”
Morcos: Complex, high-variance concepts require much more data redundancy than simple ones
“The amount of data that I need in order to properly understand dogs is going to be a lot higher than the amount of data I need to understand elephants.”
Morcos: Data curation requires compounding dozens of individually modest, conflicting techniques
“Data creation also is a hard problem to solve quote unquote, because it's not one where there's a single silver bullet. There's not just do this one trick and all of a sudden things work. It's rather here are these 50 different things that you can do, each of …”
Morcos: Proper data curation enables smaller models with equal or better performance
“Help the folks we work with to train models much faster to much better performance and to also help them train much smaller models to the same or better performance, which I actually think is some of the most exciting stuff going forward. But fundamentally, th…”
Morcos: Data curation choices fundamentally determine machine learning model performance
“There are a ton of choices you would make in that process, ranging from how you're going to filter the data, how you're going to sequence the data, what synthetic data you're going to generate, if any how you're going to batch the data, all of those things. An…”