Datology, every mention
9 scenes · ← back to Datology
tap a year for its mentions
every year anyone Shawn Wang 15Ari Morcos 9Pratyush Maini 4
Verbatim, from the transcripts: the passages where Datology comes up
⚡️ Reverse Engineering OpenAI's Training Data — Pratyush Maini, Datology
- ▶ 0:06 unnamed speaker Uh, we have Patush here from Datology, one of the founding team members. 6 times in the scene
- ▶ 19:47 Pratyush Maini And so that's the core thesis that we've been working towards in the past few months, and also very relevant to datology in general. 3 times in the scene
- ▶ 21:22 Pratyush Maini So, the NemoTron dataset is, like, one of the top datasets today, which is, like, based with a lot of synthetic data, uh, so the, and the Datology, uh, data that we, a model that we released called BeyondWeb, uh, is the blue line over here.
- ▶ 26:24 unnamed speaker and then it's like Microsoft or whatever, and it's Apple, then it's Hugging Face, then it's NVIDIA, now it's you guys, and I'm like, you know, where, where's the, the, the sort of persistence, or like, is this such a competitive field? 3 times in the scene
Better Data is All You Need — Ari Morcos, Datology
- ▶ 0:12 Shawn Wang And we're so excited to be in the studio with Ari Marcos, CEO, co-founder of Datology. 6 times in the scene
- ▶ 26:31 Ari Morcos Like, this is actually a lot of the reason why I got into data and started Datology was that the scaling laws always were terrible. 3 times in the scene
- ▶ 35:42 Ari Morcos And it's something that we thought a lot about at Datology. 6 times in the scene
- ▶ 39:20 Shawn Wang Um, you know, I, I, I figured that most of the work of Datology is filtering, but I see synthetic data as something slightly different. 2 times in the scene
- ▶ 1:10:55 Shawn Wang Uh, what data efficiency question, if somebody had an answer, they should join Datology immediately. 7 times in the scene