Everything Ce Zhang said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Zhang: The world is not running out of AI training data
“I don't think we are running out of data on earth. Right, so think about it globally... But I do think there are many organizations in the world have enough data to actually train, like, very, very good models, right? So, I mean, they are not public available,…”
Zhang: ML systems optimization is fundamentally always about data movement
“Fundamentally, the thing we are actually optimizing is actually not that different. It's always about data movement across essentially all the stacks, right? So when you do distributed, like computing, it's about communication across different machines. When y…”
Zhang: Pre-training data evolved from a model byproduct into a standalone asset
“So, so I think one fundamental thing that changed in the last year, essentially, in the beginning when people think about data, is, is always like a byproduct of the model, right? You release the model, you also release the data, right? The data side is there …”
Zhang: Targeted task training yields smaller, cheaper, and more accurate models
“The benefit you can get out of that is you could build a, Better open model, often smaller, often easier to do inference if you know what you want, right? So I think the whole trade-off would be, and the x-axis would be how generic the hosting will be. The y-a…”
Zhang: Next 10x AI inference gain requires multi-layer co-optimization
“If you only push on one direction, you are going to reach diminution return really, really quickly. Yeah, there's only that much you can do on the system side, only that much you can do on the algorithm side. And since the only big thing that's going to happen…”
Zhang: Good benchmarks must anticipate how developers over-optimize toward metrics
“A good benchmark should think about how it's going to incentivize the field to actually move forward, right? So the benchmark will become kind of standard. How are people going to over optimize to the benchmark because people are going to do that? And when peo…”
Zhang predicts much faster, diverse new embedding models within couple years
“So I think for the next couple years, yeah, we will see a whole bunch of new embeddings maybe of different sites, and much, much faster than today.”
Zhang: Combining fine-tuning and RAG provides superior performance boosts
“Combining all those techniques all together, right? So we'll give you essentially another boost, right? So that kind of one thing that we learn on the technical side.”
Zhang: RedPajama-V2 features 40 pre-computed quality signals for custom filtering
“So that's why in REST-PYRON V-II, we kind of overlay the data set, it's like, 40 different pre-computed quality signal, right? If you want to reproduce your best effort, like, C-Four filter, it's kind of like, 20 lines of code.”