Zhang: Pre-training data evolved from a model byproduct into a standalone asset
Ce Zhang · Building an open AI company - with Ce and Vipul of Together AI · Feb 8, 2024 · at 14:43
Ce Zhang, CTO of Together AI, describes the industry-wide mindset shift around open-source AI datasets and modular filtering.
“So, so I think one fundamental thing that changed in the last year, essentially, in the beginning when people think about data, is, is always like a byproduct of the model, right? You release the model, you also release the data, right? The data side is there for you to, essentially, to show people, ah, if you try on this data, you get a good model. But I think what starts to change is when people start building more than one of those models, people start to realize, like, different subsets or data set is kind of valuable for different applications, right? The data becomes something you want to play with, right?”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →