Text Data
topic on 4 shows · 5 statements across 5 episodes
A Product Market Fit Show
the MAD Podcast
Big Technology
20VC
5 statements about Text Data, every show
Izmailov: Text data carries more structural information per token than images
“So for example, we can approximate it from the scaling laws and we can, for example, say that text data has more structural information according to this measure than image data at the same kind of amount of yeah, tokens.”
Eskildsen: Vector embeddings are larger in data size than raw text
“Because they're in thousands of dimensions, the coordinate in the coordinate system is very counterintuitive, but the coordinate in the coordinate system is larger than the original data. If you take a paragraph of text, the coordinate that represents it is ac…”
LeCun: AI cannot reach human-level intelligence by training on text alone
“And it tells you clearly that we're not going to get to human level AI by just training on text. It's just not a rich enough source of information.”
Socher: Top AI models have run out of high-quality human text data
“I think a lot of the top models, you can't add multiple orders of magnitude more text to them, because there isn't that much more text around.”
Matt Clifford: Text-Based AI Scaling Is Hitting an S-Curve Limit
“I don't think there's a lot further to go in just finding more text. You know, I think we may be at the flattening out of the S-curve on text, but I suspect the next S-curve could just involve finding ways to yeah, to use video, to use, like, interactive exper…”