Q&A Data Drives VLM Detail Recognition Better Than Captioning
Eugene Xia · [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz · Oct 13, 2024 · at 43:41
Eugene Xia discusses dataset curation strategies for vision-language models during a technical Paper Club discussion on AI2's Molmo and Pixmo.
“My intuition at least is that I suspect it's really, the heavy lifting might be actually more towards the Q and A side rather than captioning side. I view captioning as a bootstrap, but at the end of the day, we want the model to generalize what they want to describe. Captioning tends to overfeed what is being described, where our question and answers kind of forces the model to pick up details that It would previously ignore like plots, for example, or the one extra finger on the hand or whatever it is. And as long as you're able to ask questions and you can answer, you're training the model to capture as much information as possible, even if you don't ask. Because even if, let's say, this data set, you didn't ask this question, another data set may ask an alternative question instead, and that will just generalize better. That's what I suspect. I think captioning is a problem in overfitting from my point of view.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →