Insight certainty 4/5 debate potential 2/5

Speech-Based Image Annotation Produces Richer Training Data Faster Than Writing

Vibhu Sapra · [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz · Oct 13, 2024 · at 6:00

Vibhu Sapra explains AI2's data collection methodology for the Pixmo dataset used to train the Molmo VLM family.

0:00 / 0:25exact quote · 25.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“We ask annotators to describe the images in speech for 60 to 90 seconds, rather than asking them to write descriptions. They prompted them to describe everything in great detail, including descriptions of spatial positioning and relationships. So stuff like, you know, there's the space needle in the middle of a bunch of buildings. With that, the modality switching trick, annotators provide far more detailed descriptions in less time.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Vibhu Sapra

Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Molmo-72B Ranks Second Behind Only GPT-4o in Human Preference Elo
“Most preference ranked was GPT-Four-O, then Momo-Semety-Two-B, then Gemini, then Sonnet, then the Seven-B.”
Vibhu Sapra Oct 13, 2024 ▶ 39:17 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Molmo 72B Beats Proprietary Models on Academic Benchmarks
“Academic benchmarks wise, Their big one is the best state-of-the-art everything, better than proprietary, but ELO-wise, it sits behind four-oh.”
Vibhu Sapra Oct 13, 2024 ▶ 45:42 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Most Open Vision Models Rely on Synthetic Data From Proprietary Models
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”
Vibhu Sapra Oct 13, 2024 ▶ 1:13 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Insight
Distilling VLMs From Proprietary Models Copies Their Spatial Pointing Failures
“For example, all this proprietary stuff sucks at clocks, so nothing that's a distillation will be good at clocks. Nothing can point if you just distill from this. If they can't point, your VLM won't point, so we show how to get good data.”
Vibhu Sapra Oct 13, 2024 ▶ 32:22 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Partly supported
Molmo 1B Matches GPT-4V Across Academic Benchmarks and Elo
“The most efficient model, the one B is based on their one B MOE. That one matches performance of four V on most academic benchmarks and their ELO ranking.”
Vibhu Sapra Oct 13, 2024 ▶ 45:06 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.