Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Most Open Vision Models Rely on Synthetic Data From Proprietary Models

Vibhu Sapra · [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz · Oct 13, 2024 · at 1:13

Vibhu Sapra discusses the training methodology of open-weight Vision-Language Models (VLMs) during a technical review of AI2's Molmo paper.

0:00 / 0:12exact quote · 12.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Vibhu Sapra

Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Insight
Speech-Based Image Annotation Produces Richer Training Data Faster Than Writing
“We ask annotators to describe the images in speech for 60 to 90 seconds, rather than asking them to write descriptions. They prompted them to describe everything in great detail, including descriptions of spatial positioning and relationships. So stuff like, y…”
Vibhu Sapra Oct 13, 2024 ▶ 6:00 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Molmo-72B Ranks Second Behind Only GPT-4o in Human Preference Elo
“Most preference ranked was GPT-Four-O, then Momo-Semety-Two-B, then Gemini, then Sonnet, then the Seven-B.”
Vibhu Sapra Oct 13, 2024 ▶ 39:17 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Molmo 72B Beats Proprietary Models on Academic Benchmarks
“Academic benchmarks wise, Their big one is the best state-of-the-art everything, better than proprietary, but ELO-wise, it sits behind four-oh.”
Vibhu Sapra Oct 13, 2024 ▶ 45:42 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Insight
Distilling VLMs From Proprietary Models Copies Their Spatial Pointing Failures
“For example, all this proprietary stuff sucks at clocks, so nothing that's a distillation will be good at clocks. Nothing can point if you just distill from this. If they can't point, your VLM won't point, so we show how to get good data.”
Vibhu Sapra Oct 13, 2024 ▶ 32:22 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Partly supported
Molmo 1B Matches GPT-4V Across Academic Benchmarks and Elo
“The most efficient model, the one B is based on their one B MOE. That one matches performance of four V on most academic benchmarks and their ELO ranking.”
Vibhu Sapra Oct 13, 2024 ▶ 45:06 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.