Vibhu Sapra

2 appearances on the record.

computed by AI from the episodes · how this works → · full disclaimer →

8statements → 6claims → 6claims resolved → 83%fully supported → 4/5average certainty → 2/5average debate potential → 6said about them ↓

5 supported 1 partly supported 0 contradicted how the 6 claims stand · each chip opens the sources

6 assertions · 2 insights · every statement was checked. The predictions and assertions are the 6 claims: statements the public record can support or contradict. 6 are resolved. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Vibhu argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz

Everything Vibhu Sapra said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Insight
Speech-Based Image Annotation Produces Richer Training Data Faster Than Writing
“We ask annotators to describe the images in speech for 60 to 90 seconds, rather than asking them to write descriptions. They prompted them to describe everything in great detail, including descriptions of spatial positioning and relationships. So stuff like, y…”
Vibhu Sapra Oct 13, 2024 ▶ 6:00 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Molmo-72B Ranks Second Behind Only GPT-4o in Human Preference Elo
“Most preference ranked was GPT-Four-O, then Momo-Semety-Two-B, then Gemini, then Sonnet, then the Seven-B.”
Vibhu Sapra Oct 13, 2024 ▶ 39:17 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Molmo 72B Beats Proprietary Models on Academic Benchmarks
“Academic benchmarks wise, Their big one is the best state-of-the-art everything, better than proprietary, but ELO-wise, it sits behind four-oh.”
Vibhu Sapra Oct 13, 2024 ▶ 45:42 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Most Open Vision Models Rely on Synthetic Data From Proprietary Models
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”
Vibhu Sapra Oct 13, 2024 ▶ 1:13 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Insight
Distilling VLMs From Proprietary Models Copies Their Spatial Pointing Failures
“For example, all this proprietary stuff sucks at clocks, so nothing that's a distillation will be good at clocks. Nothing can point if you just distill from this. If they can't point, your VLM won't point, so we show how to get good data.”
Vibhu Sapra Oct 13, 2024 ▶ 32:22 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Partly supported
Molmo 1B Matches GPT-4V Across Academic Benchmarks and Elo
“The most efficient model, the one B is based on their one B MOE. That one matches performance of four V on most academic benchmarks and their ELO ranking.”
Vibhu Sapra Oct 13, 2024 ▶ 45:06 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Assertion Supported
Pixmo Dataset Contains Approximately 1M Captions Across 700K Images
“They got about a million captions for 700,000 images, which I'm like, okay, that's kind of expensive.”
Vibhu Sapra Oct 13, 2024 ▶ 22:17 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz

The other half of the tape: Vibhu Sapra's own voice is left out of every number here. Other people bring the name up 6 times in 1 episode on Latent Space. every mention, with the transcript →

Who brings them up most Eugene Yan 1Eugene Cheah 1

Every mention by year

tap a year for its mentions
0031612024episodesmentions
0112024episodes it came up in
0030.5612024episodesmentions per episode
2024 6 mentions in 1 episode

Appearances (2)

EpisodeDateSpeaking time
[AIEWF Preview] CloudChef: Your Robot Chef - Michellin-Star food at $12/hr (w/ Kitchen tou May 31, 2025 46s
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz Oct 13, 2024 33m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.