Oct 13, 2024 · 1h 12m · latent-space

[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz

Vibhu Sapra · 33m spoken Nathan Lambert · 4m spoken Eugene Xia · 2m spoken Demo User 2 · 4s spoken Demo User 1 · 2s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this Paper Club session, speakers Vibhu Sapra, Nathan Lambert, and Amgadoz analyze the architecture and data curation pipeline behind AI2's open vision-language model family Molmo and Pixmo, followed by a technical deep-dive into the pruning and performance of OpenAI's Whisper Large v3 Turbo.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.1 Guest teaching 5.1 Guest disagreement 1.2 The hosts pushing back 0.1
05100:0015:0030:0045:001:00:0011:46–16:52 · The hosts as informed peer 1/10 Nathan Lambert on Vision Space Data Strategy Nathan Lambert demystifies the hype around Molmo, stating the vision space is heavily underdeveloped and proprietary frontier labs will simply replicate and absorb the open fine-tuning data. He also reveals internal AI2 timeline pressures and model design trade-offs regarding chat templates.16:52–22:12 · The hosts as informed peer 0/10 Pixmo Speech-Based Data Collection Pipeline Vibhu outlines Pixmo's speech-based data generation pipeline using Whisper transcription and LLM normalization. Nathan briefly corrects Vibhu regarding the distinction between MMLU and the multimodal benchmark MMMU.22:12–30:33 · The hosts as informed peer 0/10 Pixmo Subsets: Pointing, Clocks, and QA Generation Eugene and Vibhu explore the Pixmo QA generation and clock datasets. Nathan explains the pragmatic behind-the-scenes reality: the lead researcher became fixated on fixing clock reading, yet the resulting model still fails to generalize to simple dials.30:33–40:34 · The hosts as informed peer 0/10 Molmo Benchmark Performance and Human ELO Evals The panel assesses Molmo's benchmark performance and human ELO evals against proprietary models. Nathan discusses how using OpenAI's CLIP vision encoder stretches the formal definition of open-source AI and why Molmo skipped RLHF.40:33–44:44 · The hosts as informed peer 0/10 Community Q&A on Caption Granularity and Generalization Audience member Sam asks about required image caption granularity. Nathan acknowledges having no definitive answer, while Eugene offers the hypothesis that diverse QA pairs matter far more than granular captions to prevent overfitting.44:48–50:22 · The hosts as informed peer 0/10 Molmo Video Demonstration and Paper Wrap-Up Vibhu wraps up the Molmo segment by showcasing demo video clips of pointing and JSON extraction before handing over the presentation to Amgadoz for the Whisper Turbo segment.50:22–57:54 · The hosts as informed peer 0/10 Whisper Large v3 Turbo Architecture and Pruning Amgadoz explains the architectural optimization of Whisper Large v3 Turbo, dispelling the misconception that it uses distillation. He details how pruning the decoder from 32 layers to 4 targets the primary 90% latency bottleneck.57:54–1:03:23 · The hosts as informed peer 0/10 Whisper Turbo vs. Distil-Whisper and Distillation Amgadoz contrasts Whisper Turbo's multi-million hour multilingual pruning dataset with Distil-Whisper's English-only distillation. He breaks down the mechanics of KL divergence distribution matching versus standard cross-entropy loss.1:03:24–1:08:54 · The hosts as informed peer 0/10 Real-Time Streaming Audio Decoding Mechanics Amgadoz walks through the engineering details of low-latency streaming transcription, explaining how Whisper's mandatory 30-second fixed input window requires padding and small-step buffer shifts to minimize decoder execution time.11:46–16:52 · Guest teaching 6/10 Nathan Lambert on Vision Space Data Strategy Nathan Lambert demystifies the hype around Molmo, stating the vision space is heavily underdeveloped and proprietary frontier labs will simply replicate and absorb the open fine-tuning data. He also reveals internal AI2 timeline pressures and model design trade-offs regarding chat templates.16:52–22:12 · Guest teaching 4/10 Pixmo Speech-Based Data Collection Pipeline Vibhu outlines Pixmo's speech-based data generation pipeline using Whisper transcription and LLM normalization. Nathan briefly corrects Vibhu regarding the distinction between MMLU and the multimodal benchmark MMMU.22:12–30:33 · Guest teaching 6/10 Pixmo Subsets: Pointing, Clocks, and QA Generation Eugene and Vibhu explore the Pixmo QA generation and clock datasets. Nathan explains the pragmatic behind-the-scenes reality: the lead researcher became fixated on fixing clock reading, yet the resulting model still fails to generalize to simple dials.30:33–40:34 · Guest teaching 5/10 Molmo Benchmark Performance and Human ELO Evals The panel assesses Molmo's benchmark performance and human ELO evals against proprietary models. Nathan discusses how using OpenAI's CLIP vision encoder stretches the formal definition of open-source AI and why Molmo skipped RLHF.40:33–44:44 · Guest teaching 3/10 Community Q&A on Caption Granularity and Generalization Audience member Sam asks about required image caption granularity. Nathan acknowledges having no definitive answer, while Eugene offers the hypothesis that diverse QA pairs matter far more than granular captions to prevent overfitting.44:48–50:22 · Guest teaching 1/10 Molmo Video Demonstration and Paper Wrap-Up Vibhu wraps up the Molmo segment by showcasing demo video clips of pointing and JSON extraction before handing over the presentation to Amgadoz for the Whisper Turbo segment.50:22–57:54 · Guest teaching 7/10 Whisper Large v3 Turbo Architecture and Pruning Amgadoz explains the architectural optimization of Whisper Large v3 Turbo, dispelling the misconception that it uses distillation. He details how pruning the decoder from 32 layers to 4 targets the primary 90% latency bottleneck.57:54–1:03:23 · Guest teaching 7/10 Whisper Turbo vs. Distil-Whisper and Distillation Amgadoz contrasts Whisper Turbo's multi-million hour multilingual pruning dataset with Distil-Whisper's English-only distillation. He breaks down the mechanics of KL divergence distribution matching versus standard cross-entropy loss.1:03:24–1:08:54 · Guest teaching 7/10 Real-Time Streaming Audio Decoding Mechanics Amgadoz walks through the engineering details of low-latency streaming transcription, explaining how Whisper's mandatory 30-second fixed input window requires padding and small-step buffer shifts to minimize decoder execution time.11:46–16:52 · Guest disagreement 3/10 Nathan Lambert on Vision Space Data Strategy Nathan Lambert demystifies the hype around Molmo, stating the vision space is heavily underdeveloped and proprietary frontier labs will simply replicate and absorb the open fine-tuning data. He also reveals internal AI2 timeline pressures and model design trade-offs regarding chat templates.16:52–22:12 · Guest disagreement 1/10 Pixmo Speech-Based Data Collection Pipeline Vibhu outlines Pixmo's speech-based data generation pipeline using Whisper transcription and LLM normalization. Nathan briefly corrects Vibhu regarding the distinction between MMLU and the multimodal benchmark MMMU.22:12–30:33 · Guest disagreement 2/10 Pixmo Subsets: Pointing, Clocks, and QA Generation Eugene and Vibhu explore the Pixmo QA generation and clock datasets. Nathan explains the pragmatic behind-the-scenes reality: the lead researcher became fixated on fixing clock reading, yet the resulting model still fails to generalize to simple dials.30:33–40:34 · Guest disagreement 2/10 Molmo Benchmark Performance and Human ELO Evals The panel assesses Molmo's benchmark performance and human ELO evals against proprietary models. Nathan discusses how using OpenAI's CLIP vision encoder stretches the formal definition of open-source AI and why Molmo skipped RLHF.40:33–44:44 · Guest disagreement 1/10 Community Q&A on Caption Granularity and Generalization Audience member Sam asks about required image caption granularity. Nathan acknowledges having no definitive answer, while Eugene offers the hypothesis that diverse QA pairs matter far more than granular captions to prevent overfitting.44:48–50:22 · Guest disagreement 0/10 Molmo Video Demonstration and Paper Wrap-Up Vibhu wraps up the Molmo segment by showcasing demo video clips of pointing and JSON extraction before handing over the presentation to Amgadoz for the Whisper Turbo segment.50:22–57:54 · Guest disagreement 1/10 Whisper Large v3 Turbo Architecture and Pruning Amgadoz explains the architectural optimization of Whisper Large v3 Turbo, dispelling the misconception that it uses distillation. He details how pruning the decoder from 32 layers to 4 targets the primary 90% latency bottleneck.57:54–1:03:23 · Guest disagreement 1/10 Whisper Turbo vs. Distil-Whisper and Distillation Amgadoz contrasts Whisper Turbo's multi-million hour multilingual pruning dataset with Distil-Whisper's English-only distillation. He breaks down the mechanics of KL divergence distribution matching versus standard cross-entropy loss.1:03:24–1:08:54 · Guest disagreement 0/10 Real-Time Streaming Audio Decoding Mechanics Amgadoz walks through the engineering details of low-latency streaming transcription, explaining how Whisper's mandatory 30-second fixed input window requires padding and small-step buffer shifts to minimize decoder execution time.11:46–16:52 · The hosts pushing back 1/10 Nathan Lambert on Vision Space Data Strategy Nathan Lambert demystifies the hype around Molmo, stating the vision space is heavily underdeveloped and proprietary frontier labs will simply replicate and absorb the open fine-tuning data. He also reveals internal AI2 timeline pressures and model design trade-offs regarding chat templates.16:52–22:12 · The hosts pushing back 0/10 Pixmo Speech-Based Data Collection Pipeline Vibhu outlines Pixmo's speech-based data generation pipeline using Whisper transcription and LLM normalization. Nathan briefly corrects Vibhu regarding the distinction between MMLU and the multimodal benchmark MMMU.22:12–30:33 · The hosts pushing back 0/10 Pixmo Subsets: Pointing, Clocks, and QA Generation Eugene and Vibhu explore the Pixmo QA generation and clock datasets. Nathan explains the pragmatic behind-the-scenes reality: the lead researcher became fixated on fixing clock reading, yet the resulting model still fails to generalize to simple dials.30:33–40:34 · The hosts pushing back 0/10 Molmo Benchmark Performance and Human ELO Evals The panel assesses Molmo's benchmark performance and human ELO evals against proprietary models. Nathan discusses how using OpenAI's CLIP vision encoder stretches the formal definition of open-source AI and why Molmo skipped RLHF.40:33–44:44 · The hosts pushing back 0/10 Community Q&A on Caption Granularity and Generalization Audience member Sam asks about required image caption granularity. Nathan acknowledges having no definitive answer, while Eugene offers the hypothesis that diverse QA pairs matter far more than granular captions to prevent overfitting.44:48–50:22 · The hosts pushing back 0/10 Molmo Video Demonstration and Paper Wrap-Up Vibhu wraps up the Molmo segment by showcasing demo video clips of pointing and JSON extraction before handing over the presentation to Amgadoz for the Whisper Turbo segment.50:22–57:54 · The hosts pushing back 0/10 Whisper Large v3 Turbo Architecture and Pruning Amgadoz explains the architectural optimization of Whisper Large v3 Turbo, dispelling the misconception that it uses distillation. He details how pruning the decoder from 32 layers to 4 targets the primary 90% latency bottleneck.57:54–1:03:23 · The hosts pushing back 0/10 Whisper Turbo vs. Distil-Whisper and Distillation Amgadoz contrasts Whisper Turbo's multi-million hour multilingual pruning dataset with Distil-Whisper's English-only distillation. He breaks down the mechanics of KL divergence distribution matching versus standard cross-entropy loss.1:03:24–1:08:54 · The hosts pushing back 0/10 Real-Time Streaming Audio Decoding Mechanics Amgadoz walks through the engineering details of low-latency streaming transcription, explaining how Whisper's mandatory 30-second fixed input window requires padding and small-step buffer shifts to minimize decoder execution time.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 12:00 Nathan demystifies vision model breakthroughs

Nathan downplays the novelty of recent vision-language advances, arguing the field is simply underdeveloped and that frontier labs will easily ingest their fine-tuning methodology.

Hardest push from the hosts ▶ 29:04 Vibhu questions benchmark overfitting on clocks

Vibhu pushes back on the inclusion of synthetic clock datasets, questioning whether it represents genuine capability generalization or merely overfitting to a narrow benchmark.

Biggest teaching moment ▶ 1:05:01 Amgadoz explains Whisper 30s padding requirement

Amgadoz educates the group on Whisper's architectural constraint requiring 30 seconds of audio padding even for 300 millisecond streaming chunks.

The host holds their own ▶ 25:21 Vibhu clarifies synthetic QA pipeline without vision LLM

Vibhu clearly explains the synthetic bootstrapping mechanism where text-only LLMs generate QA pairs from dense OCR and speech captions to train multimodal models.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Nathan Lambert on Vision Space Data Strategy 1631 Nathan Lambert demystifies the hype around Molmo, stating the vision space is heavily underdeveloped and proprietary frontier labs will simply replicate and absorb the open fine-tuning data. He also reveals internal AI2 timeline pressures and model design trade-offs regarding chat templates.
Pixmo Speech-Based Data Collection Pipeline 0410 Vibhu outlines Pixmo's speech-based data generation pipeline using Whisper transcription and LLM normalization. Nathan briefly corrects Vibhu regarding the distinction between MMLU and the multimodal benchmark MMMU.
Pixmo Subsets: Pointing, Clocks, and QA Generation 0620 Eugene and Vibhu explore the Pixmo QA generation and clock datasets. Nathan explains the pragmatic behind-the-scenes reality: the lead researcher became fixated on fixing clock reading, yet the resulting model still fails to generalize to simple dials.
Molmo Benchmark Performance and Human ELO Evals 0520 The panel assesses Molmo's benchmark performance and human ELO evals against proprietary models. Nathan discusses how using OpenAI's CLIP vision encoder stretches the formal definition of open-source AI and why Molmo skipped RLHF.
Community Q&A on Caption Granularity and Generalization 0310 Audience member Sam asks about required image caption granularity. Nathan acknowledges having no definitive answer, while Eugene offers the hypothesis that diverse QA pairs matter far more than granular captions to prevent overfitting.
Molmo Video Demonstration and Paper Wrap-Up 0100 Vibhu wraps up the Molmo segment by showcasing demo video clips of pointing and JSON extraction before handing over the presentation to Amgadoz for the Whisper Turbo segment.
Whisper Large v3 Turbo Architecture and Pruning 0710 Amgadoz explains the architectural optimization of Whisper Large v3 Turbo, dispelling the misconception that it uses distillation. He details how pruning the decoder from 32 layers to 4 targets the primary 90% latency bottleneck.
Whisper Turbo vs. Distil-Whisper and Distillation 0710 Amgadoz contrasts Whisper Turbo's multi-million hour multilingual pruning dataset with Distil-Whisper's English-only distillation. He breaks down the mechanics of KL divergence distribution matching versus standard cross-entropy loss.
Real-Time Streaming Audio Decoding Mechanics 0700 Amgadoz walks through the engineering details of low-latency streaming transcription, explaining how Whisper's mandatory 30-second fixed input window requires padding and small-step buffer shifts to minimize decoder execution time.

Statements from this episode (13)

Assertion Supported
Most Open Vision Models Rely on Synthetic Data From Proprietary Models
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”
Vibhu Sapra Oct 13, 2024 ▶ 1:13
Assertion Supported
Molmo Outperforms Gemini 1.5 and Claude 3.5 Sonnet With 1M Samples
“They can get better than Gemini, 1.5, better than Claude, 3.5 sonnet, better than GPT for V at a much smaller size with about a million samples of data, which is very impressive, right?”
Vibhu Sapra Oct 13, 2024 ▶ 4:23
Insight
Speech-Based Image Annotation Produces Richer Training Data Faster Than Writing
“We ask annotators to describe the images in speech for 60 to 90 seconds, rather than asking them to write descriptions. They prompted them to describe everything in great detail, including descriptions of spatial positioning and relationships. So stuff like, y…”
Vibhu Sapra Oct 13, 2024 ▶ 6:00
Prediction Not checkable as stated
Major Foundation Model Companies Will Train on AI2's Vision Data
“The things that this model is good at are things that all the foundation companies, like they're just going to take our data and train on it.”
Nathan Lambert Oct 13, 2024 ▶ 12:15
Assertion Supported
Molmo Uses Base Model Without Instruction Tuning or Chat Template
“This is just, like, straight base model, no real instruction tuning. There's literally, like, no chat template for multi-turn. It just concatenates the messages together and, like, there's, like, go, look, good luck.”
Nathan Lambert Oct 13, 2024 ▶ 14:17
Disclosure
AI2 Did Not Log Raw Audio for Pixmo Annotations
“Yeah, it was like Whisper, and I don't remember the details, and I had asked, and it's like, I don't think they actually logged the audio. I was like, oh, this would be super cool for, like, other types of multimodal, but I don't think it was in the terms of t…”
Nathan Lambert Oct 13, 2024 ▶ 18:02
Assertion Supported
Pixmo Dataset Contains Approximately 1M Captions Across 700K Images
“They got about a million captions for 700,000 images, which I'm like, okay, that's kind of expensive.”
Vibhu Sapra Oct 13, 2024 ▶ 22:17
Assertion Supported
Molmo Reads Clocks but Fails to Generalize to Dials
“The model didn't work on clocks and then the lead was really on clocks and no models work on clocks. So they're like, we've got to make it work on clocks. One of the interesting things is that it doesn't work on dials, even though it works on clocks.”
Nathan Lambert Oct 13, 2024 ▶ 28:34
Insight
Distilling VLMs From Proprietary Models Copies Their Spatial Pointing Failures
“For example, all this proprietary stuff sucks at clocks, so nothing that's a distillation will be good at clocks. Nothing can point if you just distill from this. If they can't point, your VLM won't point, so we show how to get good data.”
Vibhu Sapra Oct 13, 2024 ▶ 32:22
Assertion Supported
Molmo-72B Ranks Second Behind Only GPT-4o in Human Preference Elo
“Most preference ranked was GPT-Four-O, then Momo-Semety-Two-B, then Gemini, then Sonnet, then the Seven-B.”
Vibhu Sapra Oct 13, 2024 ▶ 39:17
Insight
Q&A Data Drives VLM Detail Recognition Better Than Captioning
“My intuition at least is that I suspect it's really, the heavy lifting might be actually more towards the Q and A side rather than captioning side. I view captioning as a bootstrap, but at the end of the day, we want the model to generalize what they want to d…”
Eugene Xia Oct 13, 2024 ▶ 43:41
Assertion Partly supported
Molmo 1B Matches GPT-4V Across Academic Benchmarks and Elo
“The most efficient model, the one B is based on their one B MOE. That one matches performance of four V on most academic benchmarks and their ELO ranking.”
Vibhu Sapra Oct 13, 2024 ▶ 45:06
Assertion Supported
Molmo 72B Beats Proprietary Models on Academic Benchmarks
“Academic benchmarks wise, Their big one is the best state-of-the-art everything, better than proprietary, but ELO-wise, it sits behind four-oh.”
Vibhu Sapra Oct 13, 2024 ▶ 45:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.