vision-language models

also referred to as: vision language model · vision language models · vlm · vlms

9 statements across 5 episodes · 2 bullish · 3 bearish · 6 people on the record · first statement Oct 13, 2024 by Vibhu Sapra · across every show →

Everything said about vision-language models, oldest first

Oct 13, 2024 negative
Insight
Distilling VLMs From Proprietary Models Copies Their Spatial Pointing Failures
“For example, all this proprietary stuff sucks at clocks, so nothing that's a distillation will be good at clocks. Nothing can point if you just distill from this. If they can't point, your VLM won't point, so we show how to get good data.”
Vibhu Sapra Oct 13, 2024 ▶ 32:22 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Oct 13, 2024 positive
Insight
Q&A Data Drives VLM Detail Recognition Better Than Captioning
“My intuition at least is that I suspect it's really, the heavy lifting might be actually more towards the Q and A side rather than captioning side. I view captioning as a bootstrap, but at the end of the day, we want the model to generalize what they want to d…”
Eugene Xia Oct 13, 2024 ▶ 43:41 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Oct 13, 2024 neutral
Assertion Supported
Most Open Vision Models Rely on Synthetic Data From Proprietary Models
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”
Vibhu Sapra Oct 13, 2024 ▶ 1:13 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Dec 22, 2024 neutral
Insight
Korupati: VLMs Fail at Gauges Due to E-Commerce Training Biases
“In the case of gauges, most gauges images aren't gauges in the wild. They're product Detail images like these where it's always set to zero. It's paired with an alt text that says something like GIVTO pressure sensor PSI zero to 30 or something. And so, the mo…”
Vik Korupati Dec 22, 2024 ▶ 48:05 Best of 2024 in Vision [LS Live @ NeurIPS]
Dec 22, 2024 bearish
Opinion
Korupati: Vision-Language Models Are Lagging Behind LLMs in Reasoning
“LLMs are showing enormous progress in reasoning, especially with the latest set of models that we've seen, but we're not really seeing, I have a feeling that VLMs are lagging behind, as we can see with these tasks that should be very simple for a human to do t…”
Vik Korupati Dec 22, 2024 ▶ 53:01 Best of 2024 in Vision [LS Live @ NeurIPS]
Dec 22, 2024 bullish
Prediction Not checkable as stated
Korupati: Visual Chain-of-Thought Will Probably Generalize Across Tasks
“The real question is, is it going to generalize? Probably, like, there's some signs from text models that when you train on a broad number of tasks, it does generalize, and I'm seeing some signs with our model as well.”
Vik Korupati Dec 22, 2024 ▶ 52:10 Best of 2024 in Vision [LS Live @ NeurIPS]
Apr 2, 2026 bearish
Opinion
Manning: Vision understanding stalled; language does 90% of work in VLMs
“I mean, I think it's fair to say that, you know, vision understanding sort of stalled out, right? You got to object recognition, and then progress just wasn't being made, right? If you look at any of these vision language models, it's the language that's doing…”
Chris Manning Apr 2, 2026 ▶ 5:02 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
Jul 16, 2026
Disclosure
Beam: Lila uses a vision-language model to automate Windows 95 machines
“We actually have a vision language model controlling a Windows 95 machine. Because that's the only way to automate it.”
Andy Beam Jul 16, 2026 ▶ 1:04:57 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Aug 3, 2026
Assertion Supported
Open-source inference engines often receive model weights before official launches
“Getting to the point of I can make a token out of this model is not that hard because generally the open source inference engines, your VLMs, SGLangs of the world oftentimes even receive weights ahead of time maintainers do, or the people making the model merg…”
Philip Kiely Aug 3, 2026 ▶ 13:51 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.