Vision Language Models

topic on 7 shows · 21 statements across 15 episodes

the Y Combinator Startup Podcast In Depth Latent Space No Priors the MAD Podcast the a16z Podcast TBPN

21 statements about Vision Language Models, every show

a16z Disclosure
Moe: Developers abandon proprietary models due to false-positive safety guardrails
“A lot of our developer within Infrax and for VLM are like retreating from using Fable five because you have a two hour job and you trigger the red line, which is false positive. And then you have to lose all of your work. And so a lot of our developer are usin…”
Simon Moe Aug 5, 2026 ▶ 33:51 How Open Source Became AI's Backbone | Inferact with a16z
LATENT SPACE Assertion Supported
Open-source inference engines often receive model weights before official launches
“Getting to the point of I can make a token out of this model is not that hard because generally the open source inference engines, your VLMs, SGLangs of the world oftentimes even receive weights ahead of time maintainers do, or the people making the model merg…”
Philip Kiely Aug 3, 2026 ▶ 13:51 Next 100x in AI: Inference, Networking, & Self-Optimizing Models — Philip Kiely & Ali Taha, Baseten
LATENT SPACE Disclosure
Beam: Lila uses a vision-language model to automate Windows 95 machines
“We actually have a vision language model controlling a Windows 95 machine. Because that's the only way to automate it.”
Andy Beam Jul 16, 2026 ▶ 1:04:57 🔬 RL with Verifiable Rewards, but the Verifier is a Lab — Lila Sciences
Vuong: Vision-language models transfer semantic knowledge to low-level physical actions
“And what this two work really show is that if you start from a vision language model that is really powerful, and you kind of use robotic data to adapt this model to speak robot language, if you will then you see a lot of transfer from the kind of knowledge th…”
Quan Vuong Apr 16, 2026 ▶ 4:22 The GPT Moment for Robotics Is Here · Y Combinator
Manning: Vision understanding stalled; language does 90% of work in VLMs
“I mean, I think it's fair to say that, you know, vision understanding sort of stalled out, right? You got to object recognition, and then progress just wasn't being made, right? If you look at any of these vision language models, it's the language that's doing…”
Chris Manning Apr 2, 2026 ▶ 5:02 Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun
MAD Prediction Not checkable as stated
Patel: Most AI chips will be consumed via open-source engines, not direct programming
“I think most AI chips will not be consumed by people programming anything for it. They will download an open source inference engine, they will download an open source model, and then they will put it on the, and it's really simple to download VLM and, like, m…”
Dylan Patel Feb 5, 2026 ▶ 10:58 Dylan Patel: NVIDIA's New Moat & Why China is "Semiconductor Pilled”
MAD Prediction Held up
Patel: All major non-Nvidia AI chips will fully support vLLM by mid-2026
“All of them will have a very good UX for download model, run model on VLM by The middle of the year, I think, right? Certainly AMD is already there by the end of this quarter.”
Dylan Patel Feb 5, 2026 ▶ 15:52 Dylan Patel: NVIDIA's New Moat & Why China is "Semiconductor Pilled”
NO PRIORS Insight
Laskin: Sensory and vision-language model rewards are far more hackable than LLM rewards
“The challenge is that if we, if you think that language model rewards are hackable vision language model rewards or, you know, like other sensory signal rewards are infinitely more hackable.”
Misha Laskin Jul 17, 2025 ▶ 51:50 No Priors Ep. 123 | With ReflectionAI Co-Founder and CEO Misha Laskin
IN DEPTH Disclosure
Abraham: Reducto uses six vision models alongside a VLM
“At this point we have Six different vision models in a VLM”
Adit Abraham Apr 29, 2025 ▶ 13:11 How a weekend hack became a multimillion-dollar AI startup | Adit Abraham (CEO at Reducto)
TBPN Insight
Hausman: Robot physical actions function as another language for multimodal models
“And what we start to realize is that all of these different data sources contribute to each other. They give you just like a bigger picture of what the world is like and better understanding. And it just turns out that robot actions is just like yet another la…”
Karol Hausman Apr 26, 2025 ▶ 26:15 The Race to Create General-Purpose Robots | Karol Hausman & Lachy Groom on TBPN
TBPN Insight
Ramade: Base AI capabilities are becoming commoditized APIs
“Base cognitive functions are just becoming available as an API. So like vision is just going to be available. We shouldn't work on a vision problem, like go fine tune like a Yolo or whatever. VLM is your favorite.”
Pratap Ramade Apr 10, 2025 ▶ 8:49 How AI Can Enable AMBITION| | Pratap Ramade on TBPN April 8th
NO PRIORS Disclosure
Physical Intelligence uses transformers and pre-trained VLMs for robot models
“And in terms of the architecture, we're using transformers and we are using pre-trained models, pre-trained vision language models, and that allows you to leverage all of the rich information on the internet.”
Chelsea Finn Mar 20, 2025 ▶ 6:51 No Priors Ep. 107 | With Physical Intelligence Co-Founder Chelsea Finn
NO PRIORS Insight
Finn: Pre-trained VLMs let robots perform tasks with unseen internet concepts
“We had a research result a couple years ago where we showed that if you leverage vision language models, then you could actually get the robot to do tasks that require concepts that were never in the robot's training data, but were in the internet.”
Chelsea Finn Mar 20, 2025 ▶ 7:05 No Priors Ep. 107 | With Physical Intelligence Co-Founder Chelsea Finn
Korupati: VLMs Fail at Gauges Due to E-Commerce Training Biases
“In the case of gauges, most gauges images aren't gauges in the wild. They're product Detail images like these where it's always set to zero. It's paired with an alt text that says something like GIVTO pressure sensor PSI zero to 30 or something. And so, the mo…”
Vik Korupati Dec 22, 2024 ▶ 48:05 Best of 2024 in Vision [LS Live @ NeurIPS]
LATENT SPACE Prediction Not checkable as stated
Korupati: Visual Chain-of-Thought Will Probably Generalize Across Tasks
“The real question is, is it going to generalize? Probably, like, there's some signs from text models that when you train on a broad number of tasks, it does generalize, and I'm seeing some signs with our model as well.”
Vik Korupati Dec 22, 2024 ▶ 52:10 Best of 2024 in Vision [LS Live @ NeurIPS]
Korupati: Vision-Language Models Are Lagging Behind LLMs in Reasoning
“LLMs are showing enormous progress in reasoning, especially with the latest set of models that we've seen, but we're not really seeing, I have a feeling that VLMs are lagging behind, as we can see with these tasks that should be very simple for a human to do t…”
Vik Korupati Dec 22, 2024 ▶ 53:01 Best of 2024 in Vision [LS Live @ NeurIPS]
NO PRIORS Opinion
Dolgov: Off-the-shelf AI models cannot achieve human-beating driverless safety records
“The power of transformers, the power of realism is mind-blowing, right? So with just a little bit of effort, you get something on the road, and it works. You can, you know, drive, I don't know, 1000 of miles, and we just, it will blow your mind. But then is th…”
Dmitri Dolgov Oct 24, 2024 ▶ 39:52 No Priors Ep. 87 | With Co-CEO of Waymo Dmitri Dolgov
LATENT SPACE Assertion Supported
Most Open Vision Models Rely on Synthetic Data From Proprietary Models
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”
Vibhu Sapra Oct 13, 2024 ▶ 1:13 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Distilling VLMs From Proprietary Models Copies Their Spatial Pointing Failures
“For example, all this proprietary stuff sucks at clocks, so nothing that's a distillation will be good at clocks. Nothing can point if you just distill from this. If they can't point, your VLM won't point, so we show how to get good data.”
Vibhu Sapra Oct 13, 2024 ▶ 32:22 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
Q&A Data Drives VLM Detail Recognition Better Than Captioning
“My intuition at least is that I suspect it's really, the heavy lifting might be actually more towards the Q and A side rather than captioning side. I view captioning as a bootstrap, but at the end of the day, we want the model to generalize what they want to d…”
Eugene Xia Oct 13, 2024 ▶ 43:41 [Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
a16z Disclosure
Dolgov: Waymo is combining autonomous driving AI with vision-language models
“So that, that's what we've been very focused on that Waymo most recently is taking kind of the AI backbone and all of the AI, the Waymo AI that is over the years we've built up that is really proficient at this task of autonomous driving and combining it with …”
Dmitry Dolgov Aug 5, 2024 ▶ 8:44 How Waymo Is Using GenAI to Build a Better Driver

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.