Vision Models
topic on 5 shows · 9 statements across 8 episodes
Latent Space
the Startup Ideas Podcast
the a16z Podcast
Big Technology
20VC
9 statements about Vision Models, every show
Anandkumar: Existing video and vision world models incorrectly assume fixed resolutions
“That immediately distinguishes us from other so-called world models, whether it's video models, vision models, they all assume during training and inference, it's a fixed resolution.”
Schneider: Vision Models Can Automatically Enforce Brand Guidelines on AI Ads
“You basically can put a vision model over the outputs and be like, Hey, does this match brand style guides?
So what I'm getting back from Google Nano banana, I can be, I can, you know, basically qualify the, like, is, is the text all readable?
you know, doe…”
Acharya: Computer vision models have fallen behind hype and expectations
“I mean, I think RPA is super interesting, but vision models haven't nearly kept up with the sort of, you know, the way that we've talked about them.”
Acharya: Schools will use AI vision models for social-emotional learning
“The way we're going to get there is not by paying for two teachers in every classroom. It's probably some kind of a vision model. That's privacy first. They can observe the social interactions of children and let parents and teachers know what's going on.”
Ameisen: Language model neurons are far less directly interpretable than vision neurons
“If you look at just the neurons of a lot of vision models, you can See neurons that are curve detectors or that are edge detectors or that are high, low frequency detectors. And so you can sort of like make sense of the neurons mostly. But if you look at neuro…”
Ameisen: Superposition is more severe in language models than in vision models
“That means that like language models pack a lot more in less space than Vision models. So maybe like a kind of like really hand wavy analogy, right? It's like, well, if you want curve detectors, like you don't need that many curve detectors. You know, if each …”
Roucher: Effective web browsing agents must use vision and direct GUI inputs
“Web browsing is designed for humans. So that means it's really visual. And so a web browsing agent, a good one, should use, in my opinion a vision model and perform actions with point and click and keyboard, basically.”
Ravi: Video segmentation requires far less context than language models
“A difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi-term conversation. And so, you know, coupling this short-term spatial memory with this, like, longer-term object pointers we fou…”
Doshi: Computer vision lacks a single unified model equivalent to LLMs
“This is missing in vision, but definitely kind of exists in language and language. We can solve hundreds or thousands of different tasks. But in graphics, but in vision, it's all separated. It's kind of like where language was back three or four years ago wher…”