vision models

also referred to as: vision model

5 statements across 4 episodes · 1 bullish · 1 bearish · 4 people on the record · first statement Aug 7, 2024 by Nikhila Ravi · across every show →

Everything said about vision models, oldest first

Aug 7, 2024 neutral
Insight
Ravi: Video segmentation requires far less context than language models
“A difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi-term conversation. And so, you know, coupling this short-term spatial memory with this, like, longer-term object pointers we fou…”
Nikhila Ravi Aug 7, 2024 ▶ 44:41 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
Feb 13, 2025 positive
Opinion
Roucher: Effective web browsing agents must use vision and direct GUI inputs
“Web browsing is designed for humans. So that means it's really visual. And so a web browsing agent, a good one, should use, in my opinion a vision model and perform actions with point and click and keyboard, basically.”
Aymeric (Emmerich) Feb 13, 2025 ▶ 17:26 smol agents are all you need
Jun 6, 2025 neutral
Insight
Ameisen: Superposition is more severe in language models than in vision models
“That means that like language models pack a lot more in less space than Vision models. So maybe like a kind of like really hand wavy analogy, right? It's like, well, if you want curve detectors, like you don't need that many curve detectors. You know, if each …”
Emmanuel Ameisen Jun 6, 2025 ▶ 35:01 The Utility of Interpretability — Emmanuel Amiesen
Jun 6, 2025 neutral
Assertion Supported
Ameisen: Language model neurons are far less directly interpretable than vision neurons
“If you look at just the neurons of a lot of vision models, you can See neurons that are curve detectors or that are edge detectors or that are high, low frequency detectors. And so you can sort of like make sense of the neurons mostly. But if you look at neuro…”
Emmanuel Ameisen Jun 6, 2025 ▶ 34:31 The Utility of Interpretability — Emmanuel Amiesen
Sep 4, 2026 negative
Assertion Contradicted
Anandkumar: Existing video and vision world models incorrectly assume fixed resolutions
“That immediately distinguishes us from other so-called world models, whether it's video models, vision models, they all assume during training and inference, it's a fixed resolution.”
Anima Anandkumar Sep 4, 2026 ▶ 10:15 Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.