Vision Transformer, every mention
5 scenes across 2 shows · ← back to Vision Transformer
tap a year for its mentions
Latent Space 5
the MAD Podcast 2
every year every show
Latent Space 5
the MAD Podcast 2
Verbatim, from the transcripts: passages where Vision Transformer comes up on Latent Space, the MAD Podcast
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
- ▶ 16:46 Shawn Wang We've covered the vision transformer. 2 times in the scene
AI is Already Building AI — Google DeepMind’s Mostafa Dehghani
- ▶ 0:31 Matt Turck Today, my guest is Mustafa Degani, a top AI researcher at Google DeepMind, and a core contributor to some of the most influential architectural breakthroughs of the last decade, including Universal Transformers, the Vision Transformer, and…
- ▶ 39:56 Matt Turck Another fundamentally important contribution to the, to the field that, uh, you did was the visual transformer paper in 2022.
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 10:18 Isaac Robinson SAM can be used on a single image, in which case the only difference between SAM and SAM is that image encoder, which SAM used a standard VIT. 2 times in the scene
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 34:10 Vibhu Sapra So for the vision encoder, we release, uh, all of our released models use OpenAI's VIT large Clip model, which provides consistently good results.