Vision Transformer, every mention
3 scenes · ← back to Vision Transformer
tap a year for its mentions
every year anyone Shawn Wang 2Isaac Robinson 2Vibhu Sapra 1
Verbatim, from the transcripts: the passages where Vision Transformer comes up
Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He
- ▶ 16:46 Shawn Wang We've covered the vision transformer. 2 times in the scene
Best of 2024 in Vision [LS Live @ NeurIPS]
- ▶ 10:18 Isaac Robinson SAM can be used on a single image, in which case the only difference between SAM and SAM is that image encoder, which SAM used a standard VIT. 2 times in the scene
[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz
- ▶ 34:10 Vibhu Sapra So for the vision encoder, we release, uh, all of our released models use OpenAI's VIT large Clip model, which provides consistently good results.