Vision Transformer, every mention

3 scenes · ← back to Vision Transformer

tap a year for its mentions
002132202420252026episodesmentions
012202420252026episodes it came up in
001122202420252026episodesmentions per episode

every year anyone Shawn Wang 2Isaac Robinson 2Vibhu Sapra 1

Verbatim, from the transcripts: the passages where Vision Transformer comes up

loading…

Inside xAI: Building Grok Imagine in 3 Months, Videogen vs World Models, and Video Agents— Ethan He Jun 1, 2026 · 2 mentions

Best of 2024 in Vision [LS Live @ NeurIPS] Dec 22, 2024 · 2 mentions

  • ▶ 10:18 Isaac Robinson SAM can be used on a single image, in which case the only difference between SAM and SAM is that image encoder, which SAM used a standard VIT. 2 times in the scene

[Paper Club] Molmo + Pixmo + Whisper 3 Turbo - with Vibhu Sapra, Nathan Lambert, Amgadoz Oct 13, 2024 · 1 mention

  • ▶ 34:10 Vibhu Sapra So for the vision encoder, we release, uh, all of our released models use OpenAI's VIT large Clip model, which provides consistently good results.
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.