SAM 2

product on 1 show · 10 statements across 2 episodes · said 5 times in 2 episodes since 2024

Latent Space 5

Mentions by year, every show

tap a year for its mentions
00214120242025episodesmentions
01120242025episodes it came up in
0020.54120242025episodesmentions per episode

Latent Space 5

every mention on every show, scene by scene, with the transcript →

10 statements about SAM 2, every show

LATENT SPACE Disclosure
Ravi: Meta built SAM 3 using open-source community contributions to SAM 2
“In SAM-III we did leverage many of the open source contributions people have made on top of SAM-II. There were new data sets, there were new benchmarks, There were new kind of inference time optimizations. We adopt a lot of the things that the community builds…”
Nikhila Ravi Dec 18, 2025 ▶ 57:04 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
LATENT SPACE Assertion Supported
Ravi: Meta Achieved Fully Automated Annotation in SAM 1, Not SAM 2
“Getting to that fully automated data engine is something that we tried to do in SAM too. We actually didn't get to that fully automated approach. In SAM one, we did, we, you know, But the SA-I-B dataset that we released was fully annotated automatically. We di…”
Nikhila Ravi Dec 18, 2025 ▶ 45:38 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Unified vision models outperform composite multi-model pipelines
“Combining two models and sort of just smushing things together might not actually be as effective as if you really think about how to build things in a unified way.”
Nikhila Ravi Aug 7, 2024 ▶ 57:03 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
LATENT SPACE Assertion Supported
Nelson: Smallest SAM 2 model has 38M parameters and runs at 45 FPS
“The smallest model is thirty-eight million parameters and can run at 45 FPS on an A-one hundred, right?”
Joseph Nelson Aug 7, 2024 ▶ 52:54 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
Ravi: Video segmentation requires far less context than language models
“A difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi-term conversation. And so, you know, coupling this short-term spatial memory with this, like, longer-term object pointers we fou…”
Nikhila Ravi Aug 7, 2024 ▶ 44:41 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
LATENT SPACE Assertion Contradicted
Prior video object segmentation models lacked error recovery mechanisms
“That actually is a big limitation of current models, current video object segmentation models. Don't allow any way to recover if the model makes a mistake.”
Nikhila Ravi Aug 7, 2024 ▶ 43:46 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
LATENT SPACE Assertion Supported
Unified models enable faster error corrections via refinement clicks
“And we found that, you know, going from each phase, it both improved the efficiency and it improved the data quality. And in particular, when you get rid of this two-part model, one of the advantages is that when you make refinement clicks, so You prompt the m…”
Nikhila Ravi Aug 7, 2024 ▶ 37:45 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
LATENT SPACE Prediction Held up
Ravi: SAM 2 will soon run on-device and inside web browsers
“Like, I'm pretty sure soon we'll see like an on-device SAM-II or, you know, maybe even running in the browser or something. So I think that could definitely unlock some of these edge use cases.”
Nikhila Ravi Aug 7, 2024 ▶ 23:15 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
LATENT SPACE Assertion Supported
Ravi: SAM 2 runs roughly six times faster on video than SAM 1
“And in terms of the efficiency compared to SAM, so if we were to run SAM per frame on a video or run SAM two, it's around six times faster to run SAM two versus run SAM per frame.”
Nikhila Ravi Aug 7, 2024 ▶ 22:51 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson
LATENT SPACE Assertion Supported
Ravi: SAM 2's largest model is 224M parameters, one-third of SAM 1
“SAM-I model was around six hundred and thirty million parameters, a fraction of the size of these large language models, but very small. Actually SAM-II, the largest model is around two hundred and twenty-four million parameters. There's actually One third the…”
Nikhila Ravi Aug 7, 2024 ▶ 22:13 Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.