Aug 7, 2024 · 1h 0m · latent-space

Segment Anything 2: Memory + Vision = Object Permanence — with Nikhila Ravi and Joseph Nelson

Nikhila Ravi · 28m spoken Joseph Nelson · 23m spoken Shawn Wang · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Podcast, Meta FAIR lead author Nikhila Ravi and Roboflow's Joseph Nelson explore the release of Segment Anything 2 (SAM 2), detailing its architectural innovations, novel memory mechanisms for video object permanence, and the data engine powering zero-shot computer vision.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.8 Guest teaching 4.1 Guest disagreement 0.3 The hosts pushing back 1.5
05100:0015:0030:0045:001:00:000:04–3:36 · The hosts as informed peer 2/10 Nikhila Ravi's Engineering Journey and Transition to AI The host warmly introduces Nikhila Ravi and asks standard biographical questions about her transition from engineering at Cambridge to deep learning at Meta.3:36–13:01 · The hosts as informed peer 5/10 The Industry Impact and Generalization of Segment Anything Joseph presents production statistics from Roboflow usage, and swyx asks how zero-shot segmentation works in medical domains. Nikhila explains the class-agnostic design of the SA-1B dataset and visual prompting primitives.13:02–16:44 · The hosts as informed peer 1/10 Live Demonstration of Interactive Video Segmentation in SAM 2 Nikhila delivers a live narrated demonstration of SAM 2 on challenging video sequences, illustrating real-time tracking, occlusion recovery, and UI swim lanes.16:45–23:39 · The hosts as informed peer 6/10 Integrating UX Design with Architectural Efficiency in SAM 2 Joseph draws technical parallels between UX-driven design and model architecture, probing into the shift from ViT-H to Hiera image encoders and why browser-side querying was replaced with streaming server execution.23:39–31:57 · The hosts as informed peer 7/10 Open-Vocabulary Grounding and Class-Agnostic Model Design Choices Joseph demonstrates AutoDistill and questions why SAM 2 did not natively integrate open-vocabulary text grounding like Grounding DINO. Nikhila defends Meta's philosophy of maintaining laser focus on solving one fundamental capability well.31:57–35:53 · The hosts as informed peer 7/10 Out-of-Distribution Challenges and Handling Web Screenshot Data Joseph presents empirical screen share evidence highlighting SAM 2's failure to segment web UI elements for digital agents. Nikhila acknowledges the limitation and frames FAIR's role as providing generalist foundation tools rather than vertical niche adaptations.35:53–40:12 · The hosts as informed peer 4/10 Evolution of the Three-Phase SAM 2 Data Engine Swyx asks about scaling laws and dataset construction in SAM 2. Nikhila details the three-stage evolution of the data engine from per-frame SAM annotation to the unified model that enables non-destructive refinement clicks.40:12–47:48 · The hosts as informed peer 6/10 Memory Architecture, Object Permanence, and Error Recovery Swyx challenges why video memory is restricted to 6 frames rather than hundreds like LLM context windows. Nikhila educates the hosts on the dual-memory design combining high-res spatial memory with long-term object pointers.0:04–3:36 · Guest teaching 1/10 Nikhila Ravi's Engineering Journey and Transition to AI The host warmly introduces Nikhila Ravi and asks standard biographical questions about her transition from engineering at Cambridge to deep learning at Meta.3:36–13:01 · Guest teaching 5/10 The Industry Impact and Generalization of Segment Anything Joseph presents production statistics from Roboflow usage, and swyx asks how zero-shot segmentation works in medical domains. Nikhila explains the class-agnostic design of the SA-1B dataset and visual prompting primitives.13:02–16:44 · Guest teaching 4/10 Live Demonstration of Interactive Video Segmentation in SAM 2 Nikhila delivers a live narrated demonstration of SAM 2 on challenging video sequences, illustrating real-time tracking, occlusion recovery, and UI swim lanes.16:45–23:39 · Guest teaching 4/10 Integrating UX Design with Architectural Efficiency in SAM 2 Joseph draws technical parallels between UX-driven design and model architecture, probing into the shift from ViT-H to Hiera image encoders and why browser-side querying was replaced with streaming server execution.23:39–31:57 · Guest teaching 4/10 Open-Vocabulary Grounding and Class-Agnostic Model Design Choices Joseph demonstrates AutoDistill and questions why SAM 2 did not natively integrate open-vocabulary text grounding like Grounding DINO. Nikhila defends Meta's philosophy of maintaining laser focus on solving one fundamental capability well.31:57–35:53 · Guest teaching 3/10 Out-of-Distribution Challenges and Handling Web Screenshot Data Joseph presents empirical screen share evidence highlighting SAM 2's failure to segment web UI elements for digital agents. Nikhila acknowledges the limitation and frames FAIR's role as providing generalist foundation tools rather than vertical niche adaptations.35:53–40:12 · Guest teaching 6/10 Evolution of the Three-Phase SAM 2 Data Engine Swyx asks about scaling laws and dataset construction in SAM 2. Nikhila details the three-stage evolution of the data engine from per-frame SAM annotation to the unified model that enables non-destructive refinement clicks.40:12–47:48 · Guest teaching 6/10 Memory Architecture, Object Permanence, and Error Recovery Swyx challenges why video memory is restricted to 6 frames rather than hundreds like LLM context windows. Nikhila educates the hosts on the dual-memory design combining high-res spatial memory with long-term object pointers.0:04–3:36 · Guest disagreement 0/10 Nikhila Ravi's Engineering Journey and Transition to AI The host warmly introduces Nikhila Ravi and asks standard biographical questions about her transition from engineering at Cambridge to deep learning at Meta.3:36–13:01 · Guest disagreement 0/10 The Industry Impact and Generalization of Segment Anything Joseph presents production statistics from Roboflow usage, and swyx asks how zero-shot segmentation works in medical domains. Nikhila explains the class-agnostic design of the SA-1B dataset and visual prompting primitives.13:02–16:44 · Guest disagreement 0/10 Live Demonstration of Interactive Video Segmentation in SAM 2 Nikhila delivers a live narrated demonstration of SAM 2 on challenging video sequences, illustrating real-time tracking, occlusion recovery, and UI swim lanes.16:45–23:39 · Guest disagreement 0/10 Integrating UX Design with Architectural Efficiency in SAM 2 Joseph draws technical parallels between UX-driven design and model architecture, probing into the shift from ViT-H to Hiera image encoders and why browser-side querying was replaced with streaming server execution.23:39–31:57 · Guest disagreement 1/10 Open-Vocabulary Grounding and Class-Agnostic Model Design Choices Joseph demonstrates AutoDistill and questions why SAM 2 did not natively integrate open-vocabulary text grounding like Grounding DINO. Nikhila defends Meta's philosophy of maintaining laser focus on solving one fundamental capability well.31:57–35:53 · Guest disagreement 1/10 Out-of-Distribution Challenges and Handling Web Screenshot Data Joseph presents empirical screen share evidence highlighting SAM 2's failure to segment web UI elements for digital agents. Nikhila acknowledges the limitation and frames FAIR's role as providing generalist foundation tools rather than vertical niche adaptations.35:53–40:12 · Guest disagreement 0/10 Evolution of the Three-Phase SAM 2 Data Engine Swyx asks about scaling laws and dataset construction in SAM 2. Nikhila details the three-stage evolution of the data engine from per-frame SAM annotation to the unified model that enables non-destructive refinement clicks.40:12–47:48 · Guest disagreement 0/10 Memory Architecture, Object Permanence, and Error Recovery Swyx challenges why video memory is restricted to 6 frames rather than hundreds like LLM context windows. Nikhila educates the hosts on the dual-memory design combining high-res spatial memory with long-term object pointers.0:04–3:36 · The hosts pushing back 0/10 Nikhila Ravi's Engineering Journey and Transition to AI The host warmly introduces Nikhila Ravi and asks standard biographical questions about her transition from engineering at Cambridge to deep learning at Meta.3:36–13:01 · The hosts pushing back 1/10 The Industry Impact and Generalization of Segment Anything Joseph presents production statistics from Roboflow usage, and swyx asks how zero-shot segmentation works in medical domains. Nikhila explains the class-agnostic design of the SA-1B dataset and visual prompting primitives.13:02–16:44 · The hosts pushing back 0/10 Live Demonstration of Interactive Video Segmentation in SAM 2 Nikhila delivers a live narrated demonstration of SAM 2 on challenging video sequences, illustrating real-time tracking, occlusion recovery, and UI swim lanes.16:45–23:39 · The hosts pushing back 1/10 Integrating UX Design with Architectural Efficiency in SAM 2 Joseph draws technical parallels between UX-driven design and model architecture, probing into the shift from ViT-H to Hiera image encoders and why browser-side querying was replaced with streaming server execution.23:39–31:57 · The hosts pushing back 3/10 Open-Vocabulary Grounding and Class-Agnostic Model Design Choices Joseph demonstrates AutoDistill and questions why SAM 2 did not natively integrate open-vocabulary text grounding like Grounding DINO. Nikhila defends Meta's philosophy of maintaining laser focus on solving one fundamental capability well.31:57–35:53 · The hosts pushing back 4/10 Out-of-Distribution Challenges and Handling Web Screenshot Data Joseph presents empirical screen share evidence highlighting SAM 2's failure to segment web UI elements for digital agents. Nikhila acknowledges the limitation and frames FAIR's role as providing generalist foundation tools rather than vertical niche adaptations.35:53–40:12 · The hosts pushing back 1/10 Evolution of the Three-Phase SAM 2 Data Engine Swyx asks about scaling laws and dataset construction in SAM 2. Nikhila details the three-stage evolution of the data engine from per-frame SAM annotation to the unified model that enables non-destructive refinement clicks.40:12–47:48 · The hosts pushing back 2/10 Memory Architecture, Object Permanence, and Error Recovery Swyx challenges why video memory is restricted to 6 frames rather than hundreds like LLM context windows. Nikhila educates the hosts on the dual-memory design combining high-res spatial memory with long-term object pointers.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 30:05 Refusing feature creep in foundation models

Nikhila firmly rejects the idea that SAM 2 should have absorbed text prompting and multi-modal grounding natively, insisting on Meta's disciplined focus on delivering step-change performance on narrow primitives.

Hardest push from the hosts ▶ 34:00 Highlighting failure on digital web screenshots

Joseph shares screen evidence to directly challenge SAM 2's out-of-distribution capabilities on web screenshots and UI button parsing for autonomous agents.

Biggest teaching moment ▶ 44:40 Explaining why video vision memory differs from LLM context

Nikhila corrects swyx's assumption that vision models need 600-frame context windows like LLMs, explaining the mathematical and practical sufficiency of short-term spatial memory paired with object pointers.

The host holds their own ▶ 28:10 Showcasing AutoDistill integration pipeline

Joseph demonstrates Roboflow's AutoDistill framework combining Grounding DINO, Florence-2, and SAM to show how production workflows solve the ontology problem externally.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Nikhila Ravi's Engineering Journey and Transition to AI 2100 The host warmly introduces Nikhila Ravi and asks standard biographical questions about her transition from engineering at Cambridge to deep learning at Meta.
The Industry Impact and Generalization of Segment Anything 5501 Joseph presents production statistics from Roboflow usage, and swyx asks how zero-shot segmentation works in medical domains. Nikhila explains the class-agnostic design of the SA-1B dataset and visual prompting primitives.
Live Demonstration of Interactive Video Segmentation in SAM 2 1400 Nikhila delivers a live narrated demonstration of SAM 2 on challenging video sequences, illustrating real-time tracking, occlusion recovery, and UI swim lanes.
Integrating UX Design with Architectural Efficiency in SAM 2 6401 Joseph draws technical parallels between UX-driven design and model architecture, probing into the shift from ViT-H to Hiera image encoders and why browser-side querying was replaced with streaming server execution.
Open-Vocabulary Grounding and Class-Agnostic Model Design Choices 7413 Joseph demonstrates AutoDistill and questions why SAM 2 did not natively integrate open-vocabulary text grounding like Grounding DINO. Nikhila defends Meta's philosophy of maintaining laser focus on solving one fundamental capability well.
Out-of-Distribution Challenges and Handling Web Screenshot Data 7314 Joseph presents empirical screen share evidence highlighting SAM 2's failure to segment web UI elements for digital agents. Nikhila acknowledges the limitation and frames FAIR's role as providing generalist foundation tools rather than vertical niche adaptations.
Evolution of the Three-Phase SAM 2 Data Engine 4601 Swyx asks about scaling laws and dataset construction in SAM 2. Nikhila details the three-stage evolution of the data engine from per-frame SAM annotation to the unified model that enables non-destructive refinement clicks.
Memory Architecture, Object Permanence, and Error Recovery 6602 Swyx challenges why video memory is restricted to 6 frames rather than hundreds like LLM context windows. Nikhila educates the hosts on the dual-memory design combining high-res spatial memory with long-term object pointers.

Statements from this episode (19)

Opinion
Ravi: SAM impacts medicine more than I could have as a doctor
“Actually Sam is having so much impact in medicine, probably more than I could have ever had as a doctor myself.”
Nikhila Ravi Aug 7, 2024 ▶ 3:16
Disclosure
Nelson: Roboflow users labeled 49M images using SAM in one year
“I recently pulled statistics from the usage of Sam in RoboFlow over the course of the last year, and users have labeled about forty nine million images using segment anything on the hosted side of the RoboFlow platform, and that's like five million in the last…”
Joseph Nelson Aug 7, 2024 ▶ 6:48
Insight
Production computer vision still requires user prompting to isolate targets
“But even if you had, like, a perfect SAM, like an omniscient SAM that could see every segment in every domain with all pixels perfectly outlined, in production, you still need some way to almost, like, signal to the model what you care about.”
Joseph Nelson Aug 7, 2024 ▶ 11:53
Assertion Supported
Ravi: SAM 2 tracks moving octopus tentacles zero-shot in underwater video
“There's like underwater videos that it works actually really well for, even though we, models never really seen an octopus before. And octopus have a lot of Moving parts that SAM-II can actually quite effectively keep track of all the different tentacles.”
Nikhila Ravi Aug 7, 2024 ▶ 14:57
Insight
Public demos acting as internal annotation tools directly accelerate model quality
“And the other piece here is that the demo is actually the annotation tool. So we actually Use the demo as a way to improve our annotation tool. And so then it becomes very natural to invest in building a good demo because it speeds up your annotation and impro…”
Nikhila Ravi Aug 7, 2024 ▶ 18:24
Assertion Supported
Ravi: SAM 2's largest model is 224M parameters, one-third of SAM 1
“SAM-I model was around six hundred and thirty million parameters, a fraction of the size of these large language models, but very small. Actually SAM-II, the largest model is around two hundred and twenty-four million parameters. There's actually One third the…”
Nikhila Ravi Aug 7, 2024 ▶ 22:13
Assertion Supported
Ravi: SAM 2 runs roughly six times faster on video than SAM 1
“And in terms of the efficiency compared to SAM, so if we were to run SAM per frame on a video or run SAM two, it's around six times faster to run SAM two versus run SAM per frame.”
Nikhila Ravi Aug 7, 2024 ▶ 22:51
Prediction Held up
Ravi: SAM 2 will soon run on-device and inside web browsers
“Like, I'm pretty sure soon we'll see like an on-device SAM-II or, you know, maybe even running in the browser or something. So I think that could definitely unlock some of these edge use cases.”
Nikhila Ravi Aug 7, 2024 ▶ 23:15
Assertion Supported
Developers are combining SAM 2 with Grounding DINO and Florence-2
“We've already seen applications of using SAM-II in tandem with models like Grounding Dino or Florence-II, so that people can basically text prompt And then get the benefits of the zero shot segmentation at the same time as getting the open form querying.”
Joseph Nelson Aug 7, 2024 ▶ 27:46
Disclosure
Meta limits SAM releases to step-change breakthroughs in narrow capabilities
“So as you've probably seen with SAM and SAM-II, it's a fairly narrow problem, but we really try to make it a step change in the capability. And so with each Version. We are trying to limit the focus on one thing that we can know we can do really well. And in t…”
Nikhila Ravi Aug 7, 2024 ▶ 30:20
Opinion
Segment Anything models underperform compared to other models on screenshots
“And one place where, interestingly, segment anything may be less performant than other models is handling screenshots.”
Joseph Nelson Aug 7, 2024 ▶ 32:51
Disclosure
Meta FAIR focuses on building foundational models, not specific use cases
“Fair, we don't really build with a specific use case in mind. We try to build like these foundational models that can be applied to lots of different use cases out of the box.”
Nikhila Ravi Aug 7, 2024 ▶ 34:47
Assertion Supported
The three-phase architecture evolution of Meta's SAM 2 data engine
“We started with just SAM. We apply SAM per frame. That's like the most basic way of extending SAM to video. Then the most obvious thing to do is to take the output masks from SAM and then provide it as input into a video object segmentation model that takes th…”
Nikhila Ravi Aug 7, 2024 ▶ 37:03
Assertion Supported
Unified models enable faster error corrections via refinement clicks
“And we found that, you know, going from each phase, it both improved the efficiency and it improved the data quality. And in particular, when you get rid of this two-part model, one of the advantages is that when you make refinement clicks, so You prompt the m…”
Nikhila Ravi Aug 7, 2024 ▶ 37:45
Assertion Contradicted
Prior video object segmentation models lacked error recovery mechanisms
“That actually is a big limitation of current models, current video object segmentation models. Don't allow any way to recover if the model makes a mistake.”
Nikhila Ravi Aug 7, 2024 ▶ 43:46
Insight
Ravi: Video segmentation requires far less context than language models
“A difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi-term conversation. And so, you know, coupling this short-term spatial memory with this, like, longer-term object pointers we fou…”
Nikhila Ravi Aug 7, 2024 ▶ 44:41
Assertion Supported
Nelson: Smallest SAM 2 model has 38M parameters and runs at 45 FPS
“The smallest model is thirty-eight million parameters and can run at 45 FPS on an A-one hundred, right?”
Joseph Nelson Aug 7, 2024 ▶ 52:54
Assertion Supported
SAM 2 outperforms SAM on video frames due to differing data distributions
“And we find that actually SAM-II is a lot better than SAM when it comes to segmenting objects in video frames, because they actually have a sort of slightly different distribution than images.”
Nikhila Ravi Aug 7, 2024 ▶ 56:51
Insight
Unified vision models outperform composite multi-model pipelines
“Combining two models and sort of just smushing things together might not actually be as effective as if you really think about how to build things in a unified way.”
Nikhila Ravi Aug 7, 2024 ▶ 57:03
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.