Dec 18, 2025 · 1h 15m · latent-space

SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)

Pengchuan Zhang · 20m spoken Joseph Nelson · 20m spoken Nikhila Ravi · 17m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this podcast episode, Meta FAIR researchers and Roboflow leadership discuss the release, architecture, and real-world impact of SAM 3, Meta's open-source vision foundation model for concept segmentation and tracking. The panel explores technical innovations including automated AI verification, presence tokens, agentic multimodal grounding, and practical deployment workflows across diverse industries.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.8 Guest teaching 3.2 Guest disagreement 1.2 The hosts pushing back 1.3
05100:0020:0040:001:00:000:03–5:31 · The hosts as informed peer 5/10 Introductions and Computer Vision Backgrounds Swix mistakenly assumes SAM 3 added a 3D dimension based on the name, prompting Nikhila to immediately correct him that SAM 3, SAM 3 Objects, and SAM 3 Body are separate models. Joseph Nelson highlights Roboflow's custom RF-DETR real-time edge transformer model, establishing co-host technical credibility.5:31–10:51 · The hosts as informed peer 6/10 Live Demonstration of SAM 3 Concept Prompting Nikhila presents a live demo of concept prompting and video tracking, while Swix and Joseph drill down on real-time inference latency and parallel multi-GPU batching mechanics described in the paper.10:51–13:32 · The hosts as informed peer 5/10 Concept Segmentation and the SA-Co Benchmark Evolution The hosts inquire about how text prompting evolved from a prototype in SAM 2 to a full 200k+ concept benchmark in SA-Co. Nikhila explains the architectural shift from open-ended natural language to granular, atomic visual concepts.13:38–18:28 · The hosts as informed peer 6/10 Real-World Industrial and Scientific Deployments of SAM Joseph details Roboflow's production analytics, citing over 106 million SAM-assisted annotations across biology, aerial mapping, and industrial robotics. Nikhila welcomes the real-world validation data as the ultimate benchmark beyond synthetic test sets.18:28–24:02 · The hosts as informed peer 6/10 Domain Fine-Tuning, Negative Examples, and Presence Tokens Swix questions the balance between positive and negative training examples, prompting Nikhila to reveal that over 70% of annotations in SAM 3 are negative phrases handled via an explicit learned presence token separating recognition from localization.24:02–28:11 · The hosts as informed peer 6/10 Architectural Decoupling of Visual Detection and Tracking Swix highlights the architecture diagram's newly introduced components. Nikhila explains why decoupling the identity-agnostic detector from the identity-preserving tracker resolved fundamental task conflicts during video segmentation.28:11–37:57 · The hosts as informed peer 7/10 SAM 3 as Multimodal Agent and Live Benchmarks Swix probes the necessity of SAM tool-calling over native MLLM grounding capabilities, referencing Table 8 benchmark metrics. Joseph runs live comparisons against Gemini 3 Pro and Florence-2, demonstrating SAM 3's speed and dense mask fidelity on occluded targets.37:57–43:03 · The hosts as informed peer 5/10 Automated Data Engine and Superhuman AI Verification Pengchuan breaks down SAM 3's multi-stage data engine, explaining how fine-tuning Llama 3.2 on verification tasks achieved superhuman precision, dropping human-in-the-loop per-datapoint annotation time from 2 minutes down to 25 seconds.43:04–47:36 · The hosts as informed peer 6/10 Surpassing Human Performance and Video Pipeline Bottlenecks Swix asks what happens when training runs out of human annotators. Pengchuan cautions against blind optimism, arguing that surpassing human performance requires a transition from supervised imitation learning to RLHF-style visual preference optimization.47:36–51:05 · The hosts as informed peer 6/10 Video Temporal Smoothing and Identity Tracking Nuances Swix brings up the masklet detection score and asks why other temporal models fail to smooth across time windows. Pengchuan explains the core engineering trade-off between live streaming latency and accumulated temporal context.51:06–56:47 · The hosts as informed peer 6/10 Native Multimodal Perception Versus Tool Calling for AGI Joseph poses the core architectural debate between native multimodal visual grounding and external modular tool calling. Pengchuan argues that basic perception like counting must become native System-1 cognition, with tools reserved for dense multi-step reasoning.56:47–1:03:10 · The hosts as informed peer 5/10 Open Source Vision Ecosystem and Future Roadmaps The conversation covers open source contributions and the future roadmap for vision models. Pengchuan identifies end-to-end video training and smaller edge models as top priorities, while Swix questions how perception models interface with explicit robotics world models.1:03:10–1:12:07 · The hosts as informed peer 7/10 Roboflow Deployment, Auto-Labeling, and Human Intent Alignment Swix challenges the premise of scalar confidence sliders for handling ambiguous visual concepts like reflections, arguing that richer iterative prompting is required. Pengchuan and Joseph agree, illustrating how subjective human intent necessitates interactive refinement and fine-tuning.0:03–5:31 · Guest teaching 4/10 Introductions and Computer Vision Backgrounds Swix mistakenly assumes SAM 3 added a 3D dimension based on the name, prompting Nikhila to immediately correct him that SAM 3, SAM 3 Objects, and SAM 3 Body are separate models. Joseph Nelson highlights Roboflow's custom RF-DETR real-time edge transformer model, establishing co-host technical credibility.5:31–10:51 · Guest teaching 2/10 Live Demonstration of SAM 3 Concept Prompting Nikhila presents a live demo of concept prompting and video tracking, while Swix and Joseph drill down on real-time inference latency and parallel multi-GPU batching mechanics described in the paper.10:51–13:32 · Guest teaching 3/10 Concept Segmentation and the SA-Co Benchmark Evolution The hosts inquire about how text prompting evolved from a prototype in SAM 2 to a full 200k+ concept benchmark in SA-Co. Nikhila explains the architectural shift from open-ended natural language to granular, atomic visual concepts.13:38–18:28 · Guest teaching 1/10 Real-World Industrial and Scientific Deployments of SAM Joseph details Roboflow's production analytics, citing over 106 million SAM-assisted annotations across biology, aerial mapping, and industrial robotics. Nikhila welcomes the real-world validation data as the ultimate benchmark beyond synthetic test sets.18:28–24:02 · Guest teaching 4/10 Domain Fine-Tuning, Negative Examples, and Presence Tokens Swix questions the balance between positive and negative training examples, prompting Nikhila to reveal that over 70% of annotations in SAM 3 are negative phrases handled via an explicit learned presence token separating recognition from localization.24:02–28:11 · Guest teaching 3/10 Architectural Decoupling of Visual Detection and Tracking Swix highlights the architecture diagram's newly introduced components. Nikhila explains why decoupling the identity-agnostic detector from the identity-preserving tracker resolved fundamental task conflicts during video segmentation.28:11–37:57 · Guest teaching 3/10 SAM 3 as Multimodal Agent and Live Benchmarks Swix probes the necessity of SAM tool-calling over native MLLM grounding capabilities, referencing Table 8 benchmark metrics. Joseph runs live comparisons against Gemini 3 Pro and Florence-2, demonstrating SAM 3's speed and dense mask fidelity on occluded targets.37:57–43:03 · Guest teaching 5/10 Automated Data Engine and Superhuman AI Verification Pengchuan breaks down SAM 3's multi-stage data engine, explaining how fine-tuning Llama 3.2 on verification tasks achieved superhuman precision, dropping human-in-the-loop per-datapoint annotation time from 2 minutes down to 25 seconds.43:04–47:36 · Guest teaching 4/10 Surpassing Human Performance and Video Pipeline Bottlenecks Swix asks what happens when training runs out of human annotators. Pengchuan cautions against blind optimism, arguing that surpassing human performance requires a transition from supervised imitation learning to RLHF-style visual preference optimization.47:36–51:05 · Guest teaching 4/10 Video Temporal Smoothing and Identity Tracking Nuances Swix brings up the masklet detection score and asks why other temporal models fail to smooth across time windows. Pengchuan explains the core engineering trade-off between live streaming latency and accumulated temporal context.51:06–56:47 · Guest teaching 3/10 Native Multimodal Perception Versus Tool Calling for AGI Joseph poses the core architectural debate between native multimodal visual grounding and external modular tool calling. Pengchuan argues that basic perception like counting must become native System-1 cognition, with tools reserved for dense multi-step reasoning.56:47–1:03:10 · Guest teaching 3/10 Open Source Vision Ecosystem and Future Roadmaps The conversation covers open source contributions and the future roadmap for vision models. Pengchuan identifies end-to-end video training and smaller edge models as top priorities, while Swix questions how perception models interface with explicit robotics world models.1:03:10–1:12:07 · Guest teaching 3/10 Roboflow Deployment, Auto-Labeling, and Human Intent Alignment Swix challenges the premise of scalar confidence sliders for handling ambiguous visual concepts like reflections, arguing that richer iterative prompting is required. Pengchuan and Joseph agree, illustrating how subjective human intent necessitates interactive refinement and fine-tuning.0:03–5:31 · Guest disagreement 2/10 Introductions and Computer Vision Backgrounds Swix mistakenly assumes SAM 3 added a 3D dimension based on the name, prompting Nikhila to immediately correct him that SAM 3, SAM 3 Objects, and SAM 3 Body are separate models. Joseph Nelson highlights Roboflow's custom RF-DETR real-time edge transformer model, establishing co-host technical credibility.5:31–10:51 · Guest disagreement 1/10 Live Demonstration of SAM 3 Concept Prompting Nikhila presents a live demo of concept prompting and video tracking, while Swix and Joseph drill down on real-time inference latency and parallel multi-GPU batching mechanics described in the paper.10:51–13:32 · Guest disagreement 1/10 Concept Segmentation and the SA-Co Benchmark Evolution The hosts inquire about how text prompting evolved from a prototype in SAM 2 to a full 200k+ concept benchmark in SA-Co. Nikhila explains the architectural shift from open-ended natural language to granular, atomic visual concepts.13:38–18:28 · Guest disagreement 0/10 Real-World Industrial and Scientific Deployments of SAM Joseph details Roboflow's production analytics, citing over 106 million SAM-assisted annotations across biology, aerial mapping, and industrial robotics. Nikhila welcomes the real-world validation data as the ultimate benchmark beyond synthetic test sets.18:28–24:02 · Guest disagreement 1/10 Domain Fine-Tuning, Negative Examples, and Presence Tokens Swix questions the balance between positive and negative training examples, prompting Nikhila to reveal that over 70% of annotations in SAM 3 are negative phrases handled via an explicit learned presence token separating recognition from localization.24:02–28:11 · Guest disagreement 1/10 Architectural Decoupling of Visual Detection and Tracking Swix highlights the architecture diagram's newly introduced components. Nikhila explains why decoupling the identity-agnostic detector from the identity-preserving tracker resolved fundamental task conflicts during video segmentation.28:11–37:57 · Guest disagreement 1/10 SAM 3 as Multimodal Agent and Live Benchmarks Swix probes the necessity of SAM tool-calling over native MLLM grounding capabilities, referencing Table 8 benchmark metrics. Joseph runs live comparisons against Gemini 3 Pro and Florence-2, demonstrating SAM 3's speed and dense mask fidelity on occluded targets.37:57–43:03 · Guest disagreement 1/10 Automated Data Engine and Superhuman AI Verification Pengchuan breaks down SAM 3's multi-stage data engine, explaining how fine-tuning Llama 3.2 on verification tasks achieved superhuman precision, dropping human-in-the-loop per-datapoint annotation time from 2 minutes down to 25 seconds.43:04–47:36 · Guest disagreement 2/10 Surpassing Human Performance and Video Pipeline Bottlenecks Swix asks what happens when training runs out of human annotators. Pengchuan cautions against blind optimism, arguing that surpassing human performance requires a transition from supervised imitation learning to RLHF-style visual preference optimization.47:36–51:05 · Guest disagreement 1/10 Video Temporal Smoothing and Identity Tracking Nuances Swix brings up the masklet detection score and asks why other temporal models fail to smooth across time windows. Pengchuan explains the core engineering trade-off between live streaming latency and accumulated temporal context.51:06–56:47 · Guest disagreement 2/10 Native Multimodal Perception Versus Tool Calling for AGI Joseph poses the core architectural debate between native multimodal visual grounding and external modular tool calling. Pengchuan argues that basic perception like counting must become native System-1 cognition, with tools reserved for dense multi-step reasoning.56:47–1:03:10 · Guest disagreement 1/10 Open Source Vision Ecosystem and Future Roadmaps The conversation covers open source contributions and the future roadmap for vision models. Pengchuan identifies end-to-end video training and smaller edge models as top priorities, while Swix questions how perception models interface with explicit robotics world models.1:03:10–1:12:07 · Guest disagreement 2/10 Roboflow Deployment, Auto-Labeling, and Human Intent Alignment Swix challenges the premise of scalar confidence sliders for handling ambiguous visual concepts like reflections, arguing that richer iterative prompting is required. Pengchuan and Joseph agree, illustrating how subjective human intent necessitates interactive refinement and fine-tuning.0:03–5:31 · The hosts pushing back 1/10 Introductions and Computer Vision Backgrounds Swix mistakenly assumes SAM 3 added a 3D dimension based on the name, prompting Nikhila to immediately correct him that SAM 3, SAM 3 Objects, and SAM 3 Body are separate models. Joseph Nelson highlights Roboflow's custom RF-DETR real-time edge transformer model, establishing co-host technical credibility.5:31–10:51 · The hosts pushing back 1/10 Live Demonstration of SAM 3 Concept Prompting Nikhila presents a live demo of concept prompting and video tracking, while Swix and Joseph drill down on real-time inference latency and parallel multi-GPU batching mechanics described in the paper.10:51–13:32 · The hosts pushing back 1/10 Concept Segmentation and the SA-Co Benchmark Evolution The hosts inquire about how text prompting evolved from a prototype in SAM 2 to a full 200k+ concept benchmark in SA-Co. Nikhila explains the architectural shift from open-ended natural language to granular, atomic visual concepts.13:38–18:28 · The hosts pushing back 0/10 Real-World Industrial and Scientific Deployments of SAM Joseph details Roboflow's production analytics, citing over 106 million SAM-assisted annotations across biology, aerial mapping, and industrial robotics. Nikhila welcomes the real-world validation data as the ultimate benchmark beyond synthetic test sets.18:28–24:02 · The hosts pushing back 2/10 Domain Fine-Tuning, Negative Examples, and Presence Tokens Swix questions the balance between positive and negative training examples, prompting Nikhila to reveal that over 70% of annotations in SAM 3 are negative phrases handled via an explicit learned presence token separating recognition from localization.24:02–28:11 · The hosts pushing back 1/10 Architectural Decoupling of Visual Detection and Tracking Swix highlights the architecture diagram's newly introduced components. Nikhila explains why decoupling the identity-agnostic detector from the identity-preserving tracker resolved fundamental task conflicts during video segmentation.28:11–37:57 · The hosts pushing back 2/10 SAM 3 as Multimodal Agent and Live Benchmarks Swix probes the necessity of SAM tool-calling over native MLLM grounding capabilities, referencing Table 8 benchmark metrics. Joseph runs live comparisons against Gemini 3 Pro and Florence-2, demonstrating SAM 3's speed and dense mask fidelity on occluded targets.37:57–43:03 · The hosts pushing back 1/10 Automated Data Engine and Superhuman AI Verification Pengchuan breaks down SAM 3's multi-stage data engine, explaining how fine-tuning Llama 3.2 on verification tasks achieved superhuman precision, dropping human-in-the-loop per-datapoint annotation time from 2 minutes down to 25 seconds.43:04–47:36 · The hosts pushing back 1/10 Surpassing Human Performance and Video Pipeline Bottlenecks Swix asks what happens when training runs out of human annotators. Pengchuan cautions against blind optimism, arguing that surpassing human performance requires a transition from supervised imitation learning to RLHF-style visual preference optimization.47:36–51:05 · The hosts pushing back 2/10 Video Temporal Smoothing and Identity Tracking Nuances Swix brings up the masklet detection score and asks why other temporal models fail to smooth across time windows. Pengchuan explains the core engineering trade-off between live streaming latency and accumulated temporal context.51:06–56:47 · The hosts pushing back 1/10 Native Multimodal Perception Versus Tool Calling for AGI Joseph poses the core architectural debate between native multimodal visual grounding and external modular tool calling. Pengchuan argues that basic perception like counting must become native System-1 cognition, with tools reserved for dense multi-step reasoning.56:47–1:03:10 · The hosts pushing back 1/10 Open Source Vision Ecosystem and Future Roadmaps The conversation covers open source contributions and the future roadmap for vision models. Pengchuan identifies end-to-end video training and smaller edge models as top priorities, while Swix questions how perception models interface with explicit robotics world models.1:03:10–1:12:07 · The hosts pushing back 3/10 Roboflow Deployment, Auto-Labeling, and Human Intent Alignment Swix challenges the premise of scalar confidence sliders for handling ambiguous visual concepts like reflections, arguing that richer iterative prompting is required. Pengchuan and Joseph agree, illustrating how subjective human intent necessitates interactive refinement and fine-tuning.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 0%1:15:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 43:20 Pengchuan rejects automated data optimism

Pengchuan directly pushes back on Swix's assumption of fully automated superhuman data generation, stressing that supervised fine-tuning hits a ceiling and requires RLHF-style preference modeling.

Hardest push from the hosts ▶ 1:07:42 Swix challenges confidence sliders for visual concept labeling

Swix refuses the framing that simple confidence thresholds can resolve nuanced labeling ambiguities like reflections, demanding iterative prompting instead.

Biggest teaching moment ▶ 0:59 Nikhila clarifies SAM 3 model taxonomy

Nikhila directly corrects Swix's misunderstanding that SAM 3 is a 3D model, clarifying that SAM 3 is the image/video foundation model while Objects and Body are separate releases.

The host holds their own ▶ 31:24 Joseph demonstrates live benchmark edge over Gemini and Florence

Joseph leverages Roboflow's live benchmarking infrastructure to show SAM 3 outperforming Gemini 3 Pro and Florence-2 on speed, segmentation granularity, and occluded objects.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and Computer Vision Backgrounds 5421 Swix mistakenly assumes SAM 3 added a 3D dimension based on the name, prompting Nikhila to immediately correct him that SAM 3, SAM 3 Objects, and SAM 3 Body are separate models. Joseph Nelson highlights Roboflow's custom RF-DETR real-time edge transformer model, establishing co-host technical credibility.
Live Demonstration of SAM 3 Concept Prompting 6211 Nikhila presents a live demo of concept prompting and video tracking, while Swix and Joseph drill down on real-time inference latency and parallel multi-GPU batching mechanics described in the paper.
Concept Segmentation and the SA-Co Benchmark Evolution 5311 The hosts inquire about how text prompting evolved from a prototype in SAM 2 to a full 200k+ concept benchmark in SA-Co. Nikhila explains the architectural shift from open-ended natural language to granular, atomic visual concepts.
Real-World Industrial and Scientific Deployments of SAM 6100 Joseph details Roboflow's production analytics, citing over 106 million SAM-assisted annotations across biology, aerial mapping, and industrial robotics. Nikhila welcomes the real-world validation data as the ultimate benchmark beyond synthetic test sets.
Domain Fine-Tuning, Negative Examples, and Presence Tokens 6412 Swix questions the balance between positive and negative training examples, prompting Nikhila to reveal that over 70% of annotations in SAM 3 are negative phrases handled via an explicit learned presence token separating recognition from localization.
Architectural Decoupling of Visual Detection and Tracking 6311 Swix highlights the architecture diagram's newly introduced components. Nikhila explains why decoupling the identity-agnostic detector from the identity-preserving tracker resolved fundamental task conflicts during video segmentation.
SAM 3 as Multimodal Agent and Live Benchmarks 7312 Swix probes the necessity of SAM tool-calling over native MLLM grounding capabilities, referencing Table 8 benchmark metrics. Joseph runs live comparisons against Gemini 3 Pro and Florence-2, demonstrating SAM 3's speed and dense mask fidelity on occluded targets.
Automated Data Engine and Superhuman AI Verification 5511 Pengchuan breaks down SAM 3's multi-stage data engine, explaining how fine-tuning Llama 3.2 on verification tasks achieved superhuman precision, dropping human-in-the-loop per-datapoint annotation time from 2 minutes down to 25 seconds.
Surpassing Human Performance and Video Pipeline Bottlenecks 6421 Swix asks what happens when training runs out of human annotators. Pengchuan cautions against blind optimism, arguing that surpassing human performance requires a transition from supervised imitation learning to RLHF-style visual preference optimization.
Video Temporal Smoothing and Identity Tracking Nuances 6412 Swix brings up the masklet detection score and asks why other temporal models fail to smooth across time windows. Pengchuan explains the core engineering trade-off between live streaming latency and accumulated temporal context.
Native Multimodal Perception Versus Tool Calling for AGI 6321 Joseph poses the core architectural debate between native multimodal visual grounding and external modular tool calling. Pengchuan argues that basic perception like counting must become native System-1 cognition, with tools reserved for dense multi-step reasoning.
Open Source Vision Ecosystem and Future Roadmaps 5311 The conversation covers open source contributions and the future roadmap for vision models. Pengchuan identifies end-to-end video training and smaller edge models as top priorities, while Swix questions how perception models interface with explicit robotics world models.
Roboflow Deployment, Auto-Labeling, and Human Intent Alignment 7323 Swix challenges the premise of scalar confidence sliders for handling ambiguous visual concepts like reflections, arguing that richer iterative prompting is required. Pengchuan and Joseph agree, illustrating how subjective human intent necessitates interactive refinement and fine-tuning.

Statements from this episode (30)

Assertion Supported
Ravi: Meta launched three separate SAM 3 models, not just one
“We launched actually three separate models this time. It was SAM-III, SAM-III objects, and SAM-III body. Those were two completely separate models and SAM III is just the image and video understanding model.”
Nikhila Ravi Dec 18, 2025 ▶ 1:08
Assertion Not checkable as stated
Nelson: Millions of developers and half the Fortune 100 use Roboflow
“Now, millions of developers, half the Fortune 100, build with RoboFlow's tools and infrastructure to create and deploy models to production.”
Joseph Nelson Dec 18, 2025 ▶ 3:44
Assertion Supported
Ravi: SAM 3 concept prompting eliminates manual per-instance clicking
“Essentially, idea of a concept prompt opens up the ability to find all instances of an object category without having to manually click on every single instance, as you would have had to do if you were using SAM-II or SAM-I.”
Nikhila Ravi Dec 18, 2025 ▶ 6:21
Assertion Supported
Zhang: SAM 3 achieves real-time tracking across objects via multi-GPU parallelism
“Even for video, if you can't afford the kind of GPUs, pretty many, very kind of, do the kind of parallel inference algorithm. So even you have a lot of object to track, you can still get real-time tracking performance as long as you scale up the GPUs there.”
Pengchuan Zhang Dec 18, 2025 ▶ 9:50
Assertion Supported
Ravi: Meta's SA-Co Benchmark Has Over 200,000 Unique Concepts
“If you look at the size of these benchmarks, the previous benchmark, Peng Chuan mentioned, Elvis, that everyone uses, it has about 1.2 K unique concepts and the benchmark that we created, which we're calling segment anything with concepts or Seiko, COCO for sh…”
Nikhila Ravi Dec 18, 2025 ▶ 12:13
Insight
Ravi: True AI Advantage Comes From Data Engines Rather Than Models
“That's the advantage in AIs is not just about the models, but really about the data and maybe even more so is actually the data engine to generate that data.”
Nikhila Ravi Dec 18, 2025 ▶ 13:14
Assertion Not checkable as stated
Nelson: Roboflow Records 106 Million Annotations Powered by Meta SAM
“And we've seen basically a hundred and six million kind of smart poly created examples that are SAM one, two, or three powered.”
Joseph Nelson Dec 18, 2025 ▶ 14:30
Assertion Not checkable as stated
Nelson: Users Ran 8 Million Inferences in SAM 3's First Five Days
“I mean, in the first five days of SAM three, there was like eight million inferences of folks that were running across all diverse sets of fields.”
Joseph Nelson Dec 18, 2025 ▶ 16:57
Insight
Nelson: A Single Negative Example Goes a Long Way in Vision Fine-Tuning
“I can offer anecdotally that a single negative example goes a long way.”
Joseph Nelson Dec 18, 2025 ▶ 20:32
Assertion Supported
Ravi: Over 70% of SAM 3 Dataset Annotations Are Negative Phrases
“We have about 70, more than 70% of the annotations are these like negative phrases that are not present in the image.”
Nikhila Ravi Dec 18, 2025 ▶ 23:02
Assertion Supported
Ravi: SAM 3 Uses a Presence Token to Separate Recognition from Localization
“We basically add this presence token to the model, which explicitly separates the task of recognition and localization.”
Nikhila Ravi Dec 18, 2025 ▶ 23:25
Insight
Ravi: Detection and tracking must decouple due to conflicting representation needs
“The detector needs to be identity agnostic. So if you have a concept dog, it needs to be able to find all instances of that dog. And it needs to sort of have this representation of dog that is the same for all dogs. But when you're tracking those dogs through …”
Nikhila Ravi Dec 18, 2025 ▶ 25:31
Assertion Supported
Ravi: SAM 3 tracking compute scales with detected objects, not classes
“Each of the, it scales with the number of detected objects.”
Nikhila Ravi Dec 18, 2025 ▶ 27:24
Assertion Not publicly verifiable
Nelson: Roboflow users overwhelmingly voted for SAM 3 in blind pre-release testing
“That's actually interesting because we had blind tested SAM three before it was released, not a SAM three, just for people to try and compare. I think we call it like a potential SAG or SAG preview or something. And we allowed users to vote and they kind of un…”
Joseph Nelson Dec 18, 2025 ▶ 33:49
Disclosure
Zhang: Meta intentionally avoided OCR-heavy images during SAM 3 training data sampling
“In fact, during our data engine, we intentionally do not sample OCR-heavy images.”
Pengchuan Zhang Dec 18, 2025 ▶ 37:11
Assertion Partly supported
Zhang: Fine-Tuned Llama 3.2 Achieved Superhuman Vision Verification Performance
“We kind of fine-tune our, kind of, for example, NAMA's 3.2 with our, kind of, verification, human annotated verification data. We get, kind of, superhuman performance on these two verification tasks, and then we do not need human on these two tasks. Let's furt…”
Pengchuan Zhang Dec 18, 2025 ▶ 41:14
Insight
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan Zhang Dec 18, 2025 ▶ 44:33
Assertion Supported
Ravi: Meta Achieved Fully Automated Annotation in SAM 1, Not SAM 2
“Getting to that fully automated data engine is something that we tried to do in SAM too. We actually didn't get to that fully automated approach. In SAM one, we did, we, you know, But the SA-I-B dataset that we released was fully annotated automatically. We di…”
Nikhila Ravi Dec 18, 2025 ▶ 45:38
Insight
Zhang: Video tracking requires trading streaming latency for temporal accuracy
“So there is a trade off between kind of the kind of the latency and the accuracy here. If you care more about accuracy, then you can use kind of this kind of overall kind of information can all cause the mass net. To get kind of more robust signal about the co…”
Pengchuan Zhang Dec 18, 2025 ▶ 49:32
Insight
Ravi: Many video CV workflows require per-frame detection, not identity tracking
“I think also in many video use cases, I think, because if you were sharing on RoboFlow, Users care more about detecting the objects rather than having unique identities. So in, in some cases this, maybe it's, this isn't required to preserve the identities thro…”
Nikhila Ravi Dec 18, 2025 ▶ 49:58
Assertion Supported
Ravi: SAM 3 matches or beats single-task vision SOTA models
“We are really having a unified model that can do many different tasks in the same unified architecture. And so, you know, then the same way that LLMs can do many different tasks without needing a task-specific model. Like with SAM-III, we're able to do image-p…”
Nikhila Ravi Dec 18, 2025 ▶ 51:16
Prediction Not checkable as stated
Zhang: AI models will handle simple vision natively, using tools for complexity
“I think at least I want to bet on, you know, running their work natively together, the future for simple, I would say for simple or even intermediate difficult vision tasks. For example, kind of counting with less than 20 objects. I think for this kind of simp…”
Pengchuan Zhang Dec 18, 2025 ▶ 54:11
Disclosure
Ravi: Meta built SAM 3 using open-source community contributions to SAM 2
“In SAM-III we did leverage many of the open source contributions people have made on top of SAM-II. There were new data sets, there were new benchmarks, There were new kind of inference time optimizations. We adopt a lot of the things that the community builds…”
Nikhila Ravi Dec 18, 2025 ▶ 57:04
Assertion Not checkable as stated
Zhang: Video AI models still exhibit a large gap versus human performance
“Video is still far from, I would say, have a big gap from human performance. Right now, there's kind of still, kind of, a lot of research needs to be done there, how to do end-to-end training with video.”
Pengchuan Zhang Dec 18, 2025 ▶ 59:12
Disclosure
Zhang: SAM uses decoupled training for video instead of end-to-end
“We do not have, kind of, we have this kind of decoupled approach, but we do not end-to-end train this model. And we expect definitely kind of, it will be kind of a benefit from kind of end-to-end training.”
Pengchuan Zhang Dec 18, 2025 ▶ 59:27
Prediction Not checkable as stated
Nelson: SAM 3 Open Text Boxes Will Trigger Many Unprimed User Queries
“Now that you've kind of given this open text box for media, there's going to be a flood of the types of things users are going to want to try to do, some of which SAM is already going to be really well adapted to do, some of which not.”
Joseph Nelson Dec 18, 2025 ▶ 1:02:55
Assertion Supported
Nelson: MedSAM-3 was released within one week of SAM 3
“Within a week of releasing SAM three med SAM three came out for adapting SAM into medical contexts.”
Joseph Nelson Dec 18, 2025 ▶ 1:04:18
Assertion Not checkable as stated
Nelson: Users have created hundreds of fine-tunes of SAM 3
“And I think we're already beginning to see that with hundreds of fine tunes that users are creating for various domains.”
Joseph Nelson Dec 18, 2025 ▶ 1:04:42
Insight
Nelson: Last-mile computer vision requires aligning intention, not knowledge
“And so this is why, like, in some ways human in the loop, because identifying human intention, not necessarily human knowledge is what's going to be important for a lot of last mile use.”
Joseph Nelson Dec 18, 2025 ▶ 1:10:31
Prediction Not checkable as stated
Zhang: SA-Co Benchmark Will Likely Outlast SAM 3
“I would say that it's likely that the benchmark will last longer than our Samsung model. Maybe kind of next year there will be a stronger model, but the benchmark is kind of the one that I hope to guide the community to kind of get better and better models kin…”
Pengchuan Zhang Dec 18, 2025 ▶ 1:13:23
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.