Assertion Partly supported AI assessment confidence: 85% certainty 4/5 debate potential 2/5

Zhang: Fine-Tuned Llama 3.2 Achieved Superhuman Vision Verification Performance

Pengchuan Zhang · SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow) · Dec 18, 2025 · at 41:14

Pengchuan Zhang of Meta FAIR describes the SAM 3 data engine pipeline, explaining how automated AI verification replaced human annotators for mask checking.

0:00 / 0:19exact quote · 19.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“We kind of fine-tune our, kind of, for example, NAMA's 3.2 with our, kind of, verification, human annotated verification data. We get, kind of, superhuman performance on these two verification tasks, and then we do not need human on these two tasks. Let's further bring our, kind of, per data point annotation time to about, kind of, 25 seconds.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Pengchuan Zhang

Insight
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan Zhang Dec 18, 2025 ▶ 44:33 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Prediction Not checkable as stated
Zhang: AI models will handle simple vision natively, using tools for complexity
“I think at least I want to bet on, you know, running their work natively together, the future for simple, I would say for simple or even intermediate difficult vision tasks. For example, kind of counting with less than 20 objects. I think for this kind of simp…”
Pengchuan Zhang Dec 18, 2025 ▶ 54:11 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Assertion Supported
Zhang: SAM 3 achieves real-time tracking across objects via multi-GPU parallelism
“Even for video, if you can't afford the kind of GPUs, pretty many, very kind of, do the kind of parallel inference algorithm. So even you have a lot of object to track, you can still get real-time tracking performance as long as you scale up the GPUs there.”
Pengchuan Zhang Dec 18, 2025 ▶ 9:50 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Assertion Not checkable as stated
Zhang: Video AI models still exhibit a large gap versus human performance
“Video is still far from, I would say, have a big gap from human performance. Right now, there's kind of still, kind of, a lot of research needs to be done there, how to do end-to-end training with video.”
Pengchuan Zhang Dec 18, 2025 ▶ 59:12 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Disclosure
Zhang: Meta intentionally avoided OCR-heavy images during SAM 3 training data sampling
“In fact, during our data engine, we intentionally do not sample OCR-heavy images.”
Pengchuan Zhang Dec 18, 2025 ▶ 37:11 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Insight
Zhang: Video tracking requires trading streaming latency for temporal accuracy
“So there is a trade off between kind of the kind of the latency and the accuracy here. If you care more about accuracy, then you can use kind of this kind of overall kind of information can all cause the mass net. To get kind of more robust signal about the co…”
Pengchuan Zhang Dec 18, 2025 ▶ 49:32 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.