Prediction certainty 3/5 debate potential 3/5

Zhang: AI models will handle simple vision natively, using tools for complexity

Pengchuan Zhang · SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow) · Dec 18, 2025 · at 54:11

Pengchuan Zhang, AI Research Scientist at Meta FAIR, debates whether foundation models will natively integrate visual grounding or rely on tool-calling external vision models.

0:00 / 0:45exact quote · 45.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I think at least I want to bet on, you know, running their work natively together, the future for simple, I would say for simple or even intermediate difficult vision tasks. For example, kind of counting with less than 20 objects. I think for this kind of simple task, This is like system one, kind of visual reasoning with our brain. This should be, kind of, our brain, kind of, should do it by, kind of, by themselves. But with very, very difficult tasks, you can see that if we are counting, you know, maybe, kind of, thousands of objects, kind of, in the picture, so crowded, then we, kind of, even need to, kind of, draw something there. I would see that at that time, maybe we need, kind of, some extra model, kind of, for difficult tasks.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Pengchuan Zhang

Insight
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan Zhang Dec 18, 2025 ▶ 44:33 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Assertion Partly supported
Zhang: Fine-Tuned Llama 3.2 Achieved Superhuman Vision Verification Performance
“We kind of fine-tune our, kind of, for example, NAMA's 3.2 with our, kind of, verification, human annotated verification data. We get, kind of, superhuman performance on these two verification tasks, and then we do not need human on these two tasks. Let's furt…”
Pengchuan Zhang Dec 18, 2025 ▶ 41:14 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Assertion Supported
Zhang: SAM 3 achieves real-time tracking across objects via multi-GPU parallelism
“Even for video, if you can't afford the kind of GPUs, pretty many, very kind of, do the kind of parallel inference algorithm. So even you have a lot of object to track, you can still get real-time tracking performance as long as you scale up the GPUs there.”
Pengchuan Zhang Dec 18, 2025 ▶ 9:50 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Assertion Not checkable as stated
Zhang: Video AI models still exhibit a large gap versus human performance
“Video is still far from, I would say, have a big gap from human performance. Right now, there's kind of still, kind of, a lot of research needs to be done there, how to do end-to-end training with video.”
Pengchuan Zhang Dec 18, 2025 ▶ 59:12 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Disclosure
Zhang: Meta intentionally avoided OCR-heavy images during SAM 3 training data sampling
“In fact, during our data engine, we intentionally do not sample OCR-heavy images.”
Pengchuan Zhang Dec 18, 2025 ▶ 37:11 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Insight
Zhang: Video tracking requires trading streaming latency for temporal accuracy
“So there is a trade off between kind of the kind of the latency and the accuracy here. If you care more about accuracy, then you can use kind of this kind of overall kind of information can all cause the mass net. To get kind of more robust signal about the co…”
Pengchuan Zhang Dec 18, 2025 ▶ 49:32 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.