Pengchuan Zhang, AI Research Scientist at Meta FAIR, debates whether foundation models will natively integrate visual grounding or rely on tool-calling external vision models.
“I think at least I want to bet on, you know, running their work natively together, the future for simple, I would say for simple or even intermediate difficult vision tasks. For example, kind of counting with less than 20 objects. I think for this kind of simple task, This is like system one, kind of visual reasoning with our brain. This should be, kind of, our brain, kind of, should do it by, kind of, by themselves. But with very, very difficult tasks, you can see that if we are counting, you know, maybe, kind of, thousands of objects, kind of, in the picture, so crowded, then we, kind of, even need to, kind of, draw something there. I would see that at that time, maybe we need, kind of, some extra model, kind of, for difficult tasks.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Pengchuan Zhang
Insight
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan ZhangDec 18, 2025▶ 44:33SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
“We kind of fine-tune our, kind of, for example, NAMA's 3.2 with our, kind of, verification, human annotated verification data. We get, kind of, superhuman performance on these two verification tasks, and then we do not need human on these two tasks. Let's furt…”
Pengchuan ZhangDec 18, 2025▶ 41:14SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
AssertionSupported
Zhang: SAM 3 achieves real-time tracking across objects via multi-GPU parallelism
“Even for video, if you can't afford the kind of GPUs, pretty many, very kind of, do the kind of parallel inference algorithm. So even you have a lot of object to track, you can still get real-time tracking performance as long as you scale up the GPUs there.”
Pengchuan ZhangDec 18, 2025▶ 9:50SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
AssertionNot checkable as stated
Zhang: Video AI models still exhibit a large gap versus human performance
“Video is still far from, I would say, have a big gap from human performance. Right now, there's kind of still, kind of, a lot of research needs to be done there, how to do end-to-end training with video.”
Pengchuan ZhangDec 18, 2025▶ 59:12SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Disclosure
Zhang: Meta intentionally avoided OCR-heavy images during SAM 3 training data sampling
“In fact, during our data engine, we intentionally do not sample OCR-heavy images.”
Pengchuan ZhangDec 18, 2025▶ 37:11SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Insight
Zhang: Video tracking requires trading streaming latency for temporal accuracy
“So there is a trade off between kind of the kind of the latency and the accuracy here. If you care more about accuracy, then you can use kind of this kind of overall kind of information can all cause the mass net. To get kind of more robust signal about the co…”
Pengchuan ZhangDec 18, 2025▶ 49:32SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.