why aren't all 10 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Unified models enable faster error corrections via refinement clicks
“And we found that, you know, going from each phase, it both improved the efficiency and it improved the data quality. And in particular, when you get rid of this two-part model, one of the advantages is that when you make refinement clicks, so You prompt the m…”
Assertion Supported
Ravi: SAM 2 runs roughly six times faster on video than SAM 1
“And in terms of the efficiency compared to SAM, so if we were to run SAM per frame on a video or run SAM two, it's around six times faster to run SAM two versus run SAM per frame.”
Insight
Unified vision models outperform composite multi-model pipelines
“Combining two models and sort of just smushing things together might not actually be as effective as if you really think about how to build things in a unified way.”
Insight
Ravi: Video segmentation requires far less context than language models
“A difference to language models is probably the amount of context that you need is significantly less than maintaining a long multi-term conversation. And so, you know, coupling this short-term spatial memory with this, like, longer-term object pointers we fou…”
Prediction Held up
Ravi: SAM 2 will soon run on-device and inside web browsers
“Like, I'm pretty sure soon we'll see like an on-device SAM-II or, you know, maybe even running in the browser or something. So I think that could definitely unlock some of these edge use cases.”
Disclosure
Ravi: Meta built SAM 3 using open-source community contributions to SAM 2
“In SAM-III we did leverage many of the open source contributions people have made on top of SAM-II.
There were new data sets, there were new benchmarks,
There were new kind of inference time optimizations.
We adopt a lot of the things that the community builds…”
Assertion Supported
Ravi: Meta Achieved Fully Automated Annotation in SAM 1, Not SAM 2
“Getting to that fully automated data engine is something that we tried to do in SAM too. We actually didn't get to that fully automated approach. In SAM one, we did, we, you know, But the SA-I-B dataset that we released was fully annotated automatically. We di…”
Assertion Contradicted
Prior video object segmentation models lacked error recovery mechanisms
“That actually is a big limitation of current models, current video object segmentation models. Don't allow any way to recover if the model makes a mistake.”
Assertion Supported
Nelson: Smallest SAM 2 model has 38M parameters and runs at 45 FPS
“The smallest model is thirty-eight million parameters and can run at 45 FPS on an A-one hundred, right?”
Assertion Supported
Ravi: SAM 2's largest model is 224M parameters, one-third of SAM 1
“SAM-I model was around six hundred and thirty million parameters, a fraction of the size of these large language models, but very small. Actually SAM-II, the largest model is around two hundred and twenty-four million parameters. There's actually One third the…”