Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 1/5

Robicheaux: Florence-2 Achieves 60% mAP on COCO

Peter Robicheaux · Best of 2024 in Vision [LS Live @ NeurIPS] · Dec 22, 2024 · at 29:37

Peter Robicheaux (Roboflow) reviews Microsoft's Florence-2 architecture at a NeurIPS 2024 retrospective on vision models.

0:00 / 0:12exact quote · 12.6s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“They get 60%, 60% map on Cocoa, which is, like, approaching state of the art, and they train with... You're good. And they train with a much more much more efficiently.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Peter Robicheaux

Assertion Supported
Robicheaux: 2B PaliGemma 2 Beats ChatGPT on MMVP with 47.3%
“The big result, and one of the reasons that I was really excited about this paper, is that they blow everything else away on MMVP. I mean, 47.3, sure, that's nowhere near human accuracy, which again is 94%, but for, you know, two billion language, two billion …”
Peter Robicheaux Dec 22, 2024 ▶ 34:28 Best of 2024 in Vision [LS Live @ NeurIPS]
Assertion Supported
Robicheaux: AIMv2 Avoids Performance Saturation as Scale Increases
“And we can see that this is finally a model that doesn't saturate. It's even at the highest parameter count, it's, it appears to be Well, at the highest parameter account, it appears to be improving in performance with more and more samples seen”
Peter Robicheaux Dec 22, 2024 ▶ 36:43 Best of 2024 in Vision [LS Live @ NeurIPS]
Opinion
Robicheaux: Convolutional models fail to benefit from pre-training compared to transformers
“Essentially, I think it's kind of been shown now that convolution models, like, just don't benefit from pre-training and just don't, like, have the level of intelligence to transform models.”
Peter Robicheaux Dec 22, 2024 ▶ 42:00 Best of 2024 in Vision [LS Live @ NeurIPS]
Insight
Robicheaux: CLIP encoders lack fine-grained visual features due to caption-matching objective
“Models that have been initialized with Clip as their vision encoder, they don't have fine-grained details and the features extracted using Clip because Clip sort of doesn't need to find these fine grade details to do its job correctly, which is just to match c…”
Peter Robicheaux Dec 22, 2024 ▶ 22:21 Best of 2024 in Vision [LS Live @ NeurIPS]
Assertion Supported
Robicheaux: Adding DINOv2 features degrades multimodal language modeling performance
“As you increase the number of Dynav two features, your model does worse and worse and worse on the actual language modeling task, and that's because Dynav two features were trained completely from a self-supervised manner and completely in image space. It know…”
Peter Robicheaux Dec 22, 2024 ▶ 25:41 Best of 2024 in Vision [LS Live @ NeurIPS]
Assertion Supported
Robicheaux: PaliGemma 1 Pre-Training Saturates at 300M Examples
“One of my critiques, I guess, of polygeoma one, at least, is that you find that performance saturates as a pre-trained model after only three hundred million examples seen.”
Peter Robicheaux Dec 22, 2024 ▶ 32:52 Best of 2024 in Vision [LS Live @ NeurIPS]
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.