Dec 22, 2024 · 55m · latent-space

Best of 2024 in Vision [LS Live @ NeurIPS]

Isaac Robinson · 17m spoken Peter Robicheaux · 17m spoken Vik Korupati · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Presented live at NeurIPS, this technical retrospective highlights the defining computer vision breakthroughs of 2024, analyzing generative video architectures, the ascendance of real-time Detection Transformers over YOLO, methods for resolving vision-language model perceptual blindness, and compact edge models like Moondream.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 0.3 Guest teaching 0.3 Guest disagreement 0.1 The hosts pushing back 0.0
05100:0015:0030:0045:000:06–7:53 · The hosts as informed peer 0/10 Overview of Major Computer Vision Trends in 2024 Isaac presents a solo technical lecture on video generation models and Sora replications without host dialogue. Host-side engagement and pushback are absent during this monologue.7:55–14:47 · The hosts as informed peer 0/10 SAM 2 and Temporal Memory for Promptable Video Segmentation Isaac continues his presentation detailing SAM 2 architecture, hierarchical encoders, and temporal memory banks. There is no host involvement or disagreement.14:47–20:18 · The hosts as informed peer 0/10 The Ascendance of DETRs Over YOLO in Real-Time Object Detection Isaac presents technical benchmarks comparing real-time DETRs to YOLO architectures. The segment is entirely a presentation monologue without host participation.20:20–27:08 · The hosts as informed peer 0/10 Investigating VLM Blindness and Fine-Grained Perception with MMVP Peter begins his talk with the provocative thesis that VLMs cannot see fine details, explaining the MMVP benchmark. It is an instructional presentation delivered to an audience without conversational pushback.27:09–38:42 · The hosts as informed peer 0/10 Architectures for Spatial and Semantic Granularity: Florence-2, PaliGemma 2, and AIMv2 Peter walks through Florence-2, PaliGemma 2, and AIMv2 benchmark results in an educational format. Host metrics remain zero due to the lecture setting.38:47–42:13 · The hosts as informed peer 1/10 Audience Q&A on Foundation Models Versus Specialist Object Detectors The host opens the floor to audience questions, prompting Isaac and Peter to explain why generalist VLMs lag behind specialist object detectors. The guests collaboratively break down domain-specific architectures and pre-training dynamics.42:14–47:00 · The hosts as informed peer 0/10 Moondream: Compact Vision-Language Models and Sub-Billion Edge Pruning Vik presents Moondream's architecture and sub-billion edge pruning methodology. The segment proceeds as a solo technical demo and presentation.47:00–54:38 · The hosts as informed peer 1/10 Grounded Chain-of-Thought for Analog Gauge Reading and Visual Reasoning Vik explains grounded chain-of-thought for gauge reading before the host offers a brief concluding thought on vision modalities. The dynamic is entirely collaborative and educational.0:06–7:53 · Guest teaching 0/10 Overview of Major Computer Vision Trends in 2024 Isaac presents a solo technical lecture on video generation models and Sora replications without host dialogue. Host-side engagement and pushback are absent during this monologue.7:55–14:47 · Guest teaching 0/10 SAM 2 and Temporal Memory for Promptable Video Segmentation Isaac continues his presentation detailing SAM 2 architecture, hierarchical encoders, and temporal memory banks. There is no host involvement or disagreement.14:47–20:18 · Guest teaching 0/10 The Ascendance of DETRs Over YOLO in Real-Time Object Detection Isaac presents technical benchmarks comparing real-time DETRs to YOLO architectures. The segment is entirely a presentation monologue without host participation.20:20–27:08 · Guest teaching 0/10 Investigating VLM Blindness and Fine-Grained Perception with MMVP Peter begins his talk with the provocative thesis that VLMs cannot see fine details, explaining the MMVP benchmark. It is an instructional presentation delivered to an audience without conversational pushback.27:09–38:42 · Guest teaching 0/10 Architectures for Spatial and Semantic Granularity: Florence-2, PaliGemma 2, and AIMv2 Peter walks through Florence-2, PaliGemma 2, and AIMv2 benchmark results in an educational format. Host metrics remain zero due to the lecture setting.38:47–42:13 · Guest teaching 2/10 Audience Q&A on Foundation Models Versus Specialist Object Detectors The host opens the floor to audience questions, prompting Isaac and Peter to explain why generalist VLMs lag behind specialist object detectors. The guests collaboratively break down domain-specific architectures and pre-training dynamics.42:14–47:00 · Guest teaching 0/10 Moondream: Compact Vision-Language Models and Sub-Billion Edge Pruning Vik presents Moondream's architecture and sub-billion edge pruning methodology. The segment proceeds as a solo technical demo and presentation.47:00–54:38 · Guest teaching 0/10 Grounded Chain-of-Thought for Analog Gauge Reading and Visual Reasoning Vik explains grounded chain-of-thought for gauge reading before the host offers a brief concluding thought on vision modalities. The dynamic is entirely collaborative and educational.0:06–7:53 · Guest disagreement 0/10 Overview of Major Computer Vision Trends in 2024 Isaac presents a solo technical lecture on video generation models and Sora replications without host dialogue. Host-side engagement and pushback are absent during this monologue.7:55–14:47 · Guest disagreement 0/10 SAM 2 and Temporal Memory for Promptable Video Segmentation Isaac continues his presentation detailing SAM 2 architecture, hierarchical encoders, and temporal memory banks. There is no host involvement or disagreement.14:47–20:18 · Guest disagreement 0/10 The Ascendance of DETRs Over YOLO in Real-Time Object Detection Isaac presents technical benchmarks comparing real-time DETRs to YOLO architectures. The segment is entirely a presentation monologue without host participation.20:20–27:08 · Guest disagreement 1/10 Investigating VLM Blindness and Fine-Grained Perception with MMVP Peter begins his talk with the provocative thesis that VLMs cannot see fine details, explaining the MMVP benchmark. It is an instructional presentation delivered to an audience without conversational pushback.27:09–38:42 · Guest disagreement 0/10 Architectures for Spatial and Semantic Granularity: Florence-2, PaliGemma 2, and AIMv2 Peter walks through Florence-2, PaliGemma 2, and AIMv2 benchmark results in an educational format. Host metrics remain zero due to the lecture setting.38:47–42:13 · Guest disagreement 0/10 Audience Q&A on Foundation Models Versus Specialist Object Detectors The host opens the floor to audience questions, prompting Isaac and Peter to explain why generalist VLMs lag behind specialist object detectors. The guests collaboratively break down domain-specific architectures and pre-training dynamics.42:14–47:00 · Guest disagreement 0/10 Moondream: Compact Vision-Language Models and Sub-Billion Edge Pruning Vik presents Moondream's architecture and sub-billion edge pruning methodology. The segment proceeds as a solo technical demo and presentation.47:00–54:38 · Guest disagreement 0/10 Grounded Chain-of-Thought for Analog Gauge Reading and Visual Reasoning Vik explains grounded chain-of-thought for gauge reading before the host offers a brief concluding thought on vision modalities. The dynamic is entirely collaborative and educational.0:06–7:53 · The hosts pushing back 0/10 Overview of Major Computer Vision Trends in 2024 Isaac presents a solo technical lecture on video generation models and Sora replications without host dialogue. Host-side engagement and pushback are absent during this monologue.7:55–14:47 · The hosts pushing back 0/10 SAM 2 and Temporal Memory for Promptable Video Segmentation Isaac continues his presentation detailing SAM 2 architecture, hierarchical encoders, and temporal memory banks. There is no host involvement or disagreement.14:47–20:18 · The hosts pushing back 0/10 The Ascendance of DETRs Over YOLO in Real-Time Object Detection Isaac presents technical benchmarks comparing real-time DETRs to YOLO architectures. The segment is entirely a presentation monologue without host participation.20:20–27:08 · The hosts pushing back 0/10 Investigating VLM Blindness and Fine-Grained Perception with MMVP Peter begins his talk with the provocative thesis that VLMs cannot see fine details, explaining the MMVP benchmark. It is an instructional presentation delivered to an audience without conversational pushback.27:09–38:42 · The hosts pushing back 0/10 Architectures for Spatial and Semantic Granularity: Florence-2, PaliGemma 2, and AIMv2 Peter walks through Florence-2, PaliGemma 2, and AIMv2 benchmark results in an educational format. Host metrics remain zero due to the lecture setting.38:47–42:13 · The hosts pushing back 0/10 Audience Q&A on Foundation Models Versus Specialist Object Detectors The host opens the floor to audience questions, prompting Isaac and Peter to explain why generalist VLMs lag behind specialist object detectors. The guests collaboratively break down domain-specific architectures and pre-training dynamics.42:14–47:00 · The hosts pushing back 0/10 Moondream: Compact Vision-Language Models and Sub-Billion Edge Pruning Vik presents Moondream's architecture and sub-billion edge pruning methodology. The segment proceeds as a solo technical demo and presentation.47:00–54:38 · The hosts pushing back 0/10 Grounded Chain-of-Thought for Analog Gauge Reading and Visual Reasoning Vik explains grounded chain-of-thought for gauge reading before the host offers a brief concluding thought on vision modalities. The dynamic is entirely collaborative and educational.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 20:30 LLMs literally cannot see

Peter takes a strong contrarian stance against general VLM hype by proving top commercial models fail on simple visual clock tasks.

Hardest push from the hosts ▶ 39:10 Audience pushes on generalist model failure

An audience participant directly questions why major frontier lab models remain inferior to real-time DETRs at object detection.

Biggest teaching moment ▶ 40:06 Isaac unpacks detection architecture limits

Isaac educates the room on the domain-specific nuances of object detection architectures and historical pre-training deficits in convolutional models.

The host holds their own ▶ 54:44 Host highlights vision as first among equals

Host Alessio synthesizes the presentations by emphasizing vision's preeminent importance across multimodal AI paradigms.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Overview of Major Computer Vision Trends in 2024 0000 Isaac presents a solo technical lecture on video generation models and Sora replications without host dialogue. Host-side engagement and pushback are absent during this monologue.
SAM 2 and Temporal Memory for Promptable Video Segmentation 0000 Isaac continues his presentation detailing SAM 2 architecture, hierarchical encoders, and temporal memory banks. There is no host involvement or disagreement.
The Ascendance of DETRs Over YOLO in Real-Time Object Detection 0000 Isaac presents technical benchmarks comparing real-time DETRs to YOLO architectures. The segment is entirely a presentation monologue without host participation.
Investigating VLM Blindness and Fine-Grained Perception with MMVP 0010 Peter begins his talk with the provocative thesis that VLMs cannot see fine details, explaining the MMVP benchmark. It is an instructional presentation delivered to an audience without conversational pushback.
Architectures for Spatial and Semantic Granularity: Florence-2, PaliGemma 2, and AIMv2 0000 Peter walks through Florence-2, PaliGemma 2, and AIMv2 benchmark results in an educational format. Host metrics remain zero due to the lecture setting.
Audience Q&A on Foundation Models Versus Specialist Object Detectors 1200 The host opens the floor to audience questions, prompting Isaac and Peter to explain why generalist VLMs lag behind specialist object detectors. The guests collaboratively break down domain-specific architectures and pre-training dynamics.
Moondream: Compact Vision-Language Models and Sub-Billion Edge Pruning 0000 Vik presents Moondream's architecture and sub-billion edge pruning methodology. The segment proceeds as a solo technical demo and presentation.
Grounded Chain-of-Thought for Analog Gauge Reading and Visual Reasoning 1000 Vik explains grounded chain-of-thought for gauge reading before the host offers a brief concluding thought on vision modalities. The dynamic is entirely collaborative and educational.

Statements from this episode (23)

Assertion Not checkable as stated
Robinson: DETRs are replacing YOLO in real-time object detection
“And then also how debtors are starting to take over the real time object detection scene from YOLOs which have been dominant for years.”
Isaac Robinson Dec 22, 2024 ▶ 0:49
Assertion Partly supported
Robinson: 2024 DETR improvements are Pareto superior to YOLO
“And then how debtors are the improvements in 2024 to debtors that are making them a Pareto improvement to yellow base models.”
Isaac Robinson Dec 22, 2024 ▶ 1:28
Assertion Supported
Robinson: High-Performance Diffusion Models Are Shifting to Rectified Flows
“It's also it's also worth noting that most diffusion models today, the very high performance ones are switching away from the classic like DDPM, Denoising Diffusion Probability Modeling Framework to rectified flows.”
Isaac Robinson Dec 22, 2024 ▶ 5:39
Insight
Robinson: Rectified Flows Enable Faster Sampling by Approaching Single-Step Inference
“Rectified flows have a very interesting property of that. As they converge, they actually get closer to being able to be sampled with a single step, which means that in practice, you can actually generate high quality samples much faster.”
Isaac Robinson Dec 22, 2024 ▶ 5:54
Assertion Supported
Robinson: Original DiT Research Showed Compute Scaling Outweighed Hyperparameters
“This is so interesting because the original diffusion transformer paper from Facebook actually showed that, in fact, the specific hyperparameters of the transformer didn't really matter that much. What mattered was that you were just increasing the amount of c…”
Isaac Robinson Dec 22, 2024 ▶ 6:39
Assertion Not checkable as stated
Robinson: SAM has saved Roboflow users 75 years of labeling time
“SAM for us has saved our users 75 years of labeling time.”
Isaac Robinson Dec 22, 2024 ▶ 8:05
Assertion Supported
Robinson: SAM 2's hierarchical encoder delivers 6x faster inference than ViT
“SAM replaced that with a hierarchical encoder, which gets approximately the same results, but leads to a six times faster inference, which is excellent, especially considering how in a trend of 23 was replacing the VIT with more efficient backbones.”
Isaac Robinson Dec 22, 2024 ▶ 10:32
Opinion
Robinson: YOLO architectures have hit a performance plateau
“So, for years, yellows have been the dominant way of doing real time object detection, and we can see here that they've essentially stagnated. The performance between 10 and 11 is not meaningfully different. At least, you know, in, in this type of high level c…”
Isaac Robinson Dec 22, 2024 ▶ 15:07
Assertion Supported
Robinson: D-FINE models achieve +4.6 AP over YOLO at equivalent latency
“So, we can look here and see the yellow series has this plateau and then these RT debtor, LW debtor and define have meaningfully changed that plateau so that in fact the best defined models are plus 4.6 AP on Coco at the same latency.”
Isaac Robinson Dec 22, 2024 ▶ 15:36
Assertion Supported
Robinson: Factoring in NMS Latency Shows DETRs Outperform YOLO
“Once you include the NMS in the latency calculation, you see that, in fact, these debtors are outperforming, at least at this time, the yellows that existed.”
Isaac Robinson Dec 22, 2024 ▶ 17:10
Insight
Robicheaux: CLIP encoders lack fine-grained visual features due to caption-matching objective
“Models that have been initialized with Clip as their vision encoder, they don't have fine-grained details and the features extracted using Clip because Clip sort of doesn't need to find these fine grade details to do its job correctly, which is just to match c…”
Peter Robicheaux Dec 22, 2024 ▶ 22:21
Assertion Supported
Robicheaux: Adding DINOv2 features degrades multimodal language modeling performance
“As you increase the number of Dynav two features, your model does worse and worse and worse on the actual language modeling task, and that's because Dynav two features were trained completely from a self-supervised manner and completely in image space. It know…”
Peter Robicheaux Dec 22, 2024 ▶ 25:41
Assertion Supported
Robicheaux: Florence-2 Achieves 60% mAP on COCO
“They get 60%, 60% map on Cocoa, which is, like, approaching state of the art, and they train with... You're good. And they train with a much more much more efficiently.”
Peter Robicheaux Dec 22, 2024 ▶ 29:37
Assertion Supported
Robicheaux: PaliGemma 1 Pre-Training Saturates at 300M Examples
“One of my critiques, I guess, of polygeoma one, at least, is that you find that performance saturates as a pre-trained model after only three hundred million examples seen.”
Peter Robicheaux Dec 22, 2024 ▶ 32:52
Assertion Supported
Robicheaux: 2B PaliGemma 2 Beats ChatGPT on MMVP with 47.3%
“The big result, and one of the reasons that I was really excited about this paper, is that they blow everything else away on MMVP. I mean, 47.3, sure, that's nowhere near human accuracy, which again is 94%, but for, you know, two billion language, two billion …”
Peter Robicheaux Dec 22, 2024 ▶ 34:28
Assertion Supported
Robicheaux: AIMv2 Avoids Performance Saturation as Scale Increases
“And we can see that this is finally a model that doesn't saturate. It's even at the highest parameter count, it's, it appears to be Well, at the highest parameter account, it appears to be improving in performance with more and more samples seen”
Peter Robicheaux Dec 22, 2024 ▶ 36:43
Assertion Supported
Robinson: YOLO and real-time detectors historically gained little from pre-training
“And the other thing is, until recently, the real time object detectors didn't even really benefit from pre-training. Like, you see the Yolos that are, like, essentially saturated showing very little difference with pretraining improvements with using pretraine…”
Isaac Robinson Dec 22, 2024 ▶ 41:00
Opinion
Robicheaux: Convolutional models fail to benefit from pre-training compared to transformers
“Essentially, I think it's kind of been shown now that convolution models, like, just don't benefit from pre-training and just don't, like, have the level of intelligence to transform models.”
Peter Robicheaux Dec 22, 2024 ▶ 42:00
Disclosure
Korupati: Moondream built a 0.5B model by pruning its 2B model
“The way we built our model was to start with the two billion parameter model and Prune it while doing continual training to retain performance.”
Vik Korupati Dec 22, 2024 ▶ 45:46
Insight
Korupati: Developers Can Prototype on 2B Models and Prune for Deployment
“I think the thing that's really exciting about this is it makes it possible for developers to build using the two B per model and Just explore, build their application, and then once they're ready to deploy, figure out what exactly they need out of the model a…”
Vik Korupati Dec 22, 2024 ▶ 46:36
Insight
Korupati: VLMs Fail at Gauges Due to E-Commerce Training Biases
“In the case of gauges, most gauges images aren't gauges in the wild. They're product Detail images like these where it's always set to zero. It's paired with an alt text that says something like GIVTO pressure sensor PSI zero to 30 or something. And so, the mo…”
Vik Korupati Dec 22, 2024 ▶ 48:05
Prediction Not checkable as stated
Korupati: Visual Chain-of-Thought Will Probably Generalize Across Tasks
“The real question is, is it going to generalize? Probably, like, there's some signs from text models that when you train on a broad number of tasks, it does generalize, and I'm seeing some signs with our model as well.”
Vik Korupati Dec 22, 2024 ▶ 52:10
Opinion
Korupati: Vision-Language Models Are Lagging Behind LLMs in Reasoning
“LLMs are showing enormous progress in reasoning, especially with the latest set of models that we've seen, but we're not really seeing, I have a feeling that VLMs are lagging behind, as we can see with these tasks that should be very simple for a human to do t…”
Vik Korupati Dec 22, 2024 ▶ 53:01
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.