Dec 22, 2024 · 55m · latent-space
Best of 2024 in Vision [LS Live @ NeurIPS]
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Presented live at NeurIPS, this technical retrospective highlights the defining computer vision breakthroughs of 2024, analyzing generative video architectures, the ascendance of real-time Detection Transformers over YOLO, methods for resolving vision-language model perceptual blindness, and compact edge models like Moondream.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Peter takes a strong contrarian stance against general VLM hype by proving top commercial models fail on simple visual clock tasks.
Hardest push from the hosts ▶ 39:10 Audience pushes on generalist model failureAn audience participant directly questions why major frontier lab models remain inferior to real-time DETRs at object detection.
Biggest teaching moment ▶ 40:06 Isaac unpacks detection architecture limitsIsaac educates the room on the domain-specific nuances of object detection architectures and historical pre-training deficits in convolutional models.
The host holds their own ▶ 54:44 Host highlights vision as first among equalsHost Alessio synthesizes the presentations by emphasizing vision's preeminent importance across multimodal AI paradigms.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Overview of Major Computer Vision Trends in 2024 | 0 | 0 | 0 | 0 | Isaac presents a solo technical lecture on video generation models and Sora replications without host dialogue. Host-side engagement and pushback are absent during this monologue. | |
| SAM 2 and Temporal Memory for Promptable Video Segmentation | 0 | 0 | 0 | 0 | Isaac continues his presentation detailing SAM 2 architecture, hierarchical encoders, and temporal memory banks. There is no host involvement or disagreement. | |
| The Ascendance of DETRs Over YOLO in Real-Time Object Detection | 0 | 0 | 0 | 0 | Isaac presents technical benchmarks comparing real-time DETRs to YOLO architectures. The segment is entirely a presentation monologue without host participation. | |
| Investigating VLM Blindness and Fine-Grained Perception with MMVP | 0 | 0 | 1 | 0 | Peter begins his talk with the provocative thesis that VLMs cannot see fine details, explaining the MMVP benchmark. It is an instructional presentation delivered to an audience without conversational pushback. | |
| Architectures for Spatial and Semantic Granularity: Florence-2, PaliGemma 2, and AIMv2 | 0 | 0 | 0 | 0 | Peter walks through Florence-2, PaliGemma 2, and AIMv2 benchmark results in an educational format. Host metrics remain zero due to the lecture setting. | |
| Audience Q&A on Foundation Models Versus Specialist Object Detectors | 1 | 2 | 0 | 0 | The host opens the floor to audience questions, prompting Isaac and Peter to explain why generalist VLMs lag behind specialist object detectors. The guests collaboratively break down domain-specific architectures and pre-training dynamics. | |
| Moondream: Compact Vision-Language Models and Sub-Billion Edge Pruning | 0 | 0 | 0 | 0 | Vik presents Moondream's architecture and sub-billion edge pruning methodology. The segment proceeds as a solo technical demo and presentation. | |
| Grounded Chain-of-Thought for Analog Gauge Reading and Visual Reasoning | 1 | 0 | 0 | 0 | Vik explains grounded chain-of-thought for gauge reading before the host offers a brief concluding thought on vision modalities. The dynamic is entirely collaborative and educational. |