Jul 20, 2023 · 39m · no-priors
No Priors Ep. 24 | With Devi Parikh from Meta
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of No Priors, hosts Sarah Guo and Elad Gil interview AI researcher and artist Devi Parikh to discuss the architecture, challenges, and future of generative video and multimodal AI. Parikh details Meta's Make-A-Video system, examines technical bottlenecks in video and audio synthesis, and shares insights on how generative tools are democratizing creative expression and human-machine collaboration.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.2% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Devi rejects the conventional assumption that generative video will mirror the exponential six-month step-change curves seen in LLMs and image diffusion models, noting fundamental architectural and representation hurdles.
Hardest push from the hosts ▶ 23:55 Reframing capability roadmapsSarah intervenes to distill Devi's detailed explanation of generation versus control into the concise operational rule that generative models must produce 'random good first' before steerability becomes viable.
Biggest teaching moment ▶ 8:45 Make-A-Video decoupling architectureDevi breaks down the core breakthrough of Make-A-Video, explaining how text-image pairs provide visual semantics and diversity while motion is learned separately from unlabeled video without paired temporal captions.
The host holds their own ▶ 18:17 Historical infrastructure comparison to YouTubeElad demonstrates host domain expertise by drawing a concrete historical analogy between generative video compute bottlenecks and the streaming infrastructure constraints that drove YouTube's early sale to Google.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Academic Journey and Transition into Computer Vision | 2 | 1 | 0 | 0 | Sarah introduces Devi and prompts her on her academic background. Devi shares her transition from Rowan University pattern recognition to Carnegie Mellon for computer vision in a friendly, biographical opening. | |
| Evolution of Research: From Attributes to Generative Models | 3 | 2 | 0 | 0 | Sarah frames the shift across ML paradigms from early pattern recognition to GANs and diffusion. Devi explains the continuous thread of human-machine interaction guiding her research from attributes to VQA to generative modeling. | |
| Bridging Academia and Industry at Meta Generative AI | 2 | 2 | 0 | 0 | Sarah asks about splitting time between academia and FAIR/Meta. Devi explains the motivation behind Meta's Generative AI group and the shift from content consumption to broad-based creation. | |
| Architecture and Mechanics of Meta's Make-A-Video | 4 | 5 | 0 | 0 | Elad and Sarah drill into the mechanics of Make-A-Video. Devi delivers a detailed technical breakdown of how appearance and semantic correspondence are decoupled from motion learning using unlabeled video data. | |
| Key Bottlenecks and Challenges in Video Generation | 6 | 4 | 1 | 0 | Devi tempers timeline optimism, pointing out that video progress is fundamentally harder due to high dimensionality, data recipes, and representations. Elad adds strong domain perspective by comparing current video compute hurdles to the infrastructure issues that forced YouTube to sell to Google. | |
| Video Understanding in Robotics and Embodied AI | 5 | 4 | 0 | 0 | Sarah inquires about embodied AI and video controllability. Devi explains the active action-perception feedback loop in robotics, and Sarah cleanly synthesizes the development stages of generative control as needing 'random good first.' | |
| State of Text-to-Audio and Multimodal Synthesis | 5 | 3 | 1 | 0 | When asked about text-to-speech, Devi clarifies that she focuses on text-to-audio and sound effect synthesis, highlighting the challenge of acoustic superposition. Elad demonstrates industry knowledge regarding well-labeled sound effect libraries and legacy IP licensing businesses. | |
| Emergent Applications and Product Paradigms for Generative Media | 5 | 2 | 0 | 0 | Elad explores initial product workflows and emergent use cases, drawing an analogy to how Uber unexpectedly unlocked mobile adoption. Devi agrees on the potential for autonomous AI agents to alter media creation. | |
| Democratizing Creativity and AI as a Collaborator | 5 | 2 | 0 | 0 | Sarah pushes back against the cynical view that people do not want to create, citing Instagram, TikTok, and Midjourney. Devi explores how AI acts as both a precision tool and an unpredictable creative collaborator for artists and non-artists alike. |