Jul 20, 2023 · 39m · no-priors

No Priors Ep. 24 | With Devi Parikh from Meta

Devi Parikh · 28m spoken Sarah Guo · 4m spoken Elad Gil · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, hosts Sarah Guo and Elad Gil interview AI researcher and artist Devi Parikh to discuss the architecture, challenges, and future of generative video and multimodal AI. Parikh details Meta's Make-A-Video system, examines technical bottlenecks in video and audio synthesis, and shares insights on how generative tools are democratizing creative expression and human-machine collaboration.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.2% of the talking time here. How this is scored →

The hosts as informed peer 4.1 Guest teaching 2.8 Guest disagreement 0.2 The hosts pushing back 0.0
05100:0010:0020:0030:000:37–2:37 · The hosts as informed peer 2/10 Academic Journey and Transition into Computer Vision Sarah introduces Devi and prompts her on her academic background. Devi shares her transition from Rowan University pattern recognition to Carnegie Mellon for computer vision in a friendly, biographical opening.2:39–5:11 · The hosts as informed peer 3/10 Evolution of Research: From Attributes to Generative Models Sarah frames the shift across ML paradigms from early pattern recognition to GANs and diffusion. Devi explains the continuous thread of human-machine interaction guiding her research from attributes to VQA to generative modeling.5:12–7:52 · The hosts as informed peer 2/10 Bridging Academia and Industry at Meta Generative AI Sarah asks about splitting time between academia and FAIR/Meta. Devi explains the motivation behind Meta's Generative AI group and the shift from content consumption to broad-based creation.7:53–12:42 · The hosts as informed peer 4/10 Architecture and Mechanics of Meta's Make-A-Video Elad and Sarah drill into the mechanics of Make-A-Video. Devi delivers a detailed technical breakdown of how appearance and semantic correspondence are decoupled from motion learning using unlabeled video data.12:43–19:08 · The hosts as informed peer 6/10 Key Bottlenecks and Challenges in Video Generation Devi tempers timeline optimism, pointing out that video progress is fundamentally harder due to high dimensionality, data recipes, and representations. Elad adds strong domain perspective by comparing current video compute hurdles to the infrastructure issues that forced YouTube to sell to Google.19:09–24:52 · The hosts as informed peer 5/10 Video Understanding in Robotics and Embodied AI Sarah inquires about embodied AI and video controllability. Devi explains the active action-perception feedback loop in robotics, and Sarah cleanly synthesizes the development stages of generative control as needing 'random good first.'24:53–27:19 · The hosts as informed peer 5/10 State of Text-to-Audio and Multimodal Synthesis When asked about text-to-speech, Devi clarifies that she focuses on text-to-audio and sound effect synthesis, highlighting the challenge of acoustic superposition. Elad demonstrates industry knowledge regarding well-labeled sound effect libraries and legacy IP licensing businesses.27:21–29:21 · The hosts as informed peer 5/10 Emergent Applications and Product Paradigms for Generative Media Elad explores initial product workflows and emergent use cases, drawing an analogy to how Uber unexpectedly unlocked mobile adoption. Devi agrees on the potential for autonomous AI agents to alter media creation.29:22–34:33 · The hosts as informed peer 5/10 Democratizing Creativity and AI as a Collaborator Sarah pushes back against the cynical view that people do not want to create, citing Instagram, TikTok, and Midjourney. Devi explores how AI acts as both a precision tool and an unpredictable creative collaborator for artists and non-artists alike.0:37–2:37 · Guest teaching 1/10 Academic Journey and Transition into Computer Vision Sarah introduces Devi and prompts her on her academic background. Devi shares her transition from Rowan University pattern recognition to Carnegie Mellon for computer vision in a friendly, biographical opening.2:39–5:11 · Guest teaching 2/10 Evolution of Research: From Attributes to Generative Models Sarah frames the shift across ML paradigms from early pattern recognition to GANs and diffusion. Devi explains the continuous thread of human-machine interaction guiding her research from attributes to VQA to generative modeling.5:12–7:52 · Guest teaching 2/10 Bridging Academia and Industry at Meta Generative AI Sarah asks about splitting time between academia and FAIR/Meta. Devi explains the motivation behind Meta's Generative AI group and the shift from content consumption to broad-based creation.7:53–12:42 · Guest teaching 5/10 Architecture and Mechanics of Meta's Make-A-Video Elad and Sarah drill into the mechanics of Make-A-Video. Devi delivers a detailed technical breakdown of how appearance and semantic correspondence are decoupled from motion learning using unlabeled video data.12:43–19:08 · Guest teaching 4/10 Key Bottlenecks and Challenges in Video Generation Devi tempers timeline optimism, pointing out that video progress is fundamentally harder due to high dimensionality, data recipes, and representations. Elad adds strong domain perspective by comparing current video compute hurdles to the infrastructure issues that forced YouTube to sell to Google.19:09–24:52 · Guest teaching 4/10 Video Understanding in Robotics and Embodied AI Sarah inquires about embodied AI and video controllability. Devi explains the active action-perception feedback loop in robotics, and Sarah cleanly synthesizes the development stages of generative control as needing 'random good first.'24:53–27:19 · Guest teaching 3/10 State of Text-to-Audio and Multimodal Synthesis When asked about text-to-speech, Devi clarifies that she focuses on text-to-audio and sound effect synthesis, highlighting the challenge of acoustic superposition. Elad demonstrates industry knowledge regarding well-labeled sound effect libraries and legacy IP licensing businesses.27:21–29:21 · Guest teaching 2/10 Emergent Applications and Product Paradigms for Generative Media Elad explores initial product workflows and emergent use cases, drawing an analogy to how Uber unexpectedly unlocked mobile adoption. Devi agrees on the potential for autonomous AI agents to alter media creation.29:22–34:33 · Guest teaching 2/10 Democratizing Creativity and AI as a Collaborator Sarah pushes back against the cynical view that people do not want to create, citing Instagram, TikTok, and Midjourney. Devi explores how AI acts as both a precision tool and an unpredictable creative collaborator for artists and non-artists alike.0:37–2:37 · Guest disagreement 0/10 Academic Journey and Transition into Computer Vision Sarah introduces Devi and prompts her on her academic background. Devi shares her transition from Rowan University pattern recognition to Carnegie Mellon for computer vision in a friendly, biographical opening.2:39–5:11 · Guest disagreement 0/10 Evolution of Research: From Attributes to Generative Models Sarah frames the shift across ML paradigms from early pattern recognition to GANs and diffusion. Devi explains the continuous thread of human-machine interaction guiding her research from attributes to VQA to generative modeling.5:12–7:52 · Guest disagreement 0/10 Bridging Academia and Industry at Meta Generative AI Sarah asks about splitting time between academia and FAIR/Meta. Devi explains the motivation behind Meta's Generative AI group and the shift from content consumption to broad-based creation.7:53–12:42 · Guest disagreement 0/10 Architecture and Mechanics of Meta's Make-A-Video Elad and Sarah drill into the mechanics of Make-A-Video. Devi delivers a detailed technical breakdown of how appearance and semantic correspondence are decoupled from motion learning using unlabeled video data.12:43–19:08 · Guest disagreement 1/10 Key Bottlenecks and Challenges in Video Generation Devi tempers timeline optimism, pointing out that video progress is fundamentally harder due to high dimensionality, data recipes, and representations. Elad adds strong domain perspective by comparing current video compute hurdles to the infrastructure issues that forced YouTube to sell to Google.19:09–24:52 · Guest disagreement 0/10 Video Understanding in Robotics and Embodied AI Sarah inquires about embodied AI and video controllability. Devi explains the active action-perception feedback loop in robotics, and Sarah cleanly synthesizes the development stages of generative control as needing 'random good first.'24:53–27:19 · Guest disagreement 1/10 State of Text-to-Audio and Multimodal Synthesis When asked about text-to-speech, Devi clarifies that she focuses on text-to-audio and sound effect synthesis, highlighting the challenge of acoustic superposition. Elad demonstrates industry knowledge regarding well-labeled sound effect libraries and legacy IP licensing businesses.27:21–29:21 · Guest disagreement 0/10 Emergent Applications and Product Paradigms for Generative Media Elad explores initial product workflows and emergent use cases, drawing an analogy to how Uber unexpectedly unlocked mobile adoption. Devi agrees on the potential for autonomous AI agents to alter media creation.29:22–34:33 · Guest disagreement 0/10 Democratizing Creativity and AI as a Collaborator Sarah pushes back against the cynical view that people do not want to create, citing Instagram, TikTok, and Midjourney. Devi explores how AI acts as both a precision tool and an unpredictable creative collaborator for artists and non-artists alike.0:37–2:37 · The hosts pushing back 0/10 Academic Journey and Transition into Computer Vision Sarah introduces Devi and prompts her on her academic background. Devi shares her transition from Rowan University pattern recognition to Carnegie Mellon for computer vision in a friendly, biographical opening.2:39–5:11 · The hosts pushing back 0/10 Evolution of Research: From Attributes to Generative Models Sarah frames the shift across ML paradigms from early pattern recognition to GANs and diffusion. Devi explains the continuous thread of human-machine interaction guiding her research from attributes to VQA to generative modeling.5:12–7:52 · The hosts pushing back 0/10 Bridging Academia and Industry at Meta Generative AI Sarah asks about splitting time between academia and FAIR/Meta. Devi explains the motivation behind Meta's Generative AI group and the shift from content consumption to broad-based creation.7:53–12:42 · The hosts pushing back 0/10 Architecture and Mechanics of Meta's Make-A-Video Elad and Sarah drill into the mechanics of Make-A-Video. Devi delivers a detailed technical breakdown of how appearance and semantic correspondence are decoupled from motion learning using unlabeled video data.12:43–19:08 · The hosts pushing back 0/10 Key Bottlenecks and Challenges in Video Generation Devi tempers timeline optimism, pointing out that video progress is fundamentally harder due to high dimensionality, data recipes, and representations. Elad adds strong domain perspective by comparing current video compute hurdles to the infrastructure issues that forced YouTube to sell to Google.19:09–24:52 · The hosts pushing back 0/10 Video Understanding in Robotics and Embodied AI Sarah inquires about embodied AI and video controllability. Devi explains the active action-perception feedback loop in robotics, and Sarah cleanly synthesizes the development stages of generative control as needing 'random good first.'24:53–27:19 · The hosts pushing back 0/10 State of Text-to-Audio and Multimodal Synthesis When asked about text-to-speech, Devi clarifies that she focuses on text-to-audio and sound effect synthesis, highlighting the challenge of acoustic superposition. Elad demonstrates industry knowledge regarding well-labeled sound effect libraries and legacy IP licensing businesses.27:21–29:21 · The hosts pushing back 0/10 Emergent Applications and Product Paradigms for Generative Media Elad explores initial product workflows and emergent use cases, drawing an analogy to how Uber unexpectedly unlocked mobile adoption. Devi agrees on the potential for autonomous AI agents to alter media creation.29:22–34:33 · The hosts pushing back 0/10 Democratizing Creativity and AI as a Collaborator Sarah pushes back against the cynical view that people do not want to create, citing Instagram, TikTok, and Midjourney. Devi explores how AI acts as both a precision tool and an unpredictable creative collaborator for artists and non-artists alike.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 35.9% · guest 64.1%0:00 · the hosts 35.9% · guest 64.1%3:00 · the hosts 7.2% · guest 92.8%3:00 · the hosts 7.2% · guest 92.8%6:00 · the hosts 23.5% · guest 76.5%6:00 · the hosts 23.5% · guest 76.5%9:00 · the hosts 10.2% · guest 89.8%9:00 · the hosts 10.2% · guest 89.8%12:00 · the hosts 7.7% · guest 92.3%12:00 · the hosts 7.7% · guest 92.3%15:00 · the hosts 20.5% · guest 79.5%15:00 · the hosts 20.5% · guest 79.5%18:00 · the hosts 33.3% · guest 66.7%18:00 · the hosts 33.3% · guest 66.7%21:00 · the hosts 7.9% · guest 92.1%21:00 · the hosts 7.9% · guest 92.1%24:00 · the hosts 20.2% · guest 79.8%24:00 · the hosts 20.2% · guest 79.8%27:00 · the hosts 48.9% · guest 51.1%27:00 · the hosts 48.9% · guest 51.1%30:00 · the hosts 23.1% · guest 76.9%30:00 · the hosts 23.1% · guest 76.9%33:00 · the hosts 34.4% · guest 65.6%33:00 · the hosts 34.4% · guest 65.6%36:00 · the hosts 32.9% · guest 67.1%36:00 · the hosts 32.9% · guest 67.1%39:00 · the hosts 10% · guest 90%39:00 · the hosts 10% · guest 90%
Sharpest disagreement ▶ 14:15 Tempering video progress expectations

Devi rejects the conventional assumption that generative video will mirror the exponential six-month step-change curves seen in LLMs and image diffusion models, noting fundamental architectural and representation hurdles.

Hardest push from the hosts ▶ 23:55 Reframing capability roadmaps

Sarah intervenes to distill Devi's detailed explanation of generation versus control into the concise operational rule that generative models must produce 'random good first' before steerability becomes viable.

Biggest teaching moment ▶ 8:45 Make-A-Video decoupling architecture

Devi breaks down the core breakthrough of Make-A-Video, explaining how text-image pairs provide visual semantics and diversity while motion is learned separately from unlabeled video without paired temporal captions.

The host holds their own ▶ 18:17 Historical infrastructure comparison to YouTube

Elad demonstrates host domain expertise by drawing a concrete historical analogy between generative video compute bottlenecks and the streaming infrastructure constraints that drove YouTube's early sale to Google.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Academic Journey and Transition into Computer Vision 2100 Sarah introduces Devi and prompts her on her academic background. Devi shares her transition from Rowan University pattern recognition to Carnegie Mellon for computer vision in a friendly, biographical opening.
Evolution of Research: From Attributes to Generative Models 3200 Sarah frames the shift across ML paradigms from early pattern recognition to GANs and diffusion. Devi explains the continuous thread of human-machine interaction guiding her research from attributes to VQA to generative modeling.
Bridging Academia and Industry at Meta Generative AI 2200 Sarah asks about splitting time between academia and FAIR/Meta. Devi explains the motivation behind Meta's Generative AI group and the shift from content consumption to broad-based creation.
Architecture and Mechanics of Meta's Make-A-Video 4500 Elad and Sarah drill into the mechanics of Make-A-Video. Devi delivers a detailed technical breakdown of how appearance and semantic correspondence are decoupled from motion learning using unlabeled video data.
Key Bottlenecks and Challenges in Video Generation 6410 Devi tempers timeline optimism, pointing out that video progress is fundamentally harder due to high dimensionality, data recipes, and representations. Elad adds strong domain perspective by comparing current video compute hurdles to the infrastructure issues that forced YouTube to sell to Google.
Video Understanding in Robotics and Embodied AI 5400 Sarah inquires about embodied AI and video controllability. Devi explains the active action-perception feedback loop in robotics, and Sarah cleanly synthesizes the development stages of generative control as needing 'random good first.'
State of Text-to-Audio and Multimodal Synthesis 5310 When asked about text-to-speech, Devi clarifies that she focuses on text-to-audio and sound effect synthesis, highlighting the challenge of acoustic superposition. Elad demonstrates industry knowledge regarding well-labeled sound effect libraries and legacy IP licensing businesses.
Emergent Applications and Product Paradigms for Generative Media 5200 Elad explores initial product workflows and emergent use cases, drawing an analogy to how Uber unexpectedly unlocked mobile adoption. Devi agrees on the potential for autonomous AI agents to alter media creation.
Democratizing Creativity and AI as a Collaborator 5200 Sarah pushes back against the cynical view that people do not want to create, citing Instagram, TikTok, and Midjourney. Devi explores how AI acts as both a precision tool and an unpredictable creative collaborator for artists and non-artists alike.

Statements from this episode (13)

Insight
Parikh: Non-Visual AI Research Lacks Intuitive Feedback Compared to Vision
“I always thought that it was pretty cool that everybody gets to kind of look at the outputs of their algorithms and see what they're doing, whereas if it's kind of non-visual, then yeah, you see these metrics, but you don't really have a sense for what's, what…”
Devi Parikh Jul 20, 2023 ▶ 2:11
Insight
Parikh: Generative AI will replace search with direct content synthesis
“Almost everything that you think of images, video, you can ask this question, like for any situation where you're searching for something, trying to find something, It's relevant to ask, well, could I just create what it is that I have in my head? And so when …”
Devi Parikh Jul 20, 2023 ▶ 7:32
Insight
Parikh: Separating motion from appearance enables video models to use uncaptioned data
“There are sort of multiple advantages of thinking of it that way. One is there's less for the model to learn because you're directly bringing in everything that you already know about images to start with. The second is all of the diversity that we have in our…”
Devi Parikh Jul 20, 2023 ▶ 9:59
Assertion Supported
Parikh: Make-A-Video initializes from pretrained image model parameters before learning motion
“Concretely the way it works is that when you initialize the model, you're starting off with image generation sort of parameters that have already been learned. So before you do any training for Make a Video, you're, we set it up so that it can generate a few f…”
Devi Parikh Jul 20, 2023 ▶ 10:54
Prediction Not checkable as stated
Parikh: Generative video progress will lag behind LLM and image advancements
“I think that might be harder in video and I wonder if there is something that we are kind of fundamentally missing in terms of how we approach video generation. So it's not quite answering what you asked me, but I do think that it might be a little bit slower …”
Devi Parikh Jul 20, 2023 ▶ 14:29
Assertion Supported
Gil: Infrastructure Costs Drove YouTube's Early Acquisition by Google
“One of the reasons YouTube sold was the infrastructure point you made earlier, where just dealing with that huge amount of streaming and the costs associated with it and everything else, even in the prior generation of just, you know, can we host and stream th…”
Elad Gil Jul 20, 2023 ▶ 18:24
Opinion
Parikh: Video understanding is more relevant to robotics than video generation
“So I think there the video understanding piece is probably more relevant than the video generation piece.”
Devi Parikh Jul 20, 2023 ▶ 19:22
Insight
Parikh: Generative AI control mechanisms consistently lag core model capabilities
“I think control sort of tends to lag behind the core capability. Like even with images, I feel like we first had to get to a point where these models can actually generate nice looking images before we start worrying about, well, is it really doing what I want…”
Devi Parikh Jul 20, 2023 ▶ 23:40
Prediction Not checkable as stated
Parikh: AI video editing will see faster adoption than from-scratch generation
“We'll probably see much more of, ah, we're already seeing that, and I think we'll see more of where you already have a video that you're starting with, and then you're trying to edit it which has similarities too, but is a little bit different in my mind compa…”
Devi Parikh Jul 20, 2023 ▶ 24:36
Assertion Not checkable as stated
Parikh: SOTA text-to-audio AI currently works reasonably well one in five times
“The state of the art right now is sort of roughly sort of a few seconds to tens of seconds long audio. And I would say that roughly it probably works reasonably well one in five times or so.”
Devi Parikh Jul 20, 2023 ▶ 25:25
Opinion
Parikh: Audio and music generation remain under-invested in AI
“And I do think that audio added to visual content makes it much more expressive and much more delightful. And I do think that it tends to be under invested. Both for audio, similarly for music. I think it just makes the content much more expressive, much more …”
Devi Parikh Jul 20, 2023 ▶ 25:52
Prediction Not checkable as stated
Gil: Generative AI will eliminate the sound effect licensing industry
“Which it seems like eventually that industry is likely to go away, so.”
Elad Gil Jul 20, 2023 ▶ 26:33
Insight
Parikh: Artists view generative AI on a spectrum from tool to collaborator
“So some view them, view these models very much as tools and then others tend to view them as more of a collaborator in this process of creating, and it's always interesting to see what end of the spectrum different people lie on.”
Devi Parikh Jul 20, 2023 ▶ 34:20
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.