why aren't all 11 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
Parikh: Generative video progress will lag behind LLM and image advancements
“I think that might be harder in video and I wonder if there is something that we are kind of fundamentally missing in terms of how we approach video generation. So it's not quite answering what you asked me, but I do think that it might be a little bit slower …”
Insight
Parikh: Separating motion from appearance enables video models to use uncaptioned data
“There are sort of multiple advantages of thinking of it that way. One is there's less for the model to learn because you're directly bringing in everything that you already know about images to start with. The second is all of the diversity that we have in our…”
Insight
Parikh: Generative AI will replace search with direct content synthesis
“Almost everything that you think of images, video, you can ask this question, like for any situation where you're searching for something, trying to find something, It's relevant to ask, well, could I just create what it is that I have in my head? And so when …”
Opinion
Parikh: Video understanding is more relevant to robotics than video generation
“So I think there the video understanding piece is probably more relevant than the video generation piece.”
Insight
Parikh: Generative AI control mechanisms consistently lag core model capabilities
“I think control sort of tends to lag behind the core capability. Like even with images, I feel like we first had to get to a point where these models can actually generate nice looking images before we start worrying about, well, is it really doing what I want…”
Prediction Not checkable as stated
Parikh: AI video editing will see faster adoption than from-scratch generation
“We'll probably see much more of, ah, we're already seeing that, and I think we'll see more of where you already have a video that you're starting with, and then you're trying to edit it which has similarities too, but is a little bit different in my mind compa…”
Assertion Not checkable as stated
Parikh: SOTA text-to-audio AI currently works reasonably well one in five times
“The state of the art right now is sort of roughly sort of a few seconds to tens of seconds long audio. And I would say that roughly it probably works reasonably well one in five times or so.”
Opinion
Parikh: Audio and music generation remain under-invested in AI
“And I do think that audio added to visual content makes it much more expressive and much more delightful. And I do think that it tends to be under invested. Both for audio, similarly for music. I think it just makes the content much more expressive, much more …”
Insight
Parikh: Artists view generative AI on a spectrum from tool to collaborator
“So some view them, view these models very much as tools and then others tend to view them as more of a collaborator in this process of creating, and it's always interesting to see what end of the spectrum different people lie on.”
Insight
Parikh: Non-Visual AI Research Lacks Intuitive Feedback Compared to Vision
“I always thought that it was pretty cool that everybody gets to kind of look at the outputs of their algorithms and see what they're doing, whereas if it's kind of non-visual, then yeah, you see these metrics, but you don't really have a sense for what's, what…”
Assertion Supported
Parikh: Make-A-Video initializes from pretrained image model parameters before learning motion
“Concretely the way it works is that when you initialize the model, you're starting off with image generation sort of parameters that have already been learned. So before you do any training for Make a Video, you're, we set it up so that it can generate a few f…”