The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Devi Parikh no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q Well, what are some of the major forward-looking aspects of this sort of project and research?

A I think there's a ton to do in the context of video generation. Like if you look at make a video, it was very exciting. It was sort of first of its kind, um, capabilities at the time, but it's still, it's a four second video. It's essentially an animated image, right? It's kind of the same scene, the same set of objects that are moving around in reasonable ways, but you're not seeing objects appear, objects disappear. You're not seeing objects reappear. You're not seeing scene transitions. Um, None of this is, is, is in there. And so if you look at, if you think about the complexity of videos that you just regularly come across on various surfaces, this is far from that. And so there is a ton to be done in terms of making these videos longer, more complex, having memory so that if an object reappears, it's actually consistent. It doesn't now look entirely different. Um, things of that sort, sort of being able to tell more complex stories through videos, all of this is, uh, entirely open.

AI assessment note: “there is a ton to be done in terms of making these videos longer, more complex”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q you have these fundamental breakthroughs and sometimes it's like an, it's an architecture like transformer based models versus traditional NLP. And sometimes it's, um, you know, iterating on a lot of other things that already exist in the preexisting approaches and just sort of solving specific engineering or technical challenges. If you were to sort of list out the bottlenecks to this, what do you think they're likely to be?

A Yeah, I think there's a few different things. One is videos are just sort of from an infrastructure perspective, harder to work with, right? They're just sort of larger, more storage and sort of more expensive to process and more expensive to generate and all of that. So there's just that iteration cycle that is much slower with video than it would be with other modalities. So that is, is one. Um, the second is I don't, I don't think we've still figured out the right representations for video. Right. There is a lot of redundancy in video one frame to the next frame. There's not a whole lot that changes. Um, we still kind of approach them fairly independently as sort of independent images, even if you're generating it as kind of one after the other, or even if you're generating in parallel and then making it finer grain. Um, so I think maybe that could be something that helps with a breakthrough that if we really figured out how to represent videos efficiently. Um, and the third is this, Hierarchical architecture that if you want longer videos, they're just so many pixels that you're trying to generate, right? It's a very, very high dimensional signal, um, compared to anything else, uh, that we're doing. And so just thinking through how do we even approach that, what sort of hierarchical representation, um, makes sense, especially if you want these scene transitions, if you want…

AI assessment note: “I think there's a few different things. One is videos are just sort of from”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q recognition, and you've worked across a bunch of different, uh, modalities. How does, how has that changed your, your research path? Because like things like diffusion models and GANs and large transformers, none of that existed when you were first starting. And I, I think you have managed to sort of translate or transition your interest in a way that keeps you on the cutting edge. How has that happened?

A Yes, I think, I mean, you can always kind of look back and try and find patterns. Like when you're actually doing it, you don't necessarily have a grand strategy of anything in mind. But when I look back, I think one common theme that led to me transitioning across topics a little bit, um, was that I was interested in seeing how we can get humans to interact with machines in more meaningful ways. And so kind of Even my transition from kind of non-visual to visual modalities in hindsight, I feel like was essentially that. I felt like you can't interact with these systems too much if it's sort of these abstract modalities that you're looking at. And then when I was working in computer vision, I wanted to find ways for humans to be able to interact with these systems more. So I started looking at kind of these attributes and adjectives of like, oh, something is funny or something is shiny. And using that as a mode of communication between humans and machines, both for Humans to teach machines new concepts and for machines to be more interpretable than explaining why they're making the decisions that they're making, and that slowly led to the sort of more into natural language processing, where instead of these kind of just adjectives and attributes, looking at more natural language as a way of interacting. So a lot of my work in visual question answering, where you're answering qu…

AI assessment note: “one common theme that led to me transitioning across topics a little bit”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And you transitioned from fundamental AI research to a new sort of generative AI group. Can you talk about why that's interesting to Meta or sort of what, what kind of things you're working on now?

A Yeah. Yeah. I mean, yeah, that's a very exciting space right now. There's a lot happening both, uh, within Meta and outside as I'm sure many of the people listening to this are, are aware of. Um, but yeah, so that the new organization was created a few months ago. Um, so not, not a long, not a long time ago. Um, and it's, uh, looking at things like large language models, image generation, video generation, generating three D content, audio, music, um, Yeah, all sort of, all sorts of modalities that you might think of. Um, and Why is it interesting? I mean, like right now, if you think about all the content, there's so much content that we consume in all modalities and all sorts of surfaces. Um, and it makes a lot of sense to ask that instead of, um, maybe not instead of, but in addition to all of this consumption, can more of us be creating more of this content, right? And so, um, almost everything that you think of images, video, you can ask this question, like for any situation where you're searching for something, trying to find something, It's relevant to ask, well, could I just create what it is that I have in my head? Um, and so when you think of it that way, you can see how it can touch, um, a lot of different things across a variety of products and surfaces.

AI assessment note: “looking at things like large language models, image generation, video generation”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You just got back from CVPR and presented there. Can you mention both what you were talking about and then sort of the project or work that most inspired you there?

A Yeah, so I was at CVPR. Um, I was on a few different panels and I was giving some talks and one of it was on vision, language, and creativity at the, at the main conference. Um, and so, yeah, that was kind of what I was representing there. Um, in terms of something exciting that I saw there, um, Not necessarily a paper, but there was a workshop there called, uh, Scholars and Big Models, where the topic of discussion was as these models are getting larger and larger, making a lot of progress in what way can sort of academics or labs that don't have that sort of as many compute resources, what should their strategy be? How should they be approaching these things? Um, and that I thought was a really nice discussion in general. I tend to enjoy Venues that talk about kind of the meta, like the human aspects of the work that we do. We always, we have a lot of technical conversations, but we don't tend to talk about these other components. And so that workshop is something that I enjoyed quite a bit.

AI assessment note: “one of it was on vision, language, and creativity at the main conference”

Redirected raw tape D 2 · C 4 · P 3 · Cm 3 3.00

Q The other, um, potential, uh, part of output for videos obviously is text to speech or some sort of voice or other ways to sort of accompany the video or animate it. What is your view in terms of the state of the art of text to speech systems and how those are evolving?

A I think I've, um, I have, I haven't tracked the text to speech quite as much. What I have tracked a little bit more closely is, um, things like text to audio, where you might say that the sound of a car driving down the street and what you expect is sort of a sound of a car driving down the street, uh, to be generated. Um, and so there, uh, the state of the art right now is, um, sort of roughly sort of a few seconds to tens of seconds long audio. Um, and I would say that roughly it probably works reasonably well, uh, one in five times or so. Um, it's like, because there aren't concrete metrics, it's kind of hard to articulate where state of the art is, but, um, hopefully this is, this is helpful. And I, I do think that, um, audio added to visual content, um, makes it much more expressive and much more delightful. And I do think that it tends to be, um, under invested. Um, uh, both for audio, similarly for music. I think it just makes the content much more expressive, much more delightful. Um, but I feel like we don't do enough of that.

AI assessment note: “I haven't tracked the text to speech quite as much. What I have tracked”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.