The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Gaurav Misra argument clarity score 4.2/5 from 16 exchanges on raw tape · average scores: directness 4.5 · coherence 4.5 · precision 3.8 · compression 3.7 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
16exchanges match
16on raw tape
2redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q As, as well. Uh, so how do you think about that set of companies?

A Totally. So I think there was actually a lot of where those companies came from, like CapCut specifically, if you look at it, right? I mean, CapCut is, the main goal of CapCut is to grow TikTok and, you know, it's a virtuous cycle there. That's kind of why they have it. But they actually came up with their own set of goals because it is a separate company under ByteDance, right? There's TikTok, which is a different company than CapCut. As different companies, they have different goals, different CEOs, et cetera. CapCut's goal actually came from, became inspired for Canva, actually. And They're aspiring to be, like, a Canva, maybe a Canva for video. I think this idea has been floated around a lot because people looked at the last decade and they saw, like, basically Figma and Canva as two sort of major design companies that took out, took off. Figma went after the professional, and Canva went after the non-professional, basically, right? And it's actually a very similar story to what I'm telling you, in a way, because Canva is explicitly used by people who do not design, right? In fact, designers may hate Canva. They might think it's, you know, not It's not good. I could easily make something better, right? I could easily see a lot of designers saying that, right? But for the person who doesn't know how to design, it's amazing, right? And the main reason is it wipes away the bla…

AI assessment note: “CapCut's goal actually came from, became inspired for Canva, actually.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. And, um, what, uh, what is working very well and not working just perfectly just yet?

A So what's working well is what you would expect, which is that you can put more and more data into these models, increase the number of parameters. And they will get better, basically, right? Of course, save any bugs or anything else that might be happening. Now, what's difficult, as always, is to solve, because we are in a unique space from that sense, or we're doing this A-roll video, and so there's unique challenges in that, like audio conditioning is just, it's not a thing that really has been solved at scale, and nobody else has done it, right? And so that means that we have to be the first ones to run into these problems, Try a lot of different things until we figure out what is the best way to solve them. Right. And then scale that up. Um, right. Which is also all very expensive by the way.

AI assessment note: “So what's working well is what you would expect”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Um, how much of this is, uh, whisper, uh, and how much of this is stuff that you had to To innovate on, like, especially, um, I would assume if you have a bunch of creators recording, there might be like a noisy environment or that kind of stuff. Like, how do you solve that?

A Yeah. I mean, to be honest, like, Whisper is already trained on that type of data, right? So it does work really well for those types of things. But what Whisper isn't good at is like, it has a core set of languages that it works well on, and there's other things it doesn't work well on. And I think figuring out the limits of Whisper in different places. So to be clear, like, we don't train transcription models. Uh, but we do, you know, work on that product area very carefully. Right. And a lot of it is like figuring out what works well and for which use case and which condition. Right. And then essentially using a series of different models, correcting each other's mistakes, you know, uh, using it like, oh, what do you use for? Like, I mean, for example, whisper doesn't work in Arabic that well. Right. So we had to figure out how to make something work in Arabic, but it can't be bad because Almost all transcription is bad in Arabic, right? There's just not enough training that's been done, um, on that language. So things like that, uh, is what we figure out. There's a lot of more like product and engineering work rather than model training.

AI assessment note: “There's a lot of more like product and engineering work rather than model training.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah. And a lot of the app, um, going into some of the engineering challenges that you were describing, a lot of the app is, um, real time and very responsive, including on mobile. So how do you think about local AI versus sort of cloud AI?

A I think the more Obviously we prefer moving as much as possible to local, just because that means it gets removed from cogs and sort of, from the financial perspective, it's, it's awesome. Right. Uh, I think not all models are there yet to be able to do that. I think a lot of the easier models, I would say, like transcription and things like that, you can run locally very easily and that cost can be eliminated. But to be honest, even running that on the server is not very expensive. So it doesn't really matter that much, to be honest, except Like maybe offline access for users is more like a user benefit, but for the video generation models, we're not there yet exactly for, especially for the type of users that we're going after. Like they might have older iPhones and stuff, right? So like they may not have the latest and greatest. They might have two generations before, which may not be able to run like a large model essentially on device, right? Even if it's distilled quite a bit. So, and people do care about output. And I think like one of the main criteria for the ability to actually Use something like for it to be like actually practically useful versus entertainment and just like interesting is how photorealistic is it? Right. And it has to be past the level that the average person can tell that this was generated. If that's not the case, then it's not actually usable. Ri…

AI assessment note: “Obviously we prefer moving as much as possible to local, just because that means”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So I want to go back to, um, what you were, um, starting to talk about a few minutes ago. So do you think we passed the uncanny valley for, um, those videos and avatars?

A I think definitely, but it is kind of a moving target. I'll give you an example, right? So, and this is an example not related to us, so it'll be, like, more clear maybe. So I remember, you know, when I first discovered 11 Labs, which was, like, a while back, I heard the voice and I was like, oh my god, this is gonna be transformational. We need to, like, work with these guys. And, like, we reached out to them immediately. They didn't respond immediately, but, uh, and they were a very small shop at the time. It was truly, like, nothing that was there at the time could compare To what they had done. Now, some time has passed, obviously. And even a year later after that, the most common complaint that I got from people was like, oh, it sounds so robotic. It sounds so robotic. And it's like, wait, what? Like, this is the same thing. Nothing has changed. But I think just people just got used to hearing it so much that they just figure out its nuances. Right. And it's the same thing with video. Like a year ago, you know, The video just wasn't good enough to even be presented to users, potentially. As soon as it got good enough, I think it very quickly crossed the boundary of like, wow, like, it's very good, and we're actually past the uncanny valley. It became completely usable, right? I think very quickly people will start realizing now that, oh, there's like certain nuances, it mo…

AI assessment note: “I think definitely, but it is kind of a moving target.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So maybe walk us through that evolution and, um, you know, how do you sell currently?

A Yeah. So it is, um, a consumer application first. So a lot of our distribution happens through consumer channels. The app started off as a completely paid product. So it was a rare sort of premium only app I think it's almost like non-existent at the time, or people would call you crazy for it. It actually helped us develop the best possible product by just collecting the best possible feedback from the users who are willing to pay. And then more recently, we actually switched to the freemium model, which is like more recognized because we realized that we can offer a lot of the classic video editing stuff, right? Like literally the old way, right? What people used to do, which is manually record your video, manually edit your video. That should all be free, right? Because there's no cogs in it anyways, right? It doesn't cost us anything, right? So why should it cost the user? And we think that's the old way anyways, right? We should just give it up. And then our goal is to convince those users to try the new way, right? And hopefully we can convince them that they don't need to spend an hour recording and editing a video. They can just tap a few buttons and generate it.

AI assessment note: “more recently, we actually switched to the freemium model”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Is there a world where you would do fine tuning on a per creator basis? So like for the bigger creators that may have certain ways of doing things, you would, um, You know, fine tune the model to, to just go directly to their style.

A Potentially. Yeah. I mean, I think on video generation, it could be very interesting. Right. But I think at the end of the day, if the, if the conditioning is expressive enough, right. And we're talking like text, image, audio, right. At least these three, you should be able to basically, I mean, that's the point of foundation model, right. Is to be able to mold it into whatever is able to work on unexpected, untrained and things it wasn't trained on. Right. And it can be molded into any particular situation without any fine tuning. So, I do think that's very possible and very, very reasonable to expect from, you know, a foundation model that's large enough to train on enough data.

AI assessment note: “molded into any particular situation without any fine tuning.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So what does that mean for You know, in the future of creation and storytelling. And also, I guess as a related question, what does that mean for the future of, um, professional video editors?

A Yeah, I actually think that this is really big for the future of professional video editors. I mean, so our goal actually isn't to disrupt professional video editing, right? It's actually to enable, and it actually probably creates more professional video editors because a lot more people can get into the profession, right? Because they may not, they may have been too scared of opening Premiere Pro before. And now they can actually make cool stuff, right? With a little bit of dabbling. And of course, the next question they're going to ask is like, oh, how do I move this? And how do I change that? And that gets them into it, right? And they get better and better. And sure, maybe one day they graduate to Premiere Pro, and there's nothing wrong with that, right? That's completely fine with us, right? But I do think that the craft of video editing will change. We're not going to be the ones maybe to change it, or maybe we will. I don't know, right? The future will tell. And this is, by the way, not the first time that craft has evolved, right? Like craft evolves all the time, right? Um, a great example is just like music. Like 50 years ago, you needed to play an instrument, right? To, to play music. If you didn't play guitar, you're not gonna be a musician, right? And then digital music came wrong, right? And suddenly like anybody can like go on a computer and, you know, uh, write …

AI assessment note: “our goal actually isn't to disrupt professional video editing, right? It's actually to enable”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Uh, so, uh, jokes aside, uh, what, what, what has your experience been? Uh, why did you decide to build a company here? Why did you decide to have everyone here versus distributed?

A Yeah, I mean, to be honest, like from the very early days, like we started during peak COVID, right? Like, twenty-twenty-one, right? And, um, it was a time where you couldn't even, like, San Francisco had a hotel shut down, right? Or you couldn't even go there and get a hotel room, right? Um, and so we were a fully remote company. But as soon as COVID started exiting, we realized that to start a company, there's just nothing like being in the same room. And so we started doing the in-person thing. First start with one day, then two days, then three days. And it became five days very quickly. I think New York is uniquely positioned because a couple of reasons, right? One is that in-person companies do have an advantage over not in-person companies. No one is willing to say it. I mean, now many more people are, but like, it is very true, uh, that there is an advantage there. And the advantage is not just, you know, in terms of like, oh, the meetings are much more efficient or something like that, right? It's actually even on a personal and relationship level, people actually get to know each other. Like think about me and my co-founder Dwight. Only overlap for three months, and we kept in touch for 10 years. That would never have happened in a Zoom meeting, right? It just wouldn't. And those are the types of, like, you need trust, you need, right, like, people to be compatible wi…

AI assessment note: “New York is actually well positioned to be the best place to be in person”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q You alluded to some of the models, so let, let, let's, let's get into that. What do you run on in terms of like core models? Uh, what is, um, third party? What is maybe open source? What is proprietary?

A Yeah, so Essentially, internally, what we work on is just video generation, right? That's kind of where our core area is, like, what we can really excel at, because we have the data for it, we have the talent for it, and we have the unique ability to build that specific type of model, right? Now, for everything else, we use some sort of provider, and we've tried basically all of them at this point, so we've been able to figure out which ones work best, but It's an evolving space, and so every two or three months we have to retest everything to just understand where everybody is. There's often, you know, winners that come up over and over again, that we see over and over again, like 11 Labs, for example, comes up over and over again in the audio space.

AI assessment note: “for everything else, we use some sort of provider”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q It's a pretty, uh, magical experience, right? So you, you, you have your video, you import it, you pick a style, right? And then you basically have, like, a magic button, and on the other side of it comes out a video which is edited, where you, you can automatically cut the hesitation and sound and pauses.

A Exactly. Yep. So not just the recording part, you can say pass, so we can generate a video of you, yourself, or anybody, you have the, The license for, so if you have a license for somebody else's likeness, or, you know, we also have actors sort of available in the app that you can pick from, we can generate a video of that person saying, you know, whatever your pitch is, maybe you're pitching a product, maybe it's like an ad, maybe a social media post, right? We can generate that and you can get it perfect. You know, you can say exactly what you want to say, test different messages, whatever it might be. And then we can edit it also, right? So that it, you know, because just footage is footage, Footage isn't exactly usable most of the time.

AI assessment note: “Exactly. Yep. So not just the recording part, you can say pass”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q So the, the people that you empower with, uh, those, those tools, um, so I'm not designers and not video editors. So that's the, your TikTok creator crowd reels. Is that, is that?

A Is that a lot of creators, small businesses. So those are, I would say the two largest sections. I would say a small business is probably the biggest subsection, um, where, you know, I think what, what you end up realizing is like social media and marketing are just hand in hand now and including like ads, right? Like when you run ads, it's just social media, but you're paying for the reach. Right. And that's a massive industry with a lot of footprint, especially among small businesses. So e-commerce think like, um, Like individual businesses, like, um, I would say personal trainers, for example, real estate agents, you know, things like that, where you need to make video as part of your job for some reason or the other, and, um, you don't want to learn all of the skill set, right?

AI assessment note: “Is that a lot of creators, small businesses. So those are, I would say the two largest sections.”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q How far are we from that? What does that depend on? Does that depend on, uh, Runway or Sora doing the B-roll while you do the A-roll?

A Yeah, I mean, it definitely could be. Um, I think, obviously, for a movie, all the cinematic shots and stuff, All the B roll is going to become really important. All the A roll. And then besides that, the editing, right? So as soon as we can put all that together, there's a lot of steps along the way, like everything from character consistency, which we, you know, everybody talks about quite a bit, but really what it is today is just face consistency, right? But face is not enough, right? Everybody has one body. It doesn't change shot to shot, right? It has to be the same exact body, the same like hand size, height, or whatever. Like everything has to be exactly the same. But even location consistency, right? Like we're sitting here, the next shot better be looking exactly the same, right? Not slightly different. So these are the types of things that will have to be developed over time, um, in order for all this to be enabled. But, um, I think it's all very solvable because we've already solved it.

AI assessment note: “these are the types of things that will have to be developed over time”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q Uh, which was, uh, you know, aptly enough the, the name of the company. So, uh, why that at the beginning?

A Initially, we knew that we wanted to do Video creation with AI. Now, at that time, this is pre-GPT, pre-transform, I mean, transformers were there, but no one cared about them at this point. Essentially, we looked at all the different ways that we could apply AI to video, and we wanted something robust enough that it actually works reliably and works really well. And one of the things that was at that point was transcription, right? Speech-to-text. Something that, but, you know, interestingly, what we found at that point is like, even though it was actually like a part of daily life in many ways, like, you know, Siri and Alexa existed at that time. And you know, these were consumer products. They were out there and people understood that speech can be understood by machines. But I think people were shocked surprisingly, like I think outside of tech circles, like the everyday person was kind of surprised to see how good it was. Um, in reality, they could understand all this obscure terminology and really perfectly word for word, get something, you know, as it was said.

AI assessment note: “we wanted something robust enough that it actually works reliably and works really well”

Redirected raw tape D 3 · C 4 · P 3 · Cm 2 3.15

Q Uh, and, uh, yeah, what, what did you, how did you get started, uh, including on the tech and product front?

A We actually started with the idea of revolutionizing video. I think the long-term plan was AI plus video, but we knew that the AI part would take like 10 years, but at least that's what we thought at that time. Right. But the main driver and, you know, our framework for evaluating the initial ideas, especially, and I think, you know, there's a unique framework for figuring out truly like groundbreaking and generational companies. Like how do you build those companies? Like to start a company is much easier, I think, than to start like A truly groundbreaking company, right? And a lot of luck is involved, a lot of hard work is involved, of course, all the classic obvious ingredients, but our framework was actually to look at society and to think about what are the inevitable changes that are happening in society, right? And these could be things like, you know, five G networks are available, right? That changes how people can now upload and download video on their phones, right? Or maybe it's something like Something behavioral in society that's changing, right? Like, and there's a lot of behavioral changes actually happening in society right now, right? Which are inevitable, unstoppable, right? They will complete over a 10, 15 year period. A lot of times it's driven by technology, a lot of times it's driven by just social behavior of the masses, right? And if you identify one of…

AI assessment note: “our framework for evaluating the initial ideas”

Redirected raw tape D 2 · C 2 · P 2 · Cm 2 2.00

Q Yeah, exactly. And, uh, how did you make the decision? Why did you switch? What are the pros and cons?

A Yeah. So we actually run an evaluation every few weeks to a month, basically across all the different models. And we actually have a very, uh, unique way of doing this and probably, you know, an evaluation that Almost no other company can do not just on, you know, the LLM side, but also on the audio gen side and like, you know, other things like music and sound effects. And I mean, these are all things come together to make a video, right? At the end of the day. And I, this is like an important thing to think about is like, if you look at all the media generation foundation models, right? Like there's like music generation and image generation and even video stock video generation. And, you know, there's, uh, you know, voiceovers and Uh, you name it, sound effects, right? Like, where do you think this is all going? Like, what do you do with the sound effect, right? Like, you generated a sound effect. Now what? Like, you don't share a sound effect with a friend. Uh, you don't post a sound effect on social media.

AI assessment note: “where do you think this is all going? Like, what do you do with”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.