why aren't all 16 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Insight
Doshi: Curated data in fine-tuning matters more than vision algorithmic tricks
“I think there are all these, there are like lots of tricks that sometimes get you like 10, 20, sometimes two X improvements. But I think the number one trick is like really just like that last phase of You know, a supervised fine tune where you're finding like…”
Opinion
Doshi: Most AI benchmarks are flawed and optimized for marketing
“I think that most evals in the industry are relatively flawed. Like a lot of them are doing like benchmarks on things that maybe are valuable from the purposes of marketing, but are not necessarily well correlated with the With what maybe users care about.”
Insight
Doshi: Vision AI can access infinite real-world data unlike internet text
“Whereas like with vision, at least you can like make a robot that just like travels down the street and like just keeps taking pictures of everything. You can get like infinite training data with vision, but it might be trickier to like sort of filter and clea…”
Assertion Not checkable as stated
Doshi: Playground AI is probably number two in text-to-art
“Maybe like probably like number two, I suspect, I guess at like text to art at the moment, just because we're training these models from scratch, we're getting, we're closing the gap really rapidly as rapidly as we can around all the various like kind of use c…”
Opinion
Doshi: 3D software tools tend not to make much money
“One issue with three D is that it tends to be better to work on three D. If you're like making the content, like you're making Pixar movies the tools in three D tend to not like make as much money.”
Opinion
Doshi: Diffusion Transformers Alone Lack Utility Without Integrated Language Models
“Transformers are definitely, I think transformers are definitely like the right direction, but I don't think that we're going to get a lot of you enough utility if we're not like somewhat trying to figure out a way to combine the great, amazing knowledge of li…”
Opinion
Doshi: DiT Is Likely Not the Right Architecture for Generative Vision
“So I think the architecture is like mostly going, is most likely going to change. I don't think that DIT is like the right architecture, but transformers certainly.”
Insight
Doshi: Vocals and lyrics are the scarce resource in music, not beats
“Instrumentals in music are actually very easy to get or to make. You know, there, there's a wide variety of like quality, of course, but generally instrumentals in a song, like if you hear a song from Taylor Swift or whoever, a rap song, Those beats or those i…”
Insight
Doshi: Current text-to-image models are text-to-art with limited utility
“The difference is that these models have they don't, they haven't quite reached the potential of, like, what maybe what its utility could be. Right now, for the most part, we take, we formulate a prompt, which is really just a caption of what the image is, and…”
Insight
Doshi: Building top image models is far harder than scaling compute and data
“It turns, turns out that, you know, I think when, if, you know, a set of strong engineers, like their first thought is that like, you just take a model architecture, you find a lot of data, you get, you fund yourself with enough compute and you just sort of li…”
Insight
Doshi: EDM noise formulation fixes diffusion models' brightness and contrast issues
“And so we employed this thing called like this EDM formulation, which like samples the noise slightly differently. And it's like a really clever kind of math trick. And there's probably, there's a paper that you could probably read on it. But it, it's surprisi…”
Disclosure
Doshi: Playground AI will differentiate by focusing on image editing
“We'll probably diverge from some of the other companies in part because we care, I think we start, we're going to care a lot more about editing.”
Insight
Doshi: Current text-to-image AI feels like a high-effort loot box
“It feels a little like a loot box right now. And I think that because it's so much of a loot box, it feels like it's too much effort, I guess, to get something that you really, really want.”
Insight
Doshi: Pixels carry far higher information density than language
“So we know that like pixels have an enormous Amount of high information density compared to language and language is just really between me and you. It's like a compressed way that you and I can converse with each other at a higher, at somewhat of a higher ban…”
Disclosure
Doshi avoided music startups because the entire industry is only $26B
“Partly I didn't work on music because the music industry is like only the whole industry is twenty six billion dollars. So it was a little hard for me to figure out like how big, like a music thing could be.”
Assertion Not checkable as stated
Doshi: Video AI models are pre-trained on a billion images first
“And then the other thing with video is like videos is just extraordinarily computationally expensive to do inference or even training on. And a lot of the video models first train, like pre-train with like a billion images first anyway, to like have a rich, Se…”