Everything Suhail Doshi said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Doshi: Search engine traffic will drop over 25% before 2026
“I mean, I think there's a very high probability it is greater than that number in a shorter span of time.”
Suhail Doshi states Soham Parikh had a net negative impact at Lindy
“I am literally out of a meeting like an hour ago, where the engineer described his impact as negative.”
Doshi: Curated data in fine-tuning matters more than vision algorithmic tricks
“I think there are all these, there are like lots of tricks that sometimes get you like 10, 20, sometimes two X improvements. But I think the number one trick is like really just like that last phase of You know, a supervised fine tune where you're finding like…”
Doshi: Most AI benchmarks are flawed and optimized for marketing
“I think that most evals in the industry are relatively flawed. Like a lot of them are doing like benchmarks on things that maybe are valuable from the purposes of marketing, but are not necessarily well correlated with the With what maybe users care about.”
Doshi: Vision AI can access infinite real-world data unlike internet text
“Whereas like with vision, at least you can like make a robot that just like travels down the street and like just keeps taking pictures of everything. You can get like infinite training data with vision, but it might be trickier to like sort of filter and clea…”
Doshi: Powerful multimodal models will probably yield Westworld-style humanoid robots
“So I think that if the models get more powerful, We're probably going to see, we're probably going to see some kind of Westworld version of the world. We're going to go way beyond a chatbot.”
Suhail Doshi: Video AI models mimic training data rather than understanding real physics
“What's really doing is it's representing
The physics it understands in the videos it's being trained on, which could be incorrect physics. It's really what it understands what it's being trained on is kind of my main, the main thrust of my point. And that to a…”
Doshi: AI companies would leave NVIDIA if cheaper compute alternatives existed
“It's not CUDA that's keeping, I think, keeping a lot of us. It's actually that there is nothing really dramatically better than NVIDIA's GPUs. And so if there's nothing dramatically better than, I mean, the reality is the cost for training and inference are so…”
Doshi: AWS Quoted 5x Market Price for GPUs in 2023
“I mean, we look, we were looking for compute last year and AWS wanted to charge us five times more for the same GPU than 10 different providers all around them.”
Doshi: Image Generative AI Is Stuck in a 'GPT-2 Moment'
“So I think that we continue to feel like graphics and these foundation models for anything really related to pixels, but also definitely images continues to be very under invested. It feels a little like graphics is in like this GPT two moment, right? Like eve…”
Doshi: Image generation will ultimately converge on unified models
“I think we will make an unified model. I think it will, I think we'll certainly in the end ultimately make a unified model.”
Suhail Doshi: Midjourney is currently the best image generation model
“We, you know, we have, everyone has to acknowledge that Midjourney is very good. You know, they're, they are the best at this thing. We would, I would happily, I'm happy to admit that.”
Doshi: Playground AI is probably number two in text-to-art
“Maybe like probably like number two, I suspect, I guess at like text to art at the moment, just because we're training these models from scratch, we're getting, we're closing the gap really rapidly as rapidly as we can around all the various like kind of use c…”
Doshi: 3D software tools tend not to make much money
“One issue with three D is that it tends to be better to work on three D. If you're like making the content, like you're making Pixar movies the tools in three D tend to not like make as much money.”
Doshi: Diffusion Transformers Alone Lack Utility Without Integrated Language Models
“Transformers are definitely, I think transformers are definitely like the right direction, but I don't think that we're going to get a lot of you enough utility if we're not like somewhat trying to figure out a way to combine the great, amazing knowledge of li…”
Doshi: DiT Is Likely Not the Right Architecture for Generative Vision
“So I think the architecture is like mostly going, is most likely going to change. I don't think that DIT is like the right architecture, but transformers certainly.”
Doshi: Vocals and lyrics are the scarce resource in music, not beats
“Instrumentals in music are actually very easy to get or to make. You know, there, there's a wide variety of like quality, of course, but generally instrumentals in a song, like if you hear a song from Taylor Swift or whoever, a rap song, Those beats or those i…”
Doshi: Image AI is probably a couple years behind a GPT-4 moment
“Images is interesting because it's probably a couple years behind a GPT-IV true moment, right?”
Doshi: Robotics research hit a ceiling around 2020 and hasn't recovered
“Broadly robotics kind of asymptoted and hit a ceiling. About like three or four years ago, and the research still isn't like kind of on a trajectory. That's amazing.”
Doshi: Articulating explanations does not prove AI is genuinely reasoning
“Just because something is able to articulate its reason for doing, for getting to an answer, doesn't mean that it is necessarily reasoning.”
Doshi: The Predominant Use Case for Language Models Continues to Be Homework
“There are a lot of things posted on Twitter about how people are using language models, but the predominant use case continues to be homework.”
Doshi: Graphic Designers Won't Go Extinct From AI but Will Retool
“So did all the people that you know, that were people that drew the two D cartoons, they lose their jobs. And, you know, that was the end of the end of an era. Definitely not people retooled the stories that came from them were still really material to their c…”
Doshi: Computer vision lacks a single unified model equivalent to LLMs
“This is missing in vision, but definitely kind of exists in language and language. We can solve hundreds or thousands of different tasks. But in graphics, but in vision, it's all separated. It's kind of like where language was back three or four years ago wher…”
Doshi: Prompting AI image models by artist name will disappear this year
“I think that this thing is going to go away in the next year, this year.”