Apr 18, 2024 · 24m · no-priors

No Priors Ep. 60 | With Playground AI Founder Suhail Doshi

Suhail Doshi · 20m spoken Sarah Guo · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Playground AI founder Suhail Doshi discusses the technical breakthroughs, product philosophies, and architectural innovations driving the evolution of generative vision models from stochastic art generators into comprehensive Large Vision Models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 10.1% of the talking time here. How this is scored →

The hosts as informed peer 4.3 Guest teaching 4.4 Guest disagreement 1.1 The hosts pushing back 1.4
05100:0010:0020:001:34–4:02 · The hosts as informed peer 4/10 Choosing Image Modality and the Competitive Landscape Sarah sets the context by inquiring about Suhail's choice of image modality and the competitive dynamics relative to language. Suhail provides an analytical breakdown of why he avoided crowded language markets and distracted incumbents.4:03–8:22 · The hosts as informed peer 4/10 Beyond Text-to-Art: Expanding Utility and Image Editing Sarah prompts Suhail to distinguish text-to-art from practical utility. Suhail explains the technical nuances of diffusion architectures and how EDM noise sampling solved average brightness issues.8:23–11:53 · The hosts as informed peer 4/10 Aesthetic Optimization and Challenges with Model Evaluation The conversation shifts to model aesthetics and evaluation frameworks. Suhail explains why existing industry benchmarks fail to reflect user aesthetic preferences and human judgment.11:54–14:39 · The hosts as informed peer 5/10 Data Curation and Transitioning from Loot-Box Art to Consistency Sarah probes Playground's user-driven data curation strategy. Suhail outlines how moving away from loot-box generation toward precise editing and character consistency dictates their data approach.14:40–16:40 · The hosts as informed peer 4/10 The Vision for a Large Vision Model Suhail outlines the concept of a Large Vision Model spanning generation, editing, and perception, explaining why starting with images offers superior compute economics over video or 3D.16:41–21:31 · The hosts as informed peer 6/10 The Future of Architectures: Diffusion Transformers and Language Sarah challenges Suhail with the frontier lab hypothesis of a single omni-modal generalist model. Suhail counters by comparing the low information dimensionality of language with the dense physics data contained in raw pixels.21:32–24:07 · The hosts as informed peer 3/10 Exploring Generative Audio and AI in Music Production The hosts ask Suhail about generative audio and music tech. Suhail shares his production workflow using Suno for vocal flow generation while stripping and rebuilding custom instrumentals.1:34–4:02 · Guest teaching 3/10 Choosing Image Modality and the Competitive Landscape Sarah sets the context by inquiring about Suhail's choice of image modality and the competitive dynamics relative to language. Suhail provides an analytical breakdown of why he avoided crowded language markets and distracted incumbents.4:03–8:22 · Guest teaching 5/10 Beyond Text-to-Art: Expanding Utility and Image Editing Sarah prompts Suhail to distinguish text-to-art from practical utility. Suhail explains the technical nuances of diffusion architectures and how EDM noise sampling solved average brightness issues.8:23–11:53 · Guest teaching 5/10 Aesthetic Optimization and Challenges with Model Evaluation The conversation shifts to model aesthetics and evaluation frameworks. Suhail explains why existing industry benchmarks fail to reflect user aesthetic preferences and human judgment.11:54–14:39 · Guest teaching 4/10 Data Curation and Transitioning from Loot-Box Art to Consistency Sarah probes Playground's user-driven data curation strategy. Suhail outlines how moving away from loot-box generation toward precise editing and character consistency dictates their data approach.14:40–16:40 · Guest teaching 5/10 The Vision for a Large Vision Model Suhail outlines the concept of a Large Vision Model spanning generation, editing, and perception, explaining why starting with images offers superior compute economics over video or 3D.16:41–21:31 · Guest teaching 5/10 The Future of Architectures: Diffusion Transformers and Language Sarah challenges Suhail with the frontier lab hypothesis of a single omni-modal generalist model. Suhail counters by comparing the low information dimensionality of language with the dense physics data contained in raw pixels.21:32–24:07 · Guest teaching 4/10 Exploring Generative Audio and AI in Music Production The hosts ask Suhail about generative audio and music tech. Suhail shares his production workflow using Suno for vocal flow generation while stripping and rebuilding custom instrumentals.1:34–4:02 · Guest disagreement 1/10 Choosing Image Modality and the Competitive Landscape Sarah sets the context by inquiring about Suhail's choice of image modality and the competitive dynamics relative to language. Suhail provides an analytical breakdown of why he avoided crowded language markets and distracted incumbents.4:03–8:22 · Guest disagreement 1/10 Beyond Text-to-Art: Expanding Utility and Image Editing Sarah prompts Suhail to distinguish text-to-art from practical utility. Suhail explains the technical nuances of diffusion architectures and how EDM noise sampling solved average brightness issues.8:23–11:53 · Guest disagreement 2/10 Aesthetic Optimization and Challenges with Model Evaluation The conversation shifts to model aesthetics and evaluation frameworks. Suhail explains why existing industry benchmarks fail to reflect user aesthetic preferences and human judgment.11:54–14:39 · Guest disagreement 1/10 Data Curation and Transitioning from Loot-Box Art to Consistency Sarah probes Playground's user-driven data curation strategy. Suhail outlines how moving away from loot-box generation toward precise editing and character consistency dictates their data approach.14:40–16:40 · Guest disagreement 1/10 The Vision for a Large Vision Model Suhail outlines the concept of a Large Vision Model spanning generation, editing, and perception, explaining why starting with images offers superior compute economics over video or 3D.16:41–21:31 · Guest disagreement 2/10 The Future of Architectures: Diffusion Transformers and Language Sarah challenges Suhail with the frontier lab hypothesis of a single omni-modal generalist model. Suhail counters by comparing the low information dimensionality of language with the dense physics data contained in raw pixels.21:32–24:07 · Guest disagreement 0/10 Exploring Generative Audio and AI in Music Production The hosts ask Suhail about generative audio and music tech. Suhail shares his production workflow using Suno for vocal flow generation while stripping and rebuilding custom instrumentals.1:34–4:02 · The hosts pushing back 1/10 Choosing Image Modality and the Competitive Landscape Sarah sets the context by inquiring about Suhail's choice of image modality and the competitive dynamics relative to language. Suhail provides an analytical breakdown of why he avoided crowded language markets and distracted incumbents.4:03–8:22 · The hosts pushing back 2/10 Beyond Text-to-Art: Expanding Utility and Image Editing Sarah prompts Suhail to distinguish text-to-art from practical utility. Suhail explains the technical nuances of diffusion architectures and how EDM noise sampling solved average brightness issues.8:23–11:53 · The hosts pushing back 1/10 Aesthetic Optimization and Challenges with Model Evaluation The conversation shifts to model aesthetics and evaluation frameworks. Suhail explains why existing industry benchmarks fail to reflect user aesthetic preferences and human judgment.11:54–14:39 · The hosts pushing back 2/10 Data Curation and Transitioning from Loot-Box Art to Consistency Sarah probes Playground's user-driven data curation strategy. Suhail outlines how moving away from loot-box generation toward precise editing and character consistency dictates their data approach.14:40–16:40 · The hosts pushing back 1/10 The Vision for a Large Vision Model Suhail outlines the concept of a Large Vision Model spanning generation, editing, and perception, explaining why starting with images offers superior compute economics over video or 3D.16:41–21:31 · The hosts pushing back 3/10 The Future of Architectures: Diffusion Transformers and Language Sarah challenges Suhail with the frontier lab hypothesis of a single omni-modal generalist model. Suhail counters by comparing the low information dimensionality of language with the dense physics data contained in raw pixels.21:32–24:07 · The hosts pushing back 0/10 Exploring Generative Audio and AI in Music Production The hosts ask Suhail about generative audio and music tech. Suhail shares his production workflow using Suno for vocal flow generation while stripping and rebuilding custom instrumentals.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 30.7% · guest 69.3%0:00 · the hosts 30.7% · guest 69.3%3:00 · the hosts 0.5% · guest 99.5%3:00 · the hosts 0.5% · guest 99.5%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 3.3% · guest 96.7%9:00 · the hosts 3.3% · guest 96.7%12:00 · the hosts 17.8% · guest 82.2%12:00 · the hosts 17.8% · guest 82.2%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 19.9% · guest 80.1%18:00 · the hosts 19.9% · guest 80.1%21:00 · the hosts 2.7% · guest 97.3%21:00 · the hosts 2.7% · guest 97.3%24:00 · the hosts 58.6% · guest 41.4%24:00 · the hosts 58.6% · guest 41.4%
Sharpest disagreement ▶ 17:15 Controversial architectural take on Diffusion Transformers

Suhail presents a contrarian viewpoint arguing that pure diffusion transformers like DiT lack interpretability and knowledge without architectural integration with language models.

Hardest push from the hosts ▶ 18:26 Pushing the frontier unified multimodal hypothesis

Sarah directly presents the competing paradigm championed by major language labs: that a single omni-modal reasoning model will subsume specialized domain models.

Biggest teaching moment ▶ 4:23 Text-to-art versus practical image editing utility

Suhail re-educates the conversation on the limitations of current generative models, explaining why prompt-based generation is narrow compared to compositional editing and lighting blending.

The host holds their own ▶ 18:26 Sarah articulates the end-state omni-model thesis

Sarah demonstrates deep domain familiarity by synthesizing the prevailing long-context, multimodal architectural roadmap pursued by leading frontier AI research labs.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Choosing Image Modality and the Competitive Landscape 4311 Sarah sets the context by inquiring about Suhail's choice of image modality and the competitive dynamics relative to language. Suhail provides an analytical breakdown of why he avoided crowded language markets and distracted incumbents.
Beyond Text-to-Art: Expanding Utility and Image Editing 4512 Sarah prompts Suhail to distinguish text-to-art from practical utility. Suhail explains the technical nuances of diffusion architectures and how EDM noise sampling solved average brightness issues.
Aesthetic Optimization and Challenges with Model Evaluation 4521 The conversation shifts to model aesthetics and evaluation frameworks. Suhail explains why existing industry benchmarks fail to reflect user aesthetic preferences and human judgment.
Data Curation and Transitioning from Loot-Box Art to Consistency 5412 Sarah probes Playground's user-driven data curation strategy. Suhail outlines how moving away from loot-box generation toward precise editing and character consistency dictates their data approach.
The Vision for a Large Vision Model 4511 Suhail outlines the concept of a Large Vision Model spanning generation, editing, and perception, explaining why starting with images offers superior compute economics over video or 3D.
The Future of Architectures: Diffusion Transformers and Language 6523 Sarah challenges Suhail with the frontier lab hypothesis of a single omni-modal generalist model. Suhail counters by comparing the low information dimensionality of language with the dense physics data contained in raw pixels.
Exploring Generative Audio and AI in Music Production 3400 The hosts ask Suhail about generative audio and music tech. Suhail shares his production workflow using Suno for vocal flow generation while stripping and rebuilding custom instrumentals.

Statements from this episode (16)

Insight
Doshi: Current text-to-image models are text-to-art with limited utility
“The difference is that these models have they don't, they haven't quite reached the potential of, like, what maybe what its utility could be. Right now, for the most part, we take, we formulate a prompt, which is really just a caption of what the image is, and…”
Suhail Doshi Apr 18, 2024 ▶ 4:24
Insight
Doshi: Building top image models is far harder than scaling compute and data
“It turns, turns out that, you know, I think when, if, you know, a set of strong engineers, like their first thought is that like, you just take a model architecture, you find a lot of data, you get, you fund yourself with enough compute and you just sort of li…”
Suhail Doshi Apr 18, 2024 ▶ 5:44
Insight
Doshi: EDM noise formulation fixes diffusion models' brightness and contrast issues
“And so we employed this thing called like this EDM formulation, which like samples the noise slightly differently. And it's like a really clever kind of math trick. And there's probably, there's a paper that you could probably read on it. But it, it's surprisi…”
Suhail Doshi Apr 18, 2024 ▶ 7:36
Insight
Doshi: Curated data in fine-tuning matters more than vision algorithmic tricks
“I think there are all these, there are like lots of tricks that sometimes get you like 10, 20, sometimes two X improvements. But I think the number one trick is like really just like that last phase of You know, a supervised fine tune where you're finding like…”
Suhail Doshi Apr 18, 2024 ▶ 9:13
Opinion
Doshi: Most AI benchmarks are flawed and optimized for marketing
“I think that most evals in the industry are relatively flawed. Like a lot of them are doing like benchmarks on things that maybe are valuable from the purposes of marketing, but are not necessarily well correlated with the With what maybe users care about.”
Suhail Doshi Apr 18, 2024 ▶ 10:30
Assertion Not checkable as stated
Doshi: Playground AI is probably number two in text-to-art
“Maybe like probably like number two, I suspect, I guess at like text to art at the moment, just because we're training these models from scratch, we're getting, we're closing the gap really rapidly as rapidly as we can around all the various like kind of use c…”
Suhail Doshi Apr 18, 2024 ▶ 13:15
Disclosure
Doshi: Playground AI will differentiate by focusing on image editing
“We'll probably diverge from some of the other companies in part because we care, I think we start, we're going to care a lot more about editing.”
Suhail Doshi Apr 18, 2024 ▶ 13:29
Insight
Doshi: Current text-to-image AI feels like a high-effort loot box
“It feels a little like a loot box right now. And I think that because it's so much of a loot box, it feels like it's too much effort, I guess, to get something that you really, really want.”
Suhail Doshi Apr 18, 2024 ▶ 13:59
Opinion
Doshi: 3D software tools tend not to make much money
“One issue with three D is that it tends to be better to work on three D. If you're like making the content, like you're making Pixar movies the tools in three D tend to not like make as much money.”
Suhail Doshi Apr 18, 2024 ▶ 14:56
Assertion Not checkable as stated
Doshi: Video AI models are pre-trained on a billion images first
“And then the other thing with video is like videos is just extraordinarily computationally expensive to do inference or even training on. And a lot of the video models first train, like pre-train with like a billion images first anyway, to like have a rich, Se…”
Suhail Doshi Apr 18, 2024 ▶ 15:11
Opinion
Doshi: Diffusion Transformers Alone Lack Utility Without Integrated Language Models
“Transformers are definitely, I think transformers are definitely like the right direction, but I don't think that we're going to get a lot of you enough utility if we're not like somewhat trying to figure out a way to combine the great, amazing knowledge of li…”
Suhail Doshi Apr 18, 2024 ▶ 17:25
Opinion
Doshi: DiT Is Likely Not the Right Architecture for Generative Vision
“So I think the architecture is like mostly going, is most likely going to change. I don't think that DIT is like the right architecture, but transformers certainly.”
Suhail Doshi Apr 18, 2024 ▶ 18:12
Insight
Doshi: Pixels carry far higher information density than language
“So we know that like pixels have an enormous Amount of high information density compared to language and language is just really between me and you. It's like a compressed way that you and I can converse with each other at a higher, at somewhat of a higher ban…”
Suhail Doshi Apr 18, 2024 ▶ 19:47
Insight
Doshi: Vision AI can access infinite real-world data unlike internet text
“Whereas like with vision, at least you can like make a robot that just like travels down the street and like just keeps taking pictures of everything. You can get like infinite training data with vision, but it might be trickier to like sort of filter and clea…”
Suhail Doshi Apr 18, 2024 ▶ 21:14
Disclosure
Doshi avoided music startups because the entire industry is only $26B
“Partly I didn't work on music because the music industry is like only the whole industry is twenty six billion dollars. So it was a little hard for me to figure out like how big, like a music thing could be.”
Suhail Doshi Apr 18, 2024 ▶ 21:59
Insight
Doshi: Vocals and lyrics are the scarce resource in music, not beats
“Instrumentals in music are actually very easy to get or to make. You know, there, there's a wide variety of like quality, of course, but generally instrumentals in a song, like if you hear a song from Taylor Swift or whoever, a rap song, Those beats or those i…”
Suhail Doshi Apr 18, 2024 ▶ 22:26
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.