Apr 18, 2024 · 24m · no-priors
No Priors Ep. 60 | With Playground AI Founder Suhail Doshi
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of No Priors, Playground AI founder Suhail Doshi discusses the technical breakthroughs, product philosophies, and architectural innovations driving the evolution of generative vision models from stochastic art generators into comprehensive Large Vision Models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 10.1% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Suhail presents a contrarian viewpoint arguing that pure diffusion transformers like DiT lack interpretability and knowledge without architectural integration with language models.
Hardest push from the hosts ▶ 18:26 Pushing the frontier unified multimodal hypothesisSarah directly presents the competing paradigm championed by major language labs: that a single omni-modal reasoning model will subsume specialized domain models.
Biggest teaching moment ▶ 4:23 Text-to-art versus practical image editing utilitySuhail re-educates the conversation on the limitations of current generative models, explaining why prompt-based generation is narrow compared to compositional editing and lighting blending.
The host holds their own ▶ 18:26 Sarah articulates the end-state omni-model thesisSarah demonstrates deep domain familiarity by synthesizing the prevailing long-context, multimodal architectural roadmap pursued by leading frontier AI research labs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Choosing Image Modality and the Competitive Landscape | 4 | 3 | 1 | 1 | Sarah sets the context by inquiring about Suhail's choice of image modality and the competitive dynamics relative to language. Suhail provides an analytical breakdown of why he avoided crowded language markets and distracted incumbents. | |
| Beyond Text-to-Art: Expanding Utility and Image Editing | 4 | 5 | 1 | 2 | Sarah prompts Suhail to distinguish text-to-art from practical utility. Suhail explains the technical nuances of diffusion architectures and how EDM noise sampling solved average brightness issues. | |
| Aesthetic Optimization and Challenges with Model Evaluation | 4 | 5 | 2 | 1 | The conversation shifts to model aesthetics and evaluation frameworks. Suhail explains why existing industry benchmarks fail to reflect user aesthetic preferences and human judgment. | |
| Data Curation and Transitioning from Loot-Box Art to Consistency | 5 | 4 | 1 | 2 | Sarah probes Playground's user-driven data curation strategy. Suhail outlines how moving away from loot-box generation toward precise editing and character consistency dictates their data approach. | |
| The Vision for a Large Vision Model | 4 | 5 | 1 | 1 | Suhail outlines the concept of a Large Vision Model spanning generation, editing, and perception, explaining why starting with images offers superior compute economics over video or 3D. | |
| The Future of Architectures: Diffusion Transformers and Language | 6 | 5 | 2 | 3 | Sarah challenges Suhail with the frontier lab hypothesis of a single omni-modal generalist model. Suhail counters by comparing the low information dimensionality of language with the dense physics data contained in raw pixels. | |
| Exploring Generative Audio and AI in Music Production | 3 | 4 | 0 | 0 | The hosts ask Suhail about generative audio and music tech. Suhail shares his production workflow using Suno for vocal flow generation while stripping and rebuilding custom instrumentals. |