May 16, 2024 · 30m · no-priors

No Priors Ep. 64 | With Suno CEO and Co-Founder Mikey Shulman

Mikey Shulman · 18m spoken Sarah Guo · 4m spoken Elad Gil · 3m spoken AI Generated Song · 30s spoken Oliver McCann · 20s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Suno co-founder and CEO Mikey Shulman discusses the technical architecture, product philosophy, and cultural impact of generative AI music, exploring how accessible foundation models are turning passive listeners into active creators.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.7% of the talking time here. How this is scored →

The hosts as informed peer 4.9 Guest teaching 4.5 Guest disagreement 1.9 The hosts pushing back 1.4
05100:0010:0020:0030:001:03–3:47 · The hosts as informed peer 3/10 Mikey Shulman's Journey from Physics PhD to AI Sarah sets a welcoming tone by asking Mikey about his background from childhood piano to a Harvard physics PhD and founding AI startups. Mikey answers self-deprecatingly about his skills in physics and music, keeping the dialogue informal and friendly.3:47–8:50 · The hosts as informed peer 6/10 From Speech Transcription to Musical AI Generation Elad demonstrates technical grasp of AI audio architectures by asking about diffusion versus transformer models and scaling laws. Mikey explains why Suno opted for standard transformers and focuses its technical innovation on audio tokenization heuristics at high sample rates.8:50–11:34 · The hosts as informed peer 5/10 Avoiding Explicit Bias and Targeting Human Emotion Sarah probes the technical frontiers of audio versus video and image models. Mikey educates the hosts on why imposing music theory or fixed instrument counts creates harmful bias, comparing next-token prediction in audio directly to text LLMs.11:34–14:39 · The hosts as informed peer 7/10 Consumer Product Strategy and Generative AI Monetization Mikey challenges conventional SaaS pricing for GenAI, teasingly critiquing SaaS-oriented venture investors and founders for lacking pricing imagination. Elad matches this analysis with historical context from the 1990s web browser era comparing micropayments and ad-based models.14:39–17:37 · The hosts as informed peer 4/10 Emergent User Behaviors and Multiplayer Music Creation The conversation shifts to consumer psychology where Mikey highlights how users derive fulfillment from the creation journey itself rather than just the finished track. Sarah and Mikey share excitement over users organically inventing multiplayer collaboration loops.17:38–22:46 · The hosts as informed peer 4/10 Live Suno Song Generation and Blurring Creation with Consumption The hosts and guest engage in an interactive session prompting Suno to generate a hybrid sitar-trap statistical song live. Mikey follows up by predicting that the traditional boundary between content creators and consumers will dissolve into blurred listening and creation experiences.22:47–27:13 · The hosts as informed peer 6/10 Cultural Velocity and the Long-Term Future of Music Sarah contextualizes the evolution of digital audio workstations (DAWs) to explain accessibility in music production history. Mikey builds on this by explaining that while DAWs democratized sound design, AI models will uniquely accelerate cultural velocity and compositional structure innovations.27:13–29:55 · The hosts as informed peer 4/10 Showcase Track Playback, Synthetic Audio Reflection, and Hiring After listening to a showcase track generated on Suno, Elad notes that the music, vocals, and lyrics were entirely machine generated. Mikey reflects that the underlying model operates without an explicit concept of human voice, treating all audio as unified sound.1:03–3:47 · Guest teaching 2/10 Mikey Shulman's Journey from Physics PhD to AI Sarah sets a welcoming tone by asking Mikey about his background from childhood piano to a Harvard physics PhD and founding AI startups. Mikey answers self-deprecatingly about his skills in physics and music, keeping the dialogue informal and friendly.3:47–8:50 · Guest teaching 5/10 From Speech Transcription to Musical AI Generation Elad demonstrates technical grasp of AI audio architectures by asking about diffusion versus transformer models and scaling laws. Mikey explains why Suno opted for standard transformers and focuses its technical innovation on audio tokenization heuristics at high sample rates.8:50–11:34 · Guest teaching 6/10 Avoiding Explicit Bias and Targeting Human Emotion Sarah probes the technical frontiers of audio versus video and image models. Mikey educates the hosts on why imposing music theory or fixed instrument counts creates harmful bias, comparing next-token prediction in audio directly to text LLMs.11:34–14:39 · Guest teaching 5/10 Consumer Product Strategy and Generative AI Monetization Mikey challenges conventional SaaS pricing for GenAI, teasingly critiquing SaaS-oriented venture investors and founders for lacking pricing imagination. Elad matches this analysis with historical context from the 1990s web browser era comparing micropayments and ad-based models.14:39–17:37 · Guest teaching 4/10 Emergent User Behaviors and Multiplayer Music Creation The conversation shifts to consumer psychology where Mikey highlights how users derive fulfillment from the creation journey itself rather than just the finished track. Sarah and Mikey share excitement over users organically inventing multiplayer collaboration loops.17:38–22:46 · Guest teaching 4/10 Live Suno Song Generation and Blurring Creation with Consumption The hosts and guest engage in an interactive session prompting Suno to generate a hybrid sitar-trap statistical song live. Mikey follows up by predicting that the traditional boundary between content creators and consumers will dissolve into blurred listening and creation experiences.22:47–27:13 · Guest teaching 6/10 Cultural Velocity and the Long-Term Future of Music Sarah contextualizes the evolution of digital audio workstations (DAWs) to explain accessibility in music production history. Mikey builds on this by explaining that while DAWs democratized sound design, AI models will uniquely accelerate cultural velocity and compositional structure innovations.27:13–29:55 · Guest teaching 4/10 Showcase Track Playback, Synthetic Audio Reflection, and Hiring After listening to a showcase track generated on Suno, Elad notes that the music, vocals, and lyrics were entirely machine generated. Mikey reflects that the underlying model operates without an explicit concept of human voice, treating all audio as unified sound.1:03–3:47 · Guest disagreement 1/10 Mikey Shulman's Journey from Physics PhD to AI Sarah sets a welcoming tone by asking Mikey about his background from childhood piano to a Harvard physics PhD and founding AI startups. Mikey answers self-deprecatingly about his skills in physics and music, keeping the dialogue informal and friendly.3:47–8:50 · Guest disagreement 2/10 From Speech Transcription to Musical AI Generation Elad demonstrates technical grasp of AI audio architectures by asking about diffusion versus transformer models and scaling laws. Mikey explains why Suno opted for standard transformers and focuses its technical innovation on audio tokenization heuristics at high sample rates.8:50–11:34 · Guest disagreement 2/10 Avoiding Explicit Bias and Targeting Human Emotion Sarah probes the technical frontiers of audio versus video and image models. Mikey educates the hosts on why imposing music theory or fixed instrument counts creates harmful bias, comparing next-token prediction in audio directly to text LLMs.11:34–14:39 · Guest disagreement 4/10 Consumer Product Strategy and Generative AI Monetization Mikey challenges conventional SaaS pricing for GenAI, teasingly critiquing SaaS-oriented venture investors and founders for lacking pricing imagination. Elad matches this analysis with historical context from the 1990s web browser era comparing micropayments and ad-based models.14:39–17:37 · Guest disagreement 1/10 Emergent User Behaviors and Multiplayer Music Creation The conversation shifts to consumer psychology where Mikey highlights how users derive fulfillment from the creation journey itself rather than just the finished track. Sarah and Mikey share excitement over users organically inventing multiplayer collaboration loops.17:38–22:46 · Guest disagreement 2/10 Live Suno Song Generation and Blurring Creation with Consumption The hosts and guest engage in an interactive session prompting Suno to generate a hybrid sitar-trap statistical song live. Mikey follows up by predicting that the traditional boundary between content creators and consumers will dissolve into blurred listening and creation experiences.22:47–27:13 · Guest disagreement 2/10 Cultural Velocity and the Long-Term Future of Music Sarah contextualizes the evolution of digital audio workstations (DAWs) to explain accessibility in music production history. Mikey builds on this by explaining that while DAWs democratized sound design, AI models will uniquely accelerate cultural velocity and compositional structure innovations.27:13–29:55 · Guest disagreement 1/10 Showcase Track Playback, Synthetic Audio Reflection, and Hiring After listening to a showcase track generated on Suno, Elad notes that the music, vocals, and lyrics were entirely machine generated. Mikey reflects that the underlying model operates without an explicit concept of human voice, treating all audio as unified sound.1:03–3:47 · The hosts pushing back 1/10 Mikey Shulman's Journey from Physics PhD to AI Sarah sets a welcoming tone by asking Mikey about his background from childhood piano to a Harvard physics PhD and founding AI startups. Mikey answers self-deprecatingly about his skills in physics and music, keeping the dialogue informal and friendly.3:47–8:50 · The hosts pushing back 2/10 From Speech Transcription to Musical AI Generation Elad demonstrates technical grasp of AI audio architectures by asking about diffusion versus transformer models and scaling laws. Mikey explains why Suno opted for standard transformers and focuses its technical innovation on audio tokenization heuristics at high sample rates.8:50–11:34 · The hosts pushing back 1/10 Avoiding Explicit Bias and Targeting Human Emotion Sarah probes the technical frontiers of audio versus video and image models. Mikey educates the hosts on why imposing music theory or fixed instrument counts creates harmful bias, comparing next-token prediction in audio directly to text LLMs.11:34–14:39 · The hosts pushing back 3/10 Consumer Product Strategy and Generative AI Monetization Mikey challenges conventional SaaS pricing for GenAI, teasingly critiquing SaaS-oriented venture investors and founders for lacking pricing imagination. Elad matches this analysis with historical context from the 1990s web browser era comparing micropayments and ad-based models.14:39–17:37 · The hosts pushing back 1/10 Emergent User Behaviors and Multiplayer Music Creation The conversation shifts to consumer psychology where Mikey highlights how users derive fulfillment from the creation journey itself rather than just the finished track. Sarah and Mikey share excitement over users organically inventing multiplayer collaboration loops.17:38–22:46 · The hosts pushing back 1/10 Live Suno Song Generation and Blurring Creation with Consumption The hosts and guest engage in an interactive session prompting Suno to generate a hybrid sitar-trap statistical song live. Mikey follows up by predicting that the traditional boundary between content creators and consumers will dissolve into blurred listening and creation experiences.22:47–27:13 · The hosts pushing back 2/10 Cultural Velocity and the Long-Term Future of Music Sarah contextualizes the evolution of digital audio workstations (DAWs) to explain accessibility in music production history. Mikey builds on this by explaining that while DAWs democratized sound design, AI models will uniquely accelerate cultural velocity and compositional structure innovations.27:13–29:55 · The hosts pushing back 0/10 Showcase Track Playback, Synthetic Audio Reflection, and Hiring After listening to a showcase track generated on Suno, Elad notes that the music, vocals, and lyrics were entirely machine generated. Mikey reflects that the underlying model operates without an explicit concept of human voice, treating all audio as unified sound.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 41.6% · guest 58.4%0:00 · the hosts 41.6% · guest 58.4%3:00 · the hosts 20.1% · guest 79.9%3:00 · the hosts 20.1% · guest 79.9%6:00 · the hosts 27.9% · guest 72.1%6:00 · the hosts 27.9% · guest 72.1%9:00 · the hosts 21.9% · guest 78.1%9:00 · the hosts 21.9% · guest 78.1%12:00 · the hosts 37.5% · guest 62.5%12:00 · the hosts 37.5% · guest 62.5%15:00 · the hosts 27.3% · guest 72.7%15:00 · the hosts 27.3% · guest 72.7%18:00 · the hosts 27.9% · guest 72.1%18:00 · the hosts 27.9% · guest 72.1%21:00 · the hosts 22.2% · guest 77.8%21:00 · the hosts 22.2% · guest 77.8%24:00 · the hosts 28.3% · guest 71.7%24:00 · the hosts 28.3% · guest 71.7%27:00 · the hosts 26.7% · guest 73.3%27:00 · the hosts 26.7% · guest 73.3%30:00 · the hosts 100% · guest 0%30:00 · the hosts 100% · guest 0%
Sharpest disagreement ▶ 12:20 Mikey critiques generative AI SaaS pricing as an uninspired investor vestige

Mikey takes a contrarian position on generative AI monetization, bluntly calling standard SaaS subscription tiers a crude relic inherited from past SaaS founders and venture capitalists.

Hardest push from the hosts ▶ 13:26 Elad counters pricing skepticism with 1990s web monetization history

Elad actively challenges the simplicity of current pricing debates by offering a counter-analysis of 1990s micropayments versus ad models to illustrate why early monetization paths often look crude.

Biggest teaching moment ▶ 9:15 Mikey explains why hardcoding music theory damages model potential

Mikey educates the hosts on the philosophy of music AI, explaining that baking in human music rules like twelve tones or fifty instruments inherently constrains emergent AI generation.

The host holds their own ▶ 13:26 Elad demonstrates historical business model expertise

Elad shows deep domain fluency by linking generative AI revenue models to early browser web economics, marketplace cuts, and microtransaction platforms.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Mikey Shulman's Journey from Physics PhD to AI 3211 Sarah sets a welcoming tone by asking Mikey about his background from childhood piano to a Harvard physics PhD and founding AI startups. Mikey answers self-deprecatingly about his skills in physics and music, keeping the dialogue informal and friendly.
From Speech Transcription to Musical AI Generation 6522 Elad demonstrates technical grasp of AI audio architectures by asking about diffusion versus transformer models and scaling laws. Mikey explains why Suno opted for standard transformers and focuses its technical innovation on audio tokenization heuristics at high sample rates.
Avoiding Explicit Bias and Targeting Human Emotion 5621 Sarah probes the technical frontiers of audio versus video and image models. Mikey educates the hosts on why imposing music theory or fixed instrument counts creates harmful bias, comparing next-token prediction in audio directly to text LLMs.
Consumer Product Strategy and Generative AI Monetization 7543 Mikey challenges conventional SaaS pricing for GenAI, teasingly critiquing SaaS-oriented venture investors and founders for lacking pricing imagination. Elad matches this analysis with historical context from the 1990s web browser era comparing micropayments and ad-based models.
Emergent User Behaviors and Multiplayer Music Creation 4411 The conversation shifts to consumer psychology where Mikey highlights how users derive fulfillment from the creation journey itself rather than just the finished track. Sarah and Mikey share excitement over users organically inventing multiplayer collaboration loops.
Live Suno Song Generation and Blurring Creation with Consumption 4421 The hosts and guest engage in an interactive session prompting Suno to generate a hybrid sitar-trap statistical song live. Mikey follows up by predicting that the traditional boundary between content creators and consumers will dissolve into blurred listening and creation experiences.
Cultural Velocity and the Long-Term Future of Music 6622 Sarah contextualizes the evolution of digital audio workstations (DAWs) to explain accessibility in music production history. Mikey builds on this by explaining that while DAWs democratized sound design, AI models will uniquely accelerate cultural velocity and compositional structure innovations.
Showcase Track Playback, Synthetic Audio Reflection, and Hiring 4410 After listening to a showcase track generated on Suno, Elad notes that the music, vocals, and lyrics were entirely machine generated. Mikey reflects that the underlying model operates without an explicit concept of human voice, treating all audio as unified sound.

Statements from this episode (15)

Opinion
Audio AI lags further behind text and image models today
“We also realized that certainly compared to images and text, audio was really, really far behind, and this was in 2020. And I think that's maybe even more true now, if you just look at everything that's happened in images and text in the last couple of years.”
Mikey Shulman May 16, 2024 ▶ 4:34
Insight
Speech AI demands functional accuracy while music AI requires emotional resonance
“Speech just needs to be right. Just like read me this New York Times article. And if it's a tiny bit non-expressive or a tiny bit robotic, that'll still get the job done. And the real creativity was happening in a totally different part of audio, which is musi…”
Mikey Shulman May 16, 2024 ▶ 5:20
Disclosure
Suno builds on transformers and focuses innovation on audio tokenization
“We don't make it a secret that these are just transformers. This is somewhat our backgrounds doing text before, but also transformers scale nicely. A lot of work ends up being done for you by the open source text community, which is always really nice. We can …”
Mikey Shulman May 16, 2024 ▶ 6:07
Insight
AI evaluation benchmarks fail for audio, necessitating human aesthetic judgment
“I think in, in all branches of AI, we become slaves to our metrics, and you say, I did this accuracy on this benchmark, and this accuracy on this benchmark, and in the real world, sometimes it doesn't necessarily matter, and these benchmarks are extra terrible…”
Mikey Shulman May 16, 2024 ▶ 7:11
Insight
AI music models must learn inductively without hardcoded music theory
“The model shouldn't know about music theory. You don't tell GPT, this is a noun and this is a verb. GPT figures it out. If I tell my model, there are only 12 tones. My model will only know how to output 12 tones. If I tell my model there's 50 different instrum…”
Mikey Shulman May 16, 2024 ▶ 9:35
Disclosure
Suno aims to be a mass consumer product, not a DAW plugin
“We are trying to change how the entire globe interacts with music and to open new experiences for people. And so what that means is that this is a consumer product. This is not sprinkling AI into Ableton or Logic or Pro Tools. This isn't for the person already…”
Mikey Shulman May 16, 2024 ▶ 11:58
Opinion
Generative AI SaaS pricing models are likely outdated vestiges
“Everybody's doing kind of something that looks like SAS pricing, and it's kind of done very crudely, and we are certainly no exception to that. But I don't know if this is right in the long term, and it strikes me as probably just a vestige of, it is the same …”
Mikey Shulman May 16, 2024 ▶ 12:53
Insight
AI tools let people enjoy music creation regardless of the outcome
“Music is done now, sometimes painfully, but only in service of the final product. And I think when you open this up to people sure you definitely care about the final product about what the song sounds like on the other end, but you also really cared about the…”
Mikey Shulman May 16, 2024 ▶ 15:39
Assertion Not checkable as stated
Suno users are informally hacking multiplayer collaborative workflows into the app
“Like a video game, music is fun by yourself and maybe more fun in multiplayer mode. And so we see people enjoying this by themselves, but we see people basically hacking multiplayer mode into this in, in lots of fun ways where you can have people co-writing ly…”
Mikey Shulman May 16, 2024 ▶ 16:48
Insight
AI enables selfie-like micro-sharing previously absent in traditional music
“The first is I guess all of the sort of smaller niche micro sharing that is possible where we can make songs that the three of us are going to listen to because it is capturing a moment that three of us had the same way we might take a selfie. And that is shar…”
Mikey Shulman May 16, 2024 ▶ 18:17
Prediction Not checkable as stated
Generative AI will blur the boundary between music creation and consumption
“And I think these set of technologies have the ability to skew that much, much farther because the creation process is so enjoyable. But I actually think if we do this right in the future, these are not going to be the Terms that we use to describe what we're …”
Mikey Shulman May 16, 2024 ▶ 22:10
Prediction Not checkable as stated
Generative AI will dramatically increase time and money spent on music
“If we are correct that there are just modes of experience around music that people don't have access to, that we can get a billion people much more engaged with music than they are now, that just in terms of the number of dollars or the amount of time people a…”
Mikey Shulman May 16, 2024 ▶ 23:02
Opinion
Listeners will not lose their emotional connection to human artists
“Because music is so human and so much emotional connection involved in it, I don't really see people losing connection with their favorite artists at all. In fact, if you labor around music and you understand the process, you feel a much deeper connection with…”
Mikey Shulman May 16, 2024 ▶ 23:30
Prediction Not checkable as stated
Democratized AI tools will accelerate the evolution of musical styles
“Rate at which culture changes, the rate at which the styles of music change, the rate at which new styles of music are uncovered is likely to go up a lot.”
Mikey Shulman May 16, 2024 ▶ 24:29
Insight
Suno's AI model generates vocals without an explicit concept of human voice
“The machine doesn't know that there is even a concept of voice. Like it's just all sound and somehow it's able to produce the sounds that we have been evolved and acculturated to resonate with.”
Mikey Shulman May 16, 2024 ▶ 29:00
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.