Mar 14, 2024 · 58m · latent-space

Making Transformers Sing - with Mikey Shulman of Suno

Mikey Shulman · 32m spoken Alessio Fanelli · 7m spoken Shawn Wang · 6m spoken Generated Vocals · 4m spoken Generated Audio · 3s spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Live in Space podcast, co-hosts Alessio Fanelli and Shawn Wang interview Suno co-founder Mikey Shulman about the technical architecture, product design, and creative vision behind their state-of-the-art music generation platform. Shulman provides live v3 model demonstrations while discussing audio tokenization, consumer-focused workflows, and the future of participatory music creation.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 14% of the talking time here. How this is scored →

The hosts as informed peer 5.2 Guest teaching 5.1 Guest disagreement 1.1 The hosts pushing back 1.1
05100:0015:0030:0045:000:01–6:32 · The hosts as informed peer 5/10 Introductions and Mikey Shulman's Background Swix and Alessio open with Mikey's background and ask how music generation works compared to text LLMs and image diffusion. Mikey explains that audio lags text/vision by two years and details how transformer autoregression operates on discrete audio tokens.6:33–9:41 · The hosts as informed peer 6/10 Training Data Recipes and Multimodal Audio Learning Alessio asks about the training dataset recipe and whether high-quality music follows power laws similar to text LLM datasets. Mikey reveals that Suno trains models on non-musical human vocal audio alongside music to improve vocal synthesis realism.9:41–11:57 · The hosts as informed peer 5/10 Scaling Audio Models and Real-Time Latency Constraints Alessio queries model parameters and local execution feasibility. Mikey discusses real-time streaming latency constraints and explains why relying purely on massive scale can be a research crutch that discourages model optimization.11:58–15:20 · The hosts as informed peer 4/10 The Origins of Suno and Choosing Music Over Speech Swix asks why Mikey chose music over more commercially conventional speech applications, playfully pointing out recent speech collaborations. Mikey explains that emotional resonance and fun drove their organic focus towards music.15:20–18:40 · The hosts as informed peer 6/10 Open-Source Foundations of Bark and Latent Music Behavior Swix probes Bark's lineage, initially mischaracterizing it as speech recognition before Mikey clarifies it is text-to-speech built on nanoGPT concepts. Mikey explains how self-supervised training allowed latent musical behavior to emerge naturally.18:42–22:06 · The hosts as informed peer 5/10 Suno User Modes and Social Dynamics of Music Sharing Alessio asks about casual vs. power user behavior on Suno. Mikey reveals over half of users engage with expert mode for fine-grained lyrical tweaking and discusses how micro-sharing music brings joy to small social circles.22:07–26:38 · The hosts as informed peer 6/10 Consumer Workflows Versus Professional Music Production Alessio references Madlib's iPad production to question if Suno aims to capture professional DAW workflows. Mikey clarifies that Suno deliberately targets consumer participation rather than incremental productivity for professional audio engineers.26:38–29:29 · The hosts as informed peer 4/10 Live Demo: Generating GPU Cloud Blues and Country Songs Alessio prompts Suno to generate a country song about missing cloud GPUs. The hosts and Mikey analyze the rapid latency, vocal generation, and domain knowledge reflected in the generated output.29:30–33:05 · The hosts as informed peer 5/10 Live Demo: Generating House Music and Prompt Tuning Swix requests a house music track about podcasting, prompting Mikey to demonstrate real-time prompt modification. Swix notices special tokens like beat drops and asks if they are training artifacts, which Mikey clarifies are emergent user prompt strategies.33:06–40:20 · The hosts as informed peer 5/10 Live Demo: Style Modulation and Apple Vision Pro Blues The hosts test model boundaries with Apple Vision Pro themes and blues styles. Mikey discusses guardrails against impersonation when location names like Chicago trigger filters and explains why repeating prompt keywords too often degrades output.40:21–42:41 · The hosts as informed peer 6/10 Copyright Policy, Novelty Covers, and Original Music Alessio brings up sample flipping and remixing classic tracks in hip-hop. Mikey emphasizes that Suno intentionally blocks copyrighted lyric reuse to steer users toward creating original music rather than viral novelty covers.42:42–47:29 · The hosts as informed peer 5/10 Future Roadmap: Collaborative Concerts and Active Music Mikey contrasts the 50x larger gaming industry with music's passive consumption model. Swix suggests a Twitch Plays Pokemon radio stream while Mikey shares his vision for interactive, audience-driven collaborative concerts.47:30–51:56 · The hosts as informed peer 5/10 Model Personalization, Feedback Loops, and User Impact Swix asks how Suno handles subjective audio quality feedback loops and artifacts. Mikey details plans for personalized producer models and highlights Suno's adoption within visually impaired communities.51:57–54:24 · The hosts as informed peer 5/10 The Generative Audio Landscape and Music Production Tools Swix asks Mikey to map out the generative audio ecosystem. Mikey categorizes the landscape into royalty-free stock music, AI covers, consumer original generation, and professional DAW plugins/stem splitters.54:25–58:33 · The hosts as informed peer 6/10 Goodhart's Law, Aesthetic Evaluation, and Conclusion Alessio references Mikey's writing on Goodhart's Law. Mikey explains why quantitative ML benchmarks fall short in audio evaluation and why social scientists and economists make great ML engineers by thinking from first principles.0:01–6:32 · Guest teaching 5/10 Introductions and Mikey Shulman's Background Swix and Alessio open with Mikey's background and ask how music generation works compared to text LLMs and image diffusion. Mikey explains that audio lags text/vision by two years and details how transformer autoregression operates on discrete audio tokens.6:33–9:41 · Guest teaching 6/10 Training Data Recipes and Multimodal Audio Learning Alessio asks about the training dataset recipe and whether high-quality music follows power laws similar to text LLM datasets. Mikey reveals that Suno trains models on non-musical human vocal audio alongside music to improve vocal synthesis realism.9:41–11:57 · Guest teaching 5/10 Scaling Audio Models and Real-Time Latency Constraints Alessio queries model parameters and local execution feasibility. Mikey discusses real-time streaming latency constraints and explains why relying purely on massive scale can be a research crutch that discourages model optimization.11:58–15:20 · Guest teaching 4/10 The Origins of Suno and Choosing Music Over Speech Swix asks why Mikey chose music over more commercially conventional speech applications, playfully pointing out recent speech collaborations. Mikey explains that emotional resonance and fun drove their organic focus towards music.15:20–18:40 · Guest teaching 5/10 Open-Source Foundations of Bark and Latent Music Behavior Swix probes Bark's lineage, initially mischaracterizing it as speech recognition before Mikey clarifies it is text-to-speech built on nanoGPT concepts. Mikey explains how self-supervised training allowed latent musical behavior to emerge naturally.18:42–22:06 · Guest teaching 5/10 Suno User Modes and Social Dynamics of Music Sharing Alessio asks about casual vs. power user behavior on Suno. Mikey reveals over half of users engage with expert mode for fine-grained lyrical tweaking and discusses how micro-sharing music brings joy to small social circles.22:07–26:38 · Guest teaching 5/10 Consumer Workflows Versus Professional Music Production Alessio references Madlib's iPad production to question if Suno aims to capture professional DAW workflows. Mikey clarifies that Suno deliberately targets consumer participation rather than incremental productivity for professional audio engineers.26:38–29:29 · Guest teaching 4/10 Live Demo: Generating GPU Cloud Blues and Country Songs Alessio prompts Suno to generate a country song about missing cloud GPUs. The hosts and Mikey analyze the rapid latency, vocal generation, and domain knowledge reflected in the generated output.29:30–33:05 · Guest teaching 4/10 Live Demo: Generating House Music and Prompt Tuning Swix requests a house music track about podcasting, prompting Mikey to demonstrate real-time prompt modification. Swix notices special tokens like beat drops and asks if they are training artifacts, which Mikey clarifies are emergent user prompt strategies.33:06–40:20 · Guest teaching 5/10 Live Demo: Style Modulation and Apple Vision Pro Blues The hosts test model boundaries with Apple Vision Pro themes and blues styles. Mikey discusses guardrails against impersonation when location names like Chicago trigger filters and explains why repeating prompt keywords too often degrades output.40:21–42:41 · Guest teaching 6/10 Copyright Policy, Novelty Covers, and Original Music Alessio brings up sample flipping and remixing classic tracks in hip-hop. Mikey emphasizes that Suno intentionally blocks copyrighted lyric reuse to steer users toward creating original music rather than viral novelty covers.42:42–47:29 · Guest teaching 5/10 Future Roadmap: Collaborative Concerts and Active Music Mikey contrasts the 50x larger gaming industry with music's passive consumption model. Swix suggests a Twitch Plays Pokemon radio stream while Mikey shares his vision for interactive, audience-driven collaborative concerts.47:30–51:56 · Guest teaching 5/10 Model Personalization, Feedback Loops, and User Impact Swix asks how Suno handles subjective audio quality feedback loops and artifacts. Mikey details plans for personalized producer models and highlights Suno's adoption within visually impaired communities.51:57–54:24 · Guest teaching 6/10 The Generative Audio Landscape and Music Production Tools Swix asks Mikey to map out the generative audio ecosystem. Mikey categorizes the landscape into royalty-free stock music, AI covers, consumer original generation, and professional DAW plugins/stem splitters.54:25–58:33 · Guest teaching 6/10 Goodhart's Law, Aesthetic Evaluation, and Conclusion Alessio references Mikey's writing on Goodhart's Law. Mikey explains why quantitative ML benchmarks fall short in audio evaluation and why social scientists and economists make great ML engineers by thinking from first principles.0:01–6:32 · Guest disagreement 1/10 Introductions and Mikey Shulman's Background Swix and Alessio open with Mikey's background and ask how music generation works compared to text LLMs and image diffusion. Mikey explains that audio lags text/vision by two years and details how transformer autoregression operates on discrete audio tokens.6:33–9:41 · Guest disagreement 1/10 Training Data Recipes and Multimodal Audio Learning Alessio asks about the training dataset recipe and whether high-quality music follows power laws similar to text LLM datasets. Mikey reveals that Suno trains models on non-musical human vocal audio alongside music to improve vocal synthesis realism.9:41–11:57 · Guest disagreement 1/10 Scaling Audio Models and Real-Time Latency Constraints Alessio queries model parameters and local execution feasibility. Mikey discusses real-time streaming latency constraints and explains why relying purely on massive scale can be a research crutch that discourages model optimization.11:58–15:20 · Guest disagreement 1/10 The Origins of Suno and Choosing Music Over Speech Swix asks why Mikey chose music over more commercially conventional speech applications, playfully pointing out recent speech collaborations. Mikey explains that emotional resonance and fun drove their organic focus towards music.15:20–18:40 · Guest disagreement 2/10 Open-Source Foundations of Bark and Latent Music Behavior Swix probes Bark's lineage, initially mischaracterizing it as speech recognition before Mikey clarifies it is text-to-speech built on nanoGPT concepts. Mikey explains how self-supervised training allowed latent musical behavior to emerge naturally.18:42–22:06 · Guest disagreement 1/10 Suno User Modes and Social Dynamics of Music Sharing Alessio asks about casual vs. power user behavior on Suno. Mikey reveals over half of users engage with expert mode for fine-grained lyrical tweaking and discusses how micro-sharing music brings joy to small social circles.22:07–26:38 · Guest disagreement 1/10 Consumer Workflows Versus Professional Music Production Alessio references Madlib's iPad production to question if Suno aims to capture professional DAW workflows. Mikey clarifies that Suno deliberately targets consumer participation rather than incremental productivity for professional audio engineers.26:38–29:29 · Guest disagreement 1/10 Live Demo: Generating GPU Cloud Blues and Country Songs Alessio prompts Suno to generate a country song about missing cloud GPUs. The hosts and Mikey analyze the rapid latency, vocal generation, and domain knowledge reflected in the generated output.29:30–33:05 · Guest disagreement 1/10 Live Demo: Generating House Music and Prompt Tuning Swix requests a house music track about podcasting, prompting Mikey to demonstrate real-time prompt modification. Swix notices special tokens like beat drops and asks if they are training artifacts, which Mikey clarifies are emergent user prompt strategies.33:06–40:20 · Guest disagreement 1/10 Live Demo: Style Modulation and Apple Vision Pro Blues The hosts test model boundaries with Apple Vision Pro themes and blues styles. Mikey discusses guardrails against impersonation when location names like Chicago trigger filters and explains why repeating prompt keywords too often degrades output.40:21–42:41 · Guest disagreement 2/10 Copyright Policy, Novelty Covers, and Original Music Alessio brings up sample flipping and remixing classic tracks in hip-hop. Mikey emphasizes that Suno intentionally blocks copyrighted lyric reuse to steer users toward creating original music rather than viral novelty covers.42:42–47:29 · Guest disagreement 1/10 Future Roadmap: Collaborative Concerts and Active Music Mikey contrasts the 50x larger gaming industry with music's passive consumption model. Swix suggests a Twitch Plays Pokemon radio stream while Mikey shares his vision for interactive, audience-driven collaborative concerts.47:30–51:56 · Guest disagreement 1/10 Model Personalization, Feedback Loops, and User Impact Swix asks how Suno handles subjective audio quality feedback loops and artifacts. Mikey details plans for personalized producer models and highlights Suno's adoption within visually impaired communities.51:57–54:24 · Guest disagreement 1/10 The Generative Audio Landscape and Music Production Tools Swix asks Mikey to map out the generative audio ecosystem. Mikey categorizes the landscape into royalty-free stock music, AI covers, consumer original generation, and professional DAW plugins/stem splitters.54:25–58:33 · Guest disagreement 1/10 Goodhart's Law, Aesthetic Evaluation, and Conclusion Alessio references Mikey's writing on Goodhart's Law. Mikey explains why quantitative ML benchmarks fall short in audio evaluation and why social scientists and economists make great ML engineers by thinking from first principles.0:01–6:32 · The hosts pushing back 1/10 Introductions and Mikey Shulman's Background Swix and Alessio open with Mikey's background and ask how music generation works compared to text LLMs and image diffusion. Mikey explains that audio lags text/vision by two years and details how transformer autoregression operates on discrete audio tokens.6:33–9:41 · The hosts pushing back 1/10 Training Data Recipes and Multimodal Audio Learning Alessio asks about the training dataset recipe and whether high-quality music follows power laws similar to text LLM datasets. Mikey reveals that Suno trains models on non-musical human vocal audio alongside music to improve vocal synthesis realism.9:41–11:57 · The hosts pushing back 1/10 Scaling Audio Models and Real-Time Latency Constraints Alessio queries model parameters and local execution feasibility. Mikey discusses real-time streaming latency constraints and explains why relying purely on massive scale can be a research crutch that discourages model optimization.11:58–15:20 · The hosts pushing back 2/10 The Origins of Suno and Choosing Music Over Speech Swix asks why Mikey chose music over more commercially conventional speech applications, playfully pointing out recent speech collaborations. Mikey explains that emotional resonance and fun drove their organic focus towards music.15:20–18:40 · The hosts pushing back 1/10 Open-Source Foundations of Bark and Latent Music Behavior Swix probes Bark's lineage, initially mischaracterizing it as speech recognition before Mikey clarifies it is text-to-speech built on nanoGPT concepts. Mikey explains how self-supervised training allowed latent musical behavior to emerge naturally.18:42–22:06 · The hosts pushing back 1/10 Suno User Modes and Social Dynamics of Music Sharing Alessio asks about casual vs. power user behavior on Suno. Mikey reveals over half of users engage with expert mode for fine-grained lyrical tweaking and discusses how micro-sharing music brings joy to small social circles.22:07–26:38 · The hosts pushing back 2/10 Consumer Workflows Versus Professional Music Production Alessio references Madlib's iPad production to question if Suno aims to capture professional DAW workflows. Mikey clarifies that Suno deliberately targets consumer participation rather than incremental productivity for professional audio engineers.26:38–29:29 · The hosts pushing back 1/10 Live Demo: Generating GPU Cloud Blues and Country Songs Alessio prompts Suno to generate a country song about missing cloud GPUs. The hosts and Mikey analyze the rapid latency, vocal generation, and domain knowledge reflected in the generated output.29:30–33:05 · The hosts pushing back 1/10 Live Demo: Generating House Music and Prompt Tuning Swix requests a house music track about podcasting, prompting Mikey to demonstrate real-time prompt modification. Swix notices special tokens like beat drops and asks if they are training artifacts, which Mikey clarifies are emergent user prompt strategies.33:06–40:20 · The hosts pushing back 1/10 Live Demo: Style Modulation and Apple Vision Pro Blues The hosts test model boundaries with Apple Vision Pro themes and blues styles. Mikey discusses guardrails against impersonation when location names like Chicago trigger filters and explains why repeating prompt keywords too often degrades output.40:21–42:41 · The hosts pushing back 1/10 Copyright Policy, Novelty Covers, and Original Music Alessio brings up sample flipping and remixing classic tracks in hip-hop. Mikey emphasizes that Suno intentionally blocks copyrighted lyric reuse to steer users toward creating original music rather than viral novelty covers.42:42–47:29 · The hosts pushing back 1/10 Future Roadmap: Collaborative Concerts and Active Music Mikey contrasts the 50x larger gaming industry with music's passive consumption model. Swix suggests a Twitch Plays Pokemon radio stream while Mikey shares his vision for interactive, audience-driven collaborative concerts.47:30–51:56 · The hosts pushing back 1/10 Model Personalization, Feedback Loops, and User Impact Swix asks how Suno handles subjective audio quality feedback loops and artifacts. Mikey details plans for personalized producer models and highlights Suno's adoption within visually impaired communities.51:57–54:24 · The hosts pushing back 1/10 The Generative Audio Landscape and Music Production Tools Swix asks Mikey to map out the generative audio ecosystem. Mikey categorizes the landscape into royalty-free stock music, AI covers, consumer original generation, and professional DAW plugins/stem splitters.54:25–58:33 · The hosts pushing back 1/10 Goodhart's Law, Aesthetic Evaluation, and Conclusion Alessio references Mikey's writing on Goodhart's Law. Mikey explains why quantitative ML benchmarks fall short in audio evaluation and why social scientists and economists make great ML engineers by thinking from first principles.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 41.6% · guest 58.4%0:00 · the hosts 41.6% · guest 58.4%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 32.1% · guest 67.9%6:00 · the hosts 32.1% · guest 67.9%9:00 · the hosts 16.1% · guest 83.9%9:00 · the hosts 16.1% · guest 83.9%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 23.1% · guest 76.9%18:00 · the hosts 23.1% · guest 76.9%21:00 · the hosts 30% · guest 70%21:00 · the hosts 30% · guest 70%24:00 · the hosts 4.4% · guest 95.6%24:00 · the hosts 4.4% · guest 95.6%27:00 · the hosts 9.2% · guest 90.8%27:00 · the hosts 9.2% · guest 90.8%30:00 · the hosts 13.3% · guest 86.7%30:00 · the hosts 13.3% · guest 86.7%33:00 · the hosts 4.1% · guest 95.9%33:00 · the hosts 4.1% · guest 95.9%36:00 · the hosts 3% · guest 97%36:00 · the hosts 3% · guest 97%39:00 · the hosts 23.3% · guest 76.7%39:00 · the hosts 23.3% · guest 76.7%42:00 · the hosts 33.6% · guest 66.4%42:00 · the hosts 33.6% · guest 66.4%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 17.8% · guest 82.2%48:00 · the hosts 17.8% · guest 82.2%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 13.9% · guest 86.1%54:00 · the hosts 13.9% · guest 86.1%57:00 · the hosts 10.7% · guest 89.3%57:00 · the hosts 10.7% · guest 89.3%
Sharpest disagreement ▶ 40:56 Pushing back against AI novelty covers

Mikey firmly rejects the popular AI music trend of voice cloning and novelty covers, comparing them to disposable ChatGPT sonnets and explaining Suno's active restrictions on copyrighted lyrics.

Hardest push from the hosts ▶ 14:15 Challenging Suno's music-only positioning

Swix playfully challenges Mikey's assertion that Suno is exclusively a music company by citing their recent speech model release with NVIDIA.

Biggest teaching moment ▶ 8:30 Non-musical vocal datasets for singing synthesis

Mikey educates the hosts on multimodal training recipes, explaining that realistic singing voice generation requires incorporating non-musical human vocal data into the training corpus.

The host holds their own ▶ 54:24 Framing ML evaluation through Goodhart's Law

Alessio demonstrates deep familiarity with Mikey's technical writing by bringing up his Kensho blog post on Goodhart's Law to explore LLM benchmark limitations.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Introductions and Mikey Shulman's Background 5511 Swix and Alessio open with Mikey's background and ask how music generation works compared to text LLMs and image diffusion. Mikey explains that audio lags text/vision by two years and details how transformer autoregression operates on discrete audio tokens.
Training Data Recipes and Multimodal Audio Learning 6611 Alessio asks about the training dataset recipe and whether high-quality music follows power laws similar to text LLM datasets. Mikey reveals that Suno trains models on non-musical human vocal audio alongside music to improve vocal synthesis realism.
Scaling Audio Models and Real-Time Latency Constraints 5511 Alessio queries model parameters and local execution feasibility. Mikey discusses real-time streaming latency constraints and explains why relying purely on massive scale can be a research crutch that discourages model optimization.
The Origins of Suno and Choosing Music Over Speech 4412 Swix asks why Mikey chose music over more commercially conventional speech applications, playfully pointing out recent speech collaborations. Mikey explains that emotional resonance and fun drove their organic focus towards music.
Open-Source Foundations of Bark and Latent Music Behavior 6521 Swix probes Bark's lineage, initially mischaracterizing it as speech recognition before Mikey clarifies it is text-to-speech built on nanoGPT concepts. Mikey explains how self-supervised training allowed latent musical behavior to emerge naturally.
Suno User Modes and Social Dynamics of Music Sharing 5511 Alessio asks about casual vs. power user behavior on Suno. Mikey reveals over half of users engage with expert mode for fine-grained lyrical tweaking and discusses how micro-sharing music brings joy to small social circles.
Consumer Workflows Versus Professional Music Production 6512 Alessio references Madlib's iPad production to question if Suno aims to capture professional DAW workflows. Mikey clarifies that Suno deliberately targets consumer participation rather than incremental productivity for professional audio engineers.
Live Demo: Generating GPU Cloud Blues and Country Songs 4411 Alessio prompts Suno to generate a country song about missing cloud GPUs. The hosts and Mikey analyze the rapid latency, vocal generation, and domain knowledge reflected in the generated output.
Live Demo: Generating House Music and Prompt Tuning 5411 Swix requests a house music track about podcasting, prompting Mikey to demonstrate real-time prompt modification. Swix notices special tokens like beat drops and asks if they are training artifacts, which Mikey clarifies are emergent user prompt strategies.
Live Demo: Style Modulation and Apple Vision Pro Blues 5511 The hosts test model boundaries with Apple Vision Pro themes and blues styles. Mikey discusses guardrails against impersonation when location names like Chicago trigger filters and explains why repeating prompt keywords too often degrades output.
Copyright Policy, Novelty Covers, and Original Music 6621 Alessio brings up sample flipping and remixing classic tracks in hip-hop. Mikey emphasizes that Suno intentionally blocks copyrighted lyric reuse to steer users toward creating original music rather than viral novelty covers.
Future Roadmap: Collaborative Concerts and Active Music 5511 Mikey contrasts the 50x larger gaming industry with music's passive consumption model. Swix suggests a Twitch Plays Pokemon radio stream while Mikey shares his vision for interactive, audience-driven collaborative concerts.
Model Personalization, Feedback Loops, and User Impact 5511 Swix asks how Suno handles subjective audio quality feedback loops and artifacts. Mikey details plans for personalized producer models and highlights Suno's adoption within visually impaired communities.
The Generative Audio Landscape and Music Production Tools 5611 Swix asks Mikey to map out the generative audio ecosystem. Mikey categorizes the landscape into royalty-free stock music, AI covers, consumer original generation, and professional DAW plugins/stem splitters.
Goodhart's Law, Aesthetic Evaluation, and Conclusion 6611 Alessio references Mikey's writing on Goodhart's Law. Mikey explains why quantitative ML benchmarks fall short in audio evaluation and why social scientists and economists make great ML engineers by thinking from first principles.

Statements from this episode (23)

Opinion
Shulman: Generative audio AI lags text and images by one to two years
“So I think very roughly you can think audio is like one to two years behind images and text. And so you kind of have to think today like text was in 20, 22 or something like this.”
Mikey Shulman Mar 14, 2024 ▶ 3:00
Disclosure
Suno avoids hardcoding musical rules into its generative models
“We try not to impose anything about music or audio in general into the model, and we kind of let the models learn things by themselves.”
Mikey Shulman Mar 14, 2024 ▶ 5:42
Disclosure
Shulman: Suno focuses primarily on audio tokenization to leverage text transformers
“What we do is we benefit from all of the beautiful things people do with transformers and text, and we focus very hard basically on how do I tokenize audio in the right way.”
Mikey Shulman Mar 14, 2024 ▶ 6:54
Insight
Shulman: Understanding of music scaling laws is far behind text
“I don't think we know these things nearly as well as they're known in text. We have some notions of some of the scaling laws here, but I think yeah, we're just so, so far behind.”
Mikey Shulman Mar 14, 2024 ▶ 8:28
Disclosure
Shulman: Suno does not only train its models on music
“People are always surprised to learn that we don't only train on music.”
Mikey Shulman Mar 14, 2024 ▶ 8:40
Disclosure
Shulman: Suno is particularly bad at capturing realistic vocals
“One of the places that we are particularly bad is vocals and at capturing really realistic vocals.”
Mikey Shulman Mar 14, 2024 ▶ 9:13
Prediction Not checkable as stated
Streaming throughput constraints prevent Suno from scaling to 175 billion parameters
“We care a lot about how many tokens per second we can generate because we need to stream new music as fast as you can listen to it. And so that is a big one that I think probably has us never get to a hundred seventy-five billion parameter model, if I'm being …”
Mikey Shulman Mar 14, 2024 ▶ 10:30
Assertion Supported
Shulman: Suno's first release was open-source TTS model Bark
“In fact, we, the first thing we ever put out was a speech model. It was Bark. It was this open source text-to-speech model, and it got a lot of stars on GitHub,”
Mikey Shulman Mar 14, 2024 ▶ 13:49
Assertion Contradicted
Shulman: Bark was the first open-source transformer-based TTS model
“As far as I know there was no other certainly not in the open source text to speech that was kind of transformer based.”
Mikey Shulman Mar 14, 2024 ▶ 15:48
Disclosure
Shulman: Suno borrowed significant code from Andrej Karpathy's nanoGPT for Bark
“There's a big shout out to Andre Karpathy's nano GPT. You know, there's a lot of code borrowed from there.”
Mikey Shulman Mar 14, 2024 ▶ 16:46
Disclosure
Shulman: Suno aims to undo cultural barriers preventing people from making music
“There's actually a lot of cultural forces that kind of cue you to not think to make music and that's kind of what we're trying to undo.”
Mikey Shulman Mar 14, 2024 ▶ 18:34
Assertion Not checkable as stated
Shulman: Over half of Suno usage is in expert mode
“Yeah, actually more than half of the usage is that expert mode.”
Mikey Shulman Mar 14, 2024 ▶ 19:24
Prediction Not checkable as stated
Shulman: Accessible AI music will drive private small-group sharing
“But when you start to make that more accessible to people, they are going to share music in much smaller groups, maybe even not at all, but like with one person or three people or five people.”
Mikey Shulman Mar 14, 2024 ▶ 21:24
Assertion Supported
Fanelli: Madlib produced full Freddie Gibbs album entirely on iPad
“Madlib actually produced this whole album with him and Freddie Gibbs produced the whole thing on an iPad. He never used a computer.”
Alessio Fanelli Mar 14, 2024 ▶ 22:24
Assertion Not checkable as stated
Shulman: Professional musicians are using Suno for inspiration and sample generation
“There are lots of professionals that we know about using our stuff, whether it's for inspiration or sample generation and stuff like that.”
Mikey Shulman Mar 14, 2024 ▶ 24:10
Disclosure
Shulman: Suno restricts prompts to prevent artist impersonation
“We try to be very careful not letting you impersonate, and it is possible.”
Mikey Shulman Mar 14, 2024 ▶ 37:18
Insight
Shulman: Repeating prompt keywords too often degrades music model outputs
“It's actually, you can't really repeat too many times. You kind of, it gets like the hypothesis gets like a little too out of domain.”
Mikey Shulman Mar 14, 2024 ▶ 39:15
Disclosure
Shulman: Suno prohibits users from generating songs using copyrighted lyrics
“We actually don't let you do that. And it's because if you're taking someone else's lyrics, you didn't own those. You don't have the publishing rights to those. You can't remake that song.”
Mikey Shulman Mar 14, 2024 ▶ 40:57
Prediction Not checkable as stated
Shulman: AI song covers are viral novelties, not music's future
“And I think this stuff is very viral, but I actually really don't think that this is how people want to interact with music in the future.”
Mikey Shulman Mar 14, 2024 ▶ 41:26
Insight
Shulman: Gaming dwarfs music 50x because music is mostly passive consumption
“And, you know, as a very broad heuristic, the gaming industry is 50 times bigger than the music industry. And it's because gaming is super active. And music, too much music is just passive consumption.”
Mikey Shulman Mar 14, 2024 ▶ 45:04
Assertion Not checkable as stated
Suno is seeing significant adoption among blind and vision-impaired users
“We're fairly popular in the blind and vision impaired community”
Mikey Shulman Mar 14, 2024 ▶ 50:23
Opinion
Shulman: AI evaluation benchmarks are far worse in audio than text
“As flawed as these benchmarks are in text, they're way worse in audio.”
Mikey Shulman Mar 14, 2024 ▶ 55:42
Disclosure
Shulman: Kensho recruited from economics conferences for top ML hires
“At Kensho, we actually used to go to big econ conferences sometimes to recruit, and these were some of the best hires we ever made.”
Mikey Shulman Mar 14, 2024 ▶ 56:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.