Mar 14, 2024 · 58m · latent-space
Making Transformers Sing - with Mikey Shulman of Suno
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Live in Space podcast, co-hosts Alessio Fanelli and Shawn Wang interview Suno co-founder Mikey Shulman about the technical architecture, product design, and creative vision behind their state-of-the-art music generation platform. Shulman provides live v3 model demonstrations while discussing audio tokenization, consumer-focused workflows, and the future of participatory music creation.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 14% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Mikey firmly rejects the popular AI music trend of voice cloning and novelty covers, comparing them to disposable ChatGPT sonnets and explaining Suno's active restrictions on copyrighted lyrics.
Hardest push from the hosts ▶ 14:15 Challenging Suno's music-only positioningSwix playfully challenges Mikey's assertion that Suno is exclusively a music company by citing their recent speech model release with NVIDIA.
Biggest teaching moment ▶ 8:30 Non-musical vocal datasets for singing synthesisMikey educates the hosts on multimodal training recipes, explaining that realistic singing voice generation requires incorporating non-musical human vocal data into the training corpus.
The host holds their own ▶ 54:24 Framing ML evaluation through Goodhart's LawAlessio demonstrates deep familiarity with Mikey's technical writing by bringing up his Kensho blog post on Goodhart's Law to explore LLM benchmark limitations.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Introductions and Mikey Shulman's Background | 5 | 5 | 1 | 1 | Swix and Alessio open with Mikey's background and ask how music generation works compared to text LLMs and image diffusion. Mikey explains that audio lags text/vision by two years and details how transformer autoregression operates on discrete audio tokens. | |
| Training Data Recipes and Multimodal Audio Learning | 6 | 6 | 1 | 1 | Alessio asks about the training dataset recipe and whether high-quality music follows power laws similar to text LLM datasets. Mikey reveals that Suno trains models on non-musical human vocal audio alongside music to improve vocal synthesis realism. | |
| Scaling Audio Models and Real-Time Latency Constraints | 5 | 5 | 1 | 1 | Alessio queries model parameters and local execution feasibility. Mikey discusses real-time streaming latency constraints and explains why relying purely on massive scale can be a research crutch that discourages model optimization. | |
| The Origins of Suno and Choosing Music Over Speech | 4 | 4 | 1 | 2 | Swix asks why Mikey chose music over more commercially conventional speech applications, playfully pointing out recent speech collaborations. Mikey explains that emotional resonance and fun drove their organic focus towards music. | |
| Open-Source Foundations of Bark and Latent Music Behavior | 6 | 5 | 2 | 1 | Swix probes Bark's lineage, initially mischaracterizing it as speech recognition before Mikey clarifies it is text-to-speech built on nanoGPT concepts. Mikey explains how self-supervised training allowed latent musical behavior to emerge naturally. | |
| Suno User Modes and Social Dynamics of Music Sharing | 5 | 5 | 1 | 1 | Alessio asks about casual vs. power user behavior on Suno. Mikey reveals over half of users engage with expert mode for fine-grained lyrical tweaking and discusses how micro-sharing music brings joy to small social circles. | |
| Consumer Workflows Versus Professional Music Production | 6 | 5 | 1 | 2 | Alessio references Madlib's iPad production to question if Suno aims to capture professional DAW workflows. Mikey clarifies that Suno deliberately targets consumer participation rather than incremental productivity for professional audio engineers. | |
| Live Demo: Generating GPU Cloud Blues and Country Songs | 4 | 4 | 1 | 1 | Alessio prompts Suno to generate a country song about missing cloud GPUs. The hosts and Mikey analyze the rapid latency, vocal generation, and domain knowledge reflected in the generated output. | |
| Live Demo: Generating House Music and Prompt Tuning | 5 | 4 | 1 | 1 | Swix requests a house music track about podcasting, prompting Mikey to demonstrate real-time prompt modification. Swix notices special tokens like beat drops and asks if they are training artifacts, which Mikey clarifies are emergent user prompt strategies. | |
| Live Demo: Style Modulation and Apple Vision Pro Blues | 5 | 5 | 1 | 1 | The hosts test model boundaries with Apple Vision Pro themes and blues styles. Mikey discusses guardrails against impersonation when location names like Chicago trigger filters and explains why repeating prompt keywords too often degrades output. | |
| Copyright Policy, Novelty Covers, and Original Music | 6 | 6 | 2 | 1 | Alessio brings up sample flipping and remixing classic tracks in hip-hop. Mikey emphasizes that Suno intentionally blocks copyrighted lyric reuse to steer users toward creating original music rather than viral novelty covers. | |
| Future Roadmap: Collaborative Concerts and Active Music | 5 | 5 | 1 | 1 | Mikey contrasts the 50x larger gaming industry with music's passive consumption model. Swix suggests a Twitch Plays Pokemon radio stream while Mikey shares his vision for interactive, audience-driven collaborative concerts. | |
| Model Personalization, Feedback Loops, and User Impact | 5 | 5 | 1 | 1 | Swix asks how Suno handles subjective audio quality feedback loops and artifacts. Mikey details plans for personalized producer models and highlights Suno's adoption within visually impaired communities. | |
| The Generative Audio Landscape and Music Production Tools | 5 | 6 | 1 | 1 | Swix asks Mikey to map out the generative audio ecosystem. Mikey categorizes the landscape into royalty-free stock music, AI covers, consumer original generation, and professional DAW plugins/stem splitters. | |
| Goodhart's Law, Aesthetic Evaluation, and Conclusion | 6 | 6 | 1 | 1 | Alessio references Mikey's writing on Goodhart's Law. Mikey explains why quantitative ML benchmarks fall short in audio evaluation and why social scientists and economists make great ML engineers by thinking from first principles. |