The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mikey Shulman no published score: only 7 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 7 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
7exchanges match
7on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q can set your own lyrics, you can set your own rhythm, you can set the title of the song and whatnot. What are, how do you see users distribute themselves? You know, I'm guessing a lot of people use the easy mode. Like, are you seeing a lot of power users using the custom mode and maybe some of the favorite use cases that you've seen so far on Suno?

A Yeah, actually, um, more than half of the usage is, uh, that expert mode. And people really like to get into it and start tweaking things and adding things and playing with words or line breaks or different ad lib and people really love it. It's, it's, it's really fun. So I think, you know, there's kind of two modes that you can access now. One is that single box where you kind of just describe something, and then the other is the expert mode, and, um, those kind of fit nicely into two use cases. The first use case is what we call nice shit posting, and it's basically like something funny happened, and I'm just going to very quickly make a song about it, and the, the example I I'll usually give is like, I walk into Starbucks with one of my co-founders, he gives his name Martin, his coffee comes out, Um, with the name Margu and I can in five seconds make a song about this and it has immortalized it and that Margu song is stuck in all of our heads now and it's like funny and light and there's levity that you've brought to that moment and the other is that you got just sucked into I need there's this song that's in my head and I need to get it out and I'm going to keep tweaking it and listening and having ideas and tweaking it until I get the song that I want and um Those are very different use cases, but I think ultimately there's so much in between these two things that it's jus…

AI assessment note: “more than half of the usage is, uh, that expert mode.”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q is that, uh, like, what did, what did Bark lean off of? Like, uh, because obviously I think there was a lot of preceding TTS work that was in open source. Um, how much of that did you use? How much of that was, like, sort of brand new from, from your research? Um, what's the intellectual lineage, uh, there, just, just to cover out the, the speech recognition style?

A So it's not speech recognition. It's, it's text to speech. But, um, as far as I know, um, there was no other, uh, certainly not in the open source, uh, text to speech that was kind of transformer based. Everything else was what I would call the old style of doing things where you build these kind of single purpose models that are really good at this one narrow task. And you're kind of always data limited and the availability of high quality training data for text to speech, um, is limited. Um, I don't think we're necessarily all that inventive to say we're going to try to train in a self supervised way, a transformer based model that on kind of lots of audio, uh, and then kind of tweak it so that we can do text to speech based on that. That would be kind of the new way of doing things in a foundation model is the, is the buzzword, if you will. And so, you know, we built that up, I think from scratch, a lot of shout outs have to go to lots of different things, whether it's Uh, papers, but also, uh, it's very obvious. Uh, there's a big shout out to, um, Andre Karpathy's nano GPT. Um, you know, there's a lot of code borrowed from there. Um, I, I think I, we are huge fans of that project. It's just to show people how you don't have to be afraid of GPT type things. And it's like, um, yeah, it's actually not all that much code to make performant transformer based models. And, you kno…

AI assessment note: “a big shout out to, um, Andre Karpathy's nano GPT”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q generation? I guess that's also a funny word. It's like, what's professional music? It's like, it's all music if it's good. It becomes professional if it's good, right? But, uh, curious to see, to hear how you're thinking about Suno, too. Like, is there a, a second act of Suno that is, like, going broader into, like, the custom mode and making, making this the central hub for music generation?

A I think we intend to make, uh, Many more modes of interaction with our stuff, but we are very much not focused on quote unquote professionals right now. Um, and it's because what we're trying to do is change how most people interact with music and not necessarily make professionals a little bit better, a little bit faster. Um, it's not that that there's anything wrong with that. It's just like not what we're focused on. And I think when we think about what workflows does the average person want to use to make music, I don't think they're very similar to the way Professional musicians make music now. Like, if you pick a random person on the street and you play them a song, and then you say, like, what did you want to change about that? They're not going to say, like, you need to split out the snare drum and make it drier. Like, that's just not something that, uh, that, uh, Random person off the street is going to say. They're going to give a lot more descriptive things about the thing about the, the kind of the oeuvre of the song, like something more general. And so I don't think we know what all of the workflows are that people are going to want to use. We're just like fairly certain that the workflows that have been developed with the current set of technologies that professionals use to make beautiful music are probably not, um, what the average person wants to use. That said…

AI assessment note: “we intend to make, uh, Many more modes of interaction with our stuff”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q things. I think before we wrap, uh, you have written a blog post that can show about, uh, good hearts, law, impact NML, which is, you know, when you measure something, then, uh, the, the thing that you measure is not a good metric anymore because people optimize for it. Any thoughts on how that applies to like LLMs and benchmarks and kind of the world we're going in today?

A Yeah. I mean, I think it's maybe even more apropos than, than when I originally wrote that because, um, so much, we see so much noise about, uh, pick your favorite benchmark and this model does slightly better than that model. And then at the end of the day, actually, there is no real world difference between these things. And it is really difficult to define what real world means. And, and, um, I think to a certain extent, it's good to have these objective benchmarks. It's good to have quantitative metrics, but at the end of the day, You need some acknowledgement that you're not going to be able to capture everything. And so, um, at least at Suno, to the extent that we have corporate values, if we don't, we don't have corporate, we're too small to have corporate values written down, but something that we say a lot is aesthetics matter, that the kind of quantitative benchmarks are never going to be the be all and end all of everything that you care about. And, um, as flawed as these, uh, uh, benchmarks are in text, they're way worse in audio. And so, um, Aesthetics Matter basically is a statement that like at the end of the day, what we are trying to do is bring music to people that makes them feel a certain way. And effectively, the only good judge of that is your ears. And so you have to listen to it. Um, and it is It is a good idea to try to make better objective benchmarks,…

AI assessment note: “it's maybe even more apropos than, than when I originally wrote that”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q predict the next word. And you take a diffusion model and you basically like add noise to an image and then kind of remove the noise. But I think for music, it's hard for people to have a mental model. Like what's the, how do you turn a music model on? Like what does a music model do to generate a song? So maybe we can, we can start there.

A Yeah. Maybe I'll even take one more step back and say, um, it's not even entirely worked out. I think the same way it is in text. And so, um, It's an evolving field. If you take a giant step back, I think audio has been, uh, lagging Uh, images and text for a while. So I think very roughly you can think audio is like one to two years behind images and text. And so, um, you kind of have to think today like text was in 20, 22 or something like this. And, um, you know, the transformer was invented. It looks like it works, but it's, it's, it's far, far less established. And so, um, you know, I'll, I'll give you the way we think about the world now, but just with a big caveat that, that I'm probably wrong if we look back. Um, in a couple years from now. Um, and I think the biggest thing is you see both transformer based and diffusion based models for audio, um, in, and in ways that that is not true in text. I know people will do some diffusion for text, but I think nobody's like really doing that, uh, for real. Uh, and, uh, so we, we prefer transformers for a variety of reasons. And so you can think it's very similar to text. You have some abstract notion of a token and you train a model to Predict the probability over all of the next tokens. So it's a language model. You can think in anything language model is just something that assigns likelihoods to sequences of tokens. Sometimes…

AI assessment note: “You have some abstract notion of a token and you train a model to Predict”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q you did an announcement, I think a couple months ago. I, I, I think I saw you like, 300 times on my Twitter timeline on like the, the same day, so it was like going everywhere. Uh, what, what's coming up? What are you most excited about in this space and maybe What are some of the most interesting underexplored, um, ideas that you maybe haven't, haven't worked on yet?

A Gosh, there's, there's a lot. You know, I think, um, from the model side, um, it's still really early innings, and there's still so much low hanging fruit for us to pick to make these models much, much better, much, much more controllable, much better music, much better audio fidelity. Um, so much that we know about and so much that, um, Again, we can kind of borrow from the open source transformers community that should make these, um, just better across the board. From the product side and the, you know, we're super focused on the experiences that we can bring to people. And so, um, it's so much more than just, uh, text to music. And I think, um, you know, I'll, I'll, I'll say this nicely. I'm a machine learning person, but like machine learning people are stupid sometimes, and we can only think about like models that take X and make it into Y. And that's just not how the average human being thinks about interacting with music. And so I think what we're most excited about is All of the new ways that we can get people just much more actively participating in music, and that is making music, not only with text, maybe with other ways of, of doing stuff. That is making music together. If you want to be reductive and think about this as a video game, this is multiplayer mode, and it is the most fun that you can have with music. And, um, you know, honestly, I think, uh, there's a l…

AI assessment note: “I think what we're most excited about is All of the new ways”

Redirected raw tape D 2 · C 4 · P 2 · Cm 3 2.75

Q And what's the data recipe for turning a good music model? Like what percentage genre do you put? Like also do you split vocals and instrumentals?

A So you have to do lots of things. And I think this is the Biggest area where we have, you know, sort of our secret sauce, I think, um, To, to a large extent, what we do is we benefit from all of the beautiful things people do with transformers and text, and we focus very hard basically on how do I tokenize audio in the right way. And, uh, without divulging too much secret sauce, um, it's a, it's at least similar to how it's done in sort of the open source stuff. You will have different models that learn to encode audio in, in discrete representations. And a lot of this boils down to figuring out the right Uh, let's say implicit biases to put in those models, the right data to inject. How do I make sure that I can produce kind of all audio arbitrarily? That's, that's speech. That's background music. That's vocals. That's kind of everything to make sure that I can really capture all the behavior that I want to.

AI assessment note: “without divulging too much secret sauce, um, it's a, it's at least similar”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.