The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Mikey Shulman no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q So you guys actually started with this open source, um, model Bark. Can you talk about like what the idea was at the very beginning and how you ended up in music generation?

A We did our, we were doing all text at Kensho and we did our first audio project, um, after we were acquired by S&P Global, which was learning to transcribe earnings calls. So I'm sure both of you have read an earnings call transcript, uh, exceedingly likely it was done by S&P Global. Um, it used to be done completely manually. It was very painful and we could lend a lot of speed and scale by bringing automation to that. And we fell in love with doing audio AI. Like we happened to be musicians, but it kind of took this very honestly non-sexy project of earnings call transcription to show us how much we loved it. We also realized that certainly compared to images and text, audio was really, really far behind, and this was in 2020. And I think that's maybe even more true now, if you just look at everything that's happened in images and text in the last couple of years. Like I said, we never had a master plan. We, we made Bark and, um, as an open source project. And, um, even before we released Bark, we knew we wouldn't be focusing on speech. I think if I'm honest, a lot of people told us, go build a speech company. It is more straightforward. You'll build a Great B to B product and people will love it. And we couldn't help ourselves. We just love music too much. And so we decided to build a music company.

AI assessment note: “We just love music too much. And so we decided to build a music company.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q feel some of that, like, joy of creation with other people. Maybe you already see it in the product, but are you imagining that you get that collaboration joy from, like, or, you know, the creation joy of working with yourself, feeling like you are more skilled, you're collaborating with AI with Suno, or are people jamming? Do you see, like, mixtape, like, sharing behaviors today you can talk about?

A We see all of that, which is super cool. Like a video game, music is fun by yourself and maybe more fun in multiplayer mode. And so we see people enjoying this by themselves, but we see people basically hacking multiplayer mode, uh, into this in, in lots of fun ways where you can have people co-writing lyrics together, trading off words, trading off verses, uh, I'll write the verse, you write the chorus, or I'll write the lyrics and you pick all the styles and, uh, I'll make a song and then I'll send it to you and you'll, you know, make a song back. And so It's not surprising. I think humans really evolved to resonate strongly with music and want to do music together. Every culture basically has music. And so it really shouldn't be surprising that we see, um, all of this, but it is really fulfilling from our perspective because it really brings people together. It makes people smile. I don't pretend like we're hearing cancer at Suno, but it, it is really cool to make a lot of people smile.

AI assessment note: “We see all of that, which is super cool. Like a video game”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q That's, um, a really interesting vision of the future. I guess that has pretty deep implications as well in terms of how you think music about music, the music industry, how it permeates society. Do you have a view in terms of what all this looks like five years from now?

A If we are correct that there are just modes of experience around music that people don't have access to, that we can get a billion people much more engaged with music than they are now, that just in terms of the number of dollars or the amount of time people are spending doing music, both of those are going to go up dramatically. That I feel quite confident about. The exact nature of how this looks, um, I think is up for some more debate. So this is just an opinion. Um, I don't, Um, because, uh, music is so human and, and so much emotional connection involved in it, I don't really see people, um, losing connection with their favorite artists at all. Um, in fact, if you labor around music and you understand the process, you feel a much deeper connection, um, with the artists that you love. Um, another thing I think, uh, is likely to happen, um, if we look at like the last wave of technologies to enter music, let's say the, the DAW, Um, this really accelerates how quickly music can change and how quickly culture can change. You know, music is really just a reflection of culture. And, um, I think the way that happened is the DAW really let a lot of people start making music who could never make music. You could do this from your dorm room if you had a good pair of headphones and you had a good ear and you were willing to put in the work to learn the tool. And I think if we can giv…

AI assessment note: “just in terms of the number of dollars or the amount of time people are spending”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Mikey, what's, uh, what's hard in AI music? Like, I, I know less about, like, what this frontier looks like. Like, where do you want to push in terms of things that, um, are really hard for the model to get right? Like, you know, in visual models or video, like human hands, object permanence, like there's lots of things that are more intuitive to me there.

A Yeah, that's a really good question. I confess, I've not really thought about that too much. Um, there are the easy things, or the easy to describe things, like, You know, did you get the stereo right? Did you get the bit rate high enough, et cetera? Um, again, I think the reason music is so special is because it makes you feel a certain way. And like to the extent that any of this is difficult, it is because you are really targeting human emotions in some way. And, um, that's not terribly well understood by anyone. Um, and it is also super, super, um, diverse and super culturally dependent and super age dependent or demographic dependent. So, um, You know, I think What we're doing is so far from objective truth. Um, and it's very easy for people who spend all their days in text LLMs to be thinking about things like this is how well I did on the, on the LSAT, you know, I can pass the bar with this size model, uh, the, like the, the law bar. Um, and, um, none of that exists for us. It's really just like, I made a song and it made me feel a certain way. And it may have been grainy audio that made me feel a certain way. It may have been a long song, a short song. I think there's a lot more unanswerable questions in this domain.

AI assessment note: “to the extent that any of this is difficult, it is because you are really targeting human emotions”

Answered raw tape D 4 · C 5 · P 3 · Cm 3 3.90

Q As you said, the thing that matters is how it makes you feel. And so, like, how did you measure quality in your own models? Like, what do you know about how to train something that creates great generations? Is it just all like Mikey as human eval?

A Uh, it's definitely not all Mikey as human eval, uh, but, um, you know, one thing we say here is that aesthetics matter, and I think that is, um, a recognition that, uh, I think in, in all branches of AI, we become slaves to our metrics, and you say, I did this accuracy on this benchmark, and this accuracy on this benchmark, and in the real world, sometimes it doesn't necessarily matter, and these benchmarks are extra terrible in audio, um, Just because the field is so new. And so aesthetics matter is like a way of saying that you have to use your ears, uh, in order to evaluate things. You can look at the things like at what your final losses or something like that, but ultimately, um, it's, it's definitely more tedious to, to evaluate than, than you want it to be. I think the good news is everybody here really loves music. And so evaluating your models, which means listening to a lot of things and getting people to listen to a lot of things and doing a lot of AB tests turns out to be fun. Um, but I think we have a long way to go in this journey on how, on how we're actually going to evaluate these things, and I think we learn a lot about, um, human beings and human emotions while we learn to evaluate these things.

AI assessment note: “listening to a lot of things and getting people to listen to a lot of things”

Redirected raw tape D 2 · C 3 · P 2 · Cm 2 2.30

Q Yeah, that's cool. Are there any ways that people have started to use a product that were very unexpected for you or surprising use cases or applications or other things people have done with it?

A I think so much has been, um, Really fulfilling and cool to see and definitely surprising. And, you know, one thing I'm constantly reminding everyone is that we are eliciting a set of behaviors that are not, um, common and that are not, uh, regular for people to do. And so it's not going to be surprising when we see stuff, um, that comes out. It is maybe not surprising that people love to feel creative and they love to feel ownership over what they produce and they love to Share it with others. If you want to be a little bit more, uh, reductive about it, they love to feel famous. Um, but I think it's not the same way that, that famous people are famous. It's, it's, it's a little bit different. And so we've seen that people will spend a lot of time in front of their computers, enjoying making songs. This is really cool. And it is different from, I think the way music is done now, music is done now, sometimes painfully, but only in service of the final product. Um, and I think when you open this up to people, Um, sure you definitely care about the final product about what the song sounds like on the other end, but you also really cared about the journey and that people will really enjoy making music, um, regardless of the final product. And I can tell you, you know, um, personally, uh, the most fun I have ever had doing music is playing music with friends, jam sessions, even when…

AI assessment note: “It is maybe not surprising that people love to feel creative and they love”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.