Doug Eck, lead research scientist at Google Magenta, discusses the technical capabilities and performance bottlenecks of the NSynth generative audio model.
Opinion
Eck: Alex Graves advanced LSTMs more than anyone, including its creator
“Among the three of us, by far, Alex Graves has done the most with LSTM. So he continued, after he finished his PhD, and he continued doggedly to try to understand how recurrent neural networks worked, how to train them, and how to make them useful for sequence…”
Assertion Supported
Eck: Untrained listeners rated Magenta compositions as more Bach-like than Bach
“When we put these tunes out for, like, untrained listeners to listen to, they sometimes voted them as sounding more Bach-y”
Assertion Supported
Eck: As of 2017, AI cannot generate a single coherent text paragraph
“So, so everybody understands that's listening or watching, you know, we can, we can't generate a coherent paragraph, right? So, I don't mean we, magenta. I mean, kind of we, humanity.”
Disclosure
Eck: Magenta aims to build creative tools, not act as artists
“The reason that I put that quote there, I think, is to be honest with the division between engineering and research and artistry, and to not think that what I'm doing is being a machine learning artist, but we're trying to build interesting ways to make new ki…”
Insight
Eck: Users' first instinct with AI art models is breaking them
“The first thing you're gonna do, if you think, if someone comes to you and says, here's this really smart model that you can make art with, what are you gonna do? You're gonna try to show the world that it's a stupid model, right? But maybe the way that, maybe…”
Opinion
Eck: Magenta's generated AI music is not yet good enough for testing
“I, at least for Magenta, I haven't felt like the quality of what we've been generating has been good enough to bother, so to speak. Like you find it, you cherry pick, you find some good things, you're like, okay, this model trains, and it's interesting, and no…”