Legacy annotation firms failed at audio emotion labeling, forcing in-house creation
Mati Staniszewski · No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski · Dec 11, 2025 · at 17:15
ElevenLabs CEO Mati Staniszewski discusses the challenges of dataset creation and speech emotion labeling with Sarah Guo on No Priors.
“The understanding of how you describe audio data is still lagging in the industry. Like, when we initially started, we, of course, went into the traditional players for them to help us label not only what was said, so, like, transcription, but also how it was said. Like, what are the emotions used, accent. And most people just weren't able to do that work effectively because you kind of need to hear and have, like, a little bit of a skill set of, like, how would I describe this specific delivery? So we needed to create data ourselves.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →