Goel: Model Evaluation Is Undervalued; Audio AI Benchmarks Remain Inadequate
“The most undervalued part of this is evaluating your models. Because, like, I think, especially in things like audio. Yeah. There, when we started the benchmarks were pretty non-existent. Even today, I would say the benchmarks are not great.”
Kant: Poolside will not touch audio modality for a very long time
“We're, I don't think we'll touch audio for a very long time.”
Kant: Audio does not push AI models closer to AGI
“I don't think audio. Adds to that. I don't think it pushes us close to AGI. I think it is a necessary modality as you get close to AGI.”
Wu: Native speech-to-speech multimodal models will improve dramatically in 6-12 months
“I think they're gonna get a lot better at audio over the next six to 12 months, especially the likes, you know, the native multimodal models, the speech-to-speech ones.”
Wu: Enterprise audio AI is hugely underrated compared to text coding tools
“Audio, especially in the enterprise and in a business setting, I think is a hugely underrated domain still. Like, everyone talks about coding. It's all text. But we're talking in audio. A lot of the world's business is done via audio. A lot of services and ope…”
Charlamagne: The real money in podcasting is in audio, not video
“We all know that the money is in the audio.”
Reid: The web will become increasingly multimodal with more video and audio
“I think it is probably even more multimodal, multi-format. Like I would certainly expect video to continue to grow. I would certainly expect audio and not just, you know, text with some images to grow.”
Justine Moore: Veo 3 natively generates audio alongside video
“And what's very different about it is it generates audio natively at the same time it generates video.”
Shulman: Scaling laws will not solve AI music generation
“For music, It's very different from text, and I think people will very sloppily look at the world of OpenAI and Anthropic and the hyperscalers and say audio is just a couple years behind, which it is, but that scale is gonna solve all these things. But unlike …”
Audio AI lags further behind text and image models today
“We also realized that certainly compared to images and text, audio was really, really far behind, and this was in 2020. And I think that's maybe even more true now, if you just look at everything that's happened in images and text in the last couple of years.”
Shulman: Generative audio AI lags text and images by one to two years
“So I think very roughly you can think audio is like one to two years behind images and text. And so you kind of have to think today like text was in 20, 22 or something like this.”
Chintala: Smell and touch digitization is where images were in 1920
“When we think about audio, or images, or video, they're, like, so advanced that we have the concept of color spaces, we have the concept of, like, frequency spectrums, like, you know, we figured out how ears process, like frequencies in mouse spectrum, or what…”
Delangue: NLP, vision, and audio are the top three Hugging Face tasks
“The three main tasks right now are NLP, so text, right?
From like information extraction, text generation, text classification.
the second one is text to image and computer vision, right?
So object detection, text to image, text image generation.
the third o…”
Bob Pittman: Audio Is an Underappreciated and Undervalued Asset Class
“This company has an asset no one appreciates. Audio is so underappreciated, undervalued. People don't know what it can really do. I do.”
Bisu: Audio avoids competing directly with video for active screen time
“Audio is something that can solve this problem. And this time zone is not competing with other time zones that we spend on screen. Like, there are other content formats. We can do it on the video. We have to watch the video on the screen. Like, when you consum…”
Bisu: Listeners drop poor audio quickly because it engages only one sense
“And that's why we consume low quality content on YouTube. That's why we consume two sensors. We can consume low quality content. But in case of audio, you are listening with only one sensor. If there is a little bit noise, you will get irritated and you will l…”
Frasier: Major audio players struggle with Twitter engagement, making video essential
“If you look at some of the larger audio kind of Players in this space. They all suck at like audio engagement on Twitter and all these. Rarely do you see, you know, really taking off video is the best way to go.”
Audio will become a dominant content format within ten years
“So audio, our understanding is, is going to be one of the largest formats over the next five, 10 years. It's just that our belief also is that right now there hasn't been, or at least when we launched Patlipi FM or when we acquired IBM, it seemed like the dema…”
The US audio advertising market is worth roughly $18 billion
“I think that if you look at audio as a full category in the U S today, it's about eighteen billion.”
Blumberg: Audio only competes with other audio and music for attention
“You know, you're driving to work, whatever. You can't be looking at a screen. So all audio is competing with is other audio or music, basically.”
Mayo: Acquiring consumer audio users is brutal; B2B infrastructure scales better
“Acquiring customers is really hard in audio, and instead of competing with distributors and publishers who are our partners, we'd rather empower them and work with them to create a better user experience and a better distribution for them, so being behind the …”
Mayo: Audio adoption is slower than video, but retention is far stronger
“Audio is a product that people, once they adopt it, they don't leave it, but the initial adoption isn't as quick as You know, a YouTube video or something like that, but once you get someone, audio builds a really strong habit, as you probably know with your a…”