Assertion certainty 3/5 debate potential 2/5

Murphy: Azure OpenAI significantly beats OpenAI's hosted API 400-600ms latency.

Damien Murphy · Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI · Apr 6, 2024 · at 10:17

Deepgram engineer Damien Murphy breaks down latency bottlenecks across different LLM hosting providers when building voice agents.

0:00 / 0:12exact quote · 12.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“And then GPD, 3.5 turbo or four, you probably get, you know, 400, maybe 600 milliseconds of latency in their hosted API. And if you go into Azure and you use their services, you can get that down a lot lower.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Damien Murphy

Assertion Supported
Murphy: Five-minute voice calls cost 6.5 cents on Deepgram versus ElevenLabs.
“And then on the text-to-speech side, and doing something like this with an 11 labs would be about maybe a dollar 20. And just to give you an idea of comparison. So you can do a five minute call here for about six and a half cents.”
Damien Murphy Apr 6, 2024 ▶ 16:52 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Insight
Murphy: Voice bot latency over 1.5 seconds triggers users to repeat themselves.
“So essentially, if you go beyond, say, 1.52 seconds, a lot of people will actually say something again, right? They think that the person is no longer there on the other end.”
Damien Murphy Apr 6, 2024 ▶ 8:25 Personal AI Meetup - Bee, BasedHardware, LangChain LangFriend, Deepgram EmilyAI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.