Everything Joshua Xu said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Streaming generative video could replace real-time human conversations
“If we push forward into technology, making the performance much better, I think we will be able to create experience like generative video in a streaming way. And that actually will be potentially replace a lot of the, you know, real-time conversation we have …”
Modular AI video pipelines beat end-to-end models like Sora for enterprises
“And the other approach is that what we believe in at Hadrian is that we try to assemble the whole video into different components. Lastly, it will be A-Roll and B-Roll. B-Roll represents all different kinds of elements, like voiceover, music, transition, A-Rol…”
Avatar generation becomes real-time in two years, full video in five
“Two years from now, it would not be crazy to look at a lot of avatar generation. Asynchronization pipeline will become real-time streaming capable. And I also see the world is moving towards a way that we can probably generate the entire video in real-time as …”
AI will replace physical cameras for visual content creation
“We wanted to replace the camera because we think AI can create the content, and AI could become the new camera, and that's how we get started with H&M, and our mission is to making visual storytelling accessible to all.”
HeyGen views OpenAI's Sora as a component, not a direct competitor
“And in fact, we actually see Sora as our partner because we can, we are able to integrate that as one of the component, you know, generator, and then feed that into our acquisition engine for the business application.”
Video editing is inexpensive and standardized, but camera production is costly
“Editing, we just learned from customer that editing is not that expensive because it's pretty standard service, but camera is super expensive.”
Generative video below 90% quality threshold is unusable for enterprises
“There's an invisible line of a quality, you know, let's say that thresholds to 90 anything below 90 essentially is unusable for the customers because we cannot really replace the real life production process they have.”
HeyGen streaming avatars could become the visualization layer for real-time AI
“The streaming avatar especially with the latest release on GPD for all, really, really helped to improve the performance of the real-time interaction with text and voice, and HN avatar could become a visualization layer for all those applications.”
Multimodal AI models will converge into unified single architectures
“So I think over time, I think the whole technology trend has been moving towards to a direction. A lot of all these things will be trained together. The multi-model model, multimedia, all get into one single model.”
HeyGen strictly prohibits all political and election content on its platform
“First of all, we do not allow any political or election content on our platform today.”
Generative video is a fundamentally new dynamic format, not just video
“I have an opinion like, you know, generative image is still image. But generative video is not a video. It is a new format.”
Lower AI loss functions do not guarantee better generative video quality
“Unlike a lot of other model, I think building video model you know, being able to integrate aesthetics into the AI model is pretty hard. So, you know, video generation is not only about solving a mathematical problem. It's actually about creating something the…”
HeyGen has just 40 employees serving over 40,000 paying customers
“We are a little bit over 40 people, but we are serving over 40, 40,000 paying customers on the platform today.”
HeyGen localizes video into over 175 languages and dialects
“We can also Take existing video that localized that into a hundred, more than a 175 different languages and dialysis.”
HeyGen uses OpenAI and ElevenLabs alongside proprietary video models
“We work with OpenAI, ChatGPT on the text generation side. Obviously also serves like the brain of the orchestration engine that we build internally. And we work with you know, OpenAI and Event Lab on the voice engine, but we build the entire video stack in-hou…”
HeyGen video translation pipeline combines lip-sync, voice cloning, and ChatGPT
“One example would be when we look at, you know, the video translation technology is you know, it's a whole new way to translate a content compared to traditional dubbing. We preserve the user the natural voice and their facial expression. But if you look at re…”
HeyGen uses live video consent and dynamic passcodes to secure avatars
“So we have our safety, you know, security safeguard include very advanced user verification, include, you know, live video consent, dynamic verbal passcode, and rapid human review in the back of all the other have been created on the platform.”
Publicis generated over 100,000 personalized HeyGen videos for global employees
“And one of the use case we have seen from customers that, you know, publicist group, they generate more than a 100,000 videos, a thank you video to send to all the employee globally and in localized into different languages personalized with a name and they ar…”
HeyGen trains and evaluates video models using live A/B test data
“We have to rely on in-product signal, for example, AB test to know which model is actually better, because, you know, only the customer can be the judge for that. And this process generally is just not different from a mathematical standpoint. We kind of have …”