Jun 20, 2024 · 27m · no-priors

No Priors Ep. 69 | With HeyGen CEO and Co-Founder Joshua Xu

Joshua Xu · 18m spoken Sarah Guo · 3m spoken Elad Gil · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

HeyGen co-founder and CEO Joshua Xu joins hosts Sarah Guo and Elad Gil on No Priors to discuss how AI avatar technology, modular generative pipelines, and real-time streaming video are replacing traditional camera shoots to make personalized video production universally scalable.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.8% of the talking time here. How this is scored →

The hosts as informed peer 3.6 Guest teaching 4.0 Guest disagreement 0.1 The hosts pushing back 0.1
05100:0010:0020:000:54–3:07 · The hosts as informed peer 3/10 HeyGen Origins and Replacing the Camera Sarah asks Joshua about HeyGen's founding story and what motivated replacing the traditional physical camera. Joshua explains his background in CMU robotics and Snapchat's AI camera team, emphasizing the goal to lower content creation barriers.3:07–5:55 · The hosts as informed peer 3/10 Deconstructing Video Production into Avatars and Editing Elad inquires why HeyGen started with virtual avatars over other video components. Joshua breaks down the video production bottleneck, noting that camera crews and scheduling are significantly more expensive than post-production editing.5:56–8:50 · The hosts as informed peer 3/10 Core HeyGen Use Cases: Create, Localize, and Personalize Joshua categorizes HeyGen's core offerings into creation, localization, and personalization, highlighting an enterprise campaign with McDonald's. He outlines his quality framework, emphasizing that generation quality must surpass a strict threshold to replace real cameras.8:50–11:03 · The hosts as informed peer 3/10 Full-Body Generation and Motion Dynamics Roadmap Sarah asks about upcoming product capabilities and customer demand for full-body generation. Joshua explains the spectrum of needs ranging from static educational training to dynamic marketing and ad creatives.11:03–14:49 · The hosts as informed peer 4/10 HeyGen Tech Stack and Multimodal Synchronization Elad and Sarah ask about the internal tech stack and how HeyGen's approach compares to end-to-end models like OpenAI's Sora. Joshua details why enterprise video demands modular component orchestration for brand consistency rather than monolithic generation.14:50–17:53 · The hosts as informed peer 3/10 Designing Around Model Limitations and Research Strategy Sarah prompts Joshua on research methodology and raises the critical issue of deepfakes and misuse. Joshua describes designing around model limitations, such as lip-sync composition, and details strict platform safeguards like dynamic verbal passcodes.17:53–20:07 · The hosts as informed peer 5/10 Transforming Business Communication via Scaled Personalization Elad introduces the positive framing of hyper-personalized video in political campaigns, asking how generative video shifts business communication. Joshua agrees, noting that AI video unlocks entirely new customer capabilities rather than just cost savings.20:08–24:00 · The hosts as informed peer 5/10 Generative Video as a Dynamic New Media Format Joshua posits that generative video represents a completely new, real-time dynamic format rather than static MP4 files. Sarah builds upon this thesis by citing educational research on personalized tutoring efficacy.24:01–26:11 · The hosts as informed peer 4/10 Evaluating Visual Aesthetics and Snapchat Lessons Elad asks about video model research challenges and Joshua's learnings from Snapchat. Joshua explains that optimizing mathematical loss functions fails to capture visual aesthetics, requiring in-product AB testing similar to mobile camera tuning.26:12–27:00 · The hosts as informed peer 3/10 HeyGen Traction, Customer Scale, and Hiring Sarah and Elad ask about HeyGen's team scale and hiring needs. Joshua shares that a 40-person team serves over 40,000 paying customers across diverse mainstream industries, and Elad highlights the strong efficiency ratio.0:54–3:07 · Guest teaching 3/10 HeyGen Origins and Replacing the Camera Sarah asks Joshua about HeyGen's founding story and what motivated replacing the traditional physical camera. Joshua explains his background in CMU robotics and Snapchat's AI camera team, emphasizing the goal to lower content creation barriers.3:07–5:55 · Guest teaching 4/10 Deconstructing Video Production into Avatars and Editing Elad inquires why HeyGen started with virtual avatars over other video components. Joshua breaks down the video production bottleneck, noting that camera crews and scheduling are significantly more expensive than post-production editing.5:56–8:50 · Guest teaching 4/10 Core HeyGen Use Cases: Create, Localize, and Personalize Joshua categorizes HeyGen's core offerings into creation, localization, and personalization, highlighting an enterprise campaign with McDonald's. He outlines his quality framework, emphasizing that generation quality must surpass a strict threshold to replace real cameras.8:50–11:03 · Guest teaching 4/10 Full-Body Generation and Motion Dynamics Roadmap Sarah asks about upcoming product capabilities and customer demand for full-body generation. Joshua explains the spectrum of needs ranging from static educational training to dynamic marketing and ad creatives.11:03–14:49 · Guest teaching 5/10 HeyGen Tech Stack and Multimodal Synchronization Elad and Sarah ask about the internal tech stack and how HeyGen's approach compares to end-to-end models like OpenAI's Sora. Joshua details why enterprise video demands modular component orchestration for brand consistency rather than monolithic generation.14:50–17:53 · Guest teaching 4/10 Designing Around Model Limitations and Research Strategy Sarah prompts Joshua on research methodology and raises the critical issue of deepfakes and misuse. Joshua describes designing around model limitations, such as lip-sync composition, and details strict platform safeguards like dynamic verbal passcodes.17:53–20:07 · Guest teaching 3/10 Transforming Business Communication via Scaled Personalization Elad introduces the positive framing of hyper-personalized video in political campaigns, asking how generative video shifts business communication. Joshua agrees, noting that AI video unlocks entirely new customer capabilities rather than just cost savings.20:08–24:00 · Guest teaching 5/10 Generative Video as a Dynamic New Media Format Joshua posits that generative video represents a completely new, real-time dynamic format rather than static MP4 files. Sarah builds upon this thesis by citing educational research on personalized tutoring efficacy.24:01–26:11 · Guest teaching 5/10 Evaluating Visual Aesthetics and Snapchat Lessons Elad asks about video model research challenges and Joshua's learnings from Snapchat. Joshua explains that optimizing mathematical loss functions fails to capture visual aesthetics, requiring in-product AB testing similar to mobile camera tuning.26:12–27:00 · Guest teaching 3/10 HeyGen Traction, Customer Scale, and Hiring Sarah and Elad ask about HeyGen's team scale and hiring needs. Joshua shares that a 40-person team serves over 40,000 paying customers across diverse mainstream industries, and Elad highlights the strong efficiency ratio.0:54–3:07 · Guest disagreement 0/10 HeyGen Origins and Replacing the Camera Sarah asks Joshua about HeyGen's founding story and what motivated replacing the traditional physical camera. Joshua explains his background in CMU robotics and Snapchat's AI camera team, emphasizing the goal to lower content creation barriers.3:07–5:55 · Guest disagreement 0/10 Deconstructing Video Production into Avatars and Editing Elad inquires why HeyGen started with virtual avatars over other video components. Joshua breaks down the video production bottleneck, noting that camera crews and scheduling are significantly more expensive than post-production editing.5:56–8:50 · Guest disagreement 0/10 Core HeyGen Use Cases: Create, Localize, and Personalize Joshua categorizes HeyGen's core offerings into creation, localization, and personalization, highlighting an enterprise campaign with McDonald's. He outlines his quality framework, emphasizing that generation quality must surpass a strict threshold to replace real cameras.8:50–11:03 · Guest disagreement 0/10 Full-Body Generation and Motion Dynamics Roadmap Sarah asks about upcoming product capabilities and customer demand for full-body generation. Joshua explains the spectrum of needs ranging from static educational training to dynamic marketing and ad creatives.11:03–14:49 · Guest disagreement 1/10 HeyGen Tech Stack and Multimodal Synchronization Elad and Sarah ask about the internal tech stack and how HeyGen's approach compares to end-to-end models like OpenAI's Sora. Joshua details why enterprise video demands modular component orchestration for brand consistency rather than monolithic generation.14:50–17:53 · Guest disagreement 0/10 Designing Around Model Limitations and Research Strategy Sarah prompts Joshua on research methodology and raises the critical issue of deepfakes and misuse. Joshua describes designing around model limitations, such as lip-sync composition, and details strict platform safeguards like dynamic verbal passcodes.17:53–20:07 · Guest disagreement 0/10 Transforming Business Communication via Scaled Personalization Elad introduces the positive framing of hyper-personalized video in political campaigns, asking how generative video shifts business communication. Joshua agrees, noting that AI video unlocks entirely new customer capabilities rather than just cost savings.20:08–24:00 · Guest disagreement 0/10 Generative Video as a Dynamic New Media Format Joshua posits that generative video represents a completely new, real-time dynamic format rather than static MP4 files. Sarah builds upon this thesis by citing educational research on personalized tutoring efficacy.24:01–26:11 · Guest disagreement 0/10 Evaluating Visual Aesthetics and Snapchat Lessons Elad asks about video model research challenges and Joshua's learnings from Snapchat. Joshua explains that optimizing mathematical loss functions fails to capture visual aesthetics, requiring in-product AB testing similar to mobile camera tuning.26:12–27:00 · Guest disagreement 0/10 HeyGen Traction, Customer Scale, and Hiring Sarah and Elad ask about HeyGen's team scale and hiring needs. Joshua shares that a 40-person team serves over 40,000 paying customers across diverse mainstream industries, and Elad highlights the strong efficiency ratio.0:54–3:07 · The hosts pushing back 0/10 HeyGen Origins and Replacing the Camera Sarah asks Joshua about HeyGen's founding story and what motivated replacing the traditional physical camera. Joshua explains his background in CMU robotics and Snapchat's AI camera team, emphasizing the goal to lower content creation barriers.3:07–5:55 · The hosts pushing back 0/10 Deconstructing Video Production into Avatars and Editing Elad inquires why HeyGen started with virtual avatars over other video components. Joshua breaks down the video production bottleneck, noting that camera crews and scheduling are significantly more expensive than post-production editing.5:56–8:50 · The hosts pushing back 0/10 Core HeyGen Use Cases: Create, Localize, and Personalize Joshua categorizes HeyGen's core offerings into creation, localization, and personalization, highlighting an enterprise campaign with McDonald's. He outlines his quality framework, emphasizing that generation quality must surpass a strict threshold to replace real cameras.8:50–11:03 · The hosts pushing back 0/10 Full-Body Generation and Motion Dynamics Roadmap Sarah asks about upcoming product capabilities and customer demand for full-body generation. Joshua explains the spectrum of needs ranging from static educational training to dynamic marketing and ad creatives.11:03–14:49 · The hosts pushing back 0/10 HeyGen Tech Stack and Multimodal Synchronization Elad and Sarah ask about the internal tech stack and how HeyGen's approach compares to end-to-end models like OpenAI's Sora. Joshua details why enterprise video demands modular component orchestration for brand consistency rather than monolithic generation.14:50–17:53 · The hosts pushing back 1/10 Designing Around Model Limitations and Research Strategy Sarah prompts Joshua on research methodology and raises the critical issue of deepfakes and misuse. Joshua describes designing around model limitations, such as lip-sync composition, and details strict platform safeguards like dynamic verbal passcodes.17:53–20:07 · The hosts pushing back 0/10 Transforming Business Communication via Scaled Personalization Elad introduces the positive framing of hyper-personalized video in political campaigns, asking how generative video shifts business communication. Joshua agrees, noting that AI video unlocks entirely new customer capabilities rather than just cost savings.20:08–24:00 · The hosts pushing back 0/10 Generative Video as a Dynamic New Media Format Joshua posits that generative video represents a completely new, real-time dynamic format rather than static MP4 files. Sarah builds upon this thesis by citing educational research on personalized tutoring efficacy.24:01–26:11 · The hosts pushing back 0/10 Evaluating Visual Aesthetics and Snapchat Lessons Elad asks about video model research challenges and Joshua's learnings from Snapchat. Joshua explains that optimizing mathematical loss functions fails to capture visual aesthetics, requiring in-product AB testing similar to mobile camera tuning.26:12–27:00 · The hosts pushing back 0/10 HeyGen Traction, Customer Scale, and Hiring Sarah and Elad ask about HeyGen's team scale and hiring needs. Joshua shares that a 40-person team serves over 40,000 paying customers across diverse mainstream industries, and Elad highlights the strong efficiency ratio.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 31.2% · guest 68.8%0:00 · the hosts 31.2% · guest 68.8%3:00 · the hosts 29.7% · guest 70.3%3:00 · the hosts 29.7% · guest 70.3%6:00 · the hosts 17.6% · guest 82.4%6:00 · the hosts 17.6% · guest 82.4%9:00 · the hosts 23% · guest 77%9:00 · the hosts 23% · guest 77%12:00 · the hosts 12.7% · guest 87.3%12:00 · the hosts 12.7% · guest 87.3%15:00 · the hosts 22.8% · guest 77.2%15:00 · the hosts 22.8% · guest 77.2%18:00 · the hosts 35% · guest 65%18:00 · the hosts 35% · guest 65%21:00 · the hosts 19.4% · guest 80.6%21:00 · the hosts 19.4% · guest 80.6%24:00 · the hosts 15.3% · guest 84.7%24:00 · the hosts 15.3% · guest 84.7%27:00 · the hosts 89.5% · guest 10.5%27:00 · the hosts 89.5% · guest 10.5%
Sharpest disagreement ▶ 13:02 Rejecting end-to-end generation in favor of modular assembly

Joshua politely rejects the premise that end-to-end models like Sora are the right solution for enterprise video, arguing that brand consistency requires modular A-roll and B-roll orchestration.

Hardest push from the hosts ▶ 16:36 Sarah confronts the risks of deepfakes and misuse

Sarah directly questions Joshua on the concerning safety risks and deepfake abuses associated with likeness and voice replication.

Biggest teaching moment ▶ 21:15 Generative video is not a static MP4

Joshua reframes how to conceptualize generative video, explaining to the hosts that it is not an immutable MP4 file but a real-time, dynamic media stream tailored to individual viewer attributes.

The host holds their own ▶ 22:45 Sarah connects dynamic video to personalized learning studies

Sarah demonstrates domain expertise by connecting Joshua's thesis on dynamic generative video directly to Bloom's educational research on the superiority of personalized tutoring.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
HeyGen Origins and Replacing the Camera 3300 Sarah asks Joshua about HeyGen's founding story and what motivated replacing the traditional physical camera. Joshua explains his background in CMU robotics and Snapchat's AI camera team, emphasizing the goal to lower content creation barriers.
Deconstructing Video Production into Avatars and Editing 3400 Elad inquires why HeyGen started with virtual avatars over other video components. Joshua breaks down the video production bottleneck, noting that camera crews and scheduling are significantly more expensive than post-production editing.
Core HeyGen Use Cases: Create, Localize, and Personalize 3400 Joshua categorizes HeyGen's core offerings into creation, localization, and personalization, highlighting an enterprise campaign with McDonald's. He outlines his quality framework, emphasizing that generation quality must surpass a strict threshold to replace real cameras.
Full-Body Generation and Motion Dynamics Roadmap 3400 Sarah asks about upcoming product capabilities and customer demand for full-body generation. Joshua explains the spectrum of needs ranging from static educational training to dynamic marketing and ad creatives.
HeyGen Tech Stack and Multimodal Synchronization 4510 Elad and Sarah ask about the internal tech stack and how HeyGen's approach compares to end-to-end models like OpenAI's Sora. Joshua details why enterprise video demands modular component orchestration for brand consistency rather than monolithic generation.
Designing Around Model Limitations and Research Strategy 3401 Sarah prompts Joshua on research methodology and raises the critical issue of deepfakes and misuse. Joshua describes designing around model limitations, such as lip-sync composition, and details strict platform safeguards like dynamic verbal passcodes.
Transforming Business Communication via Scaled Personalization 5300 Elad introduces the positive framing of hyper-personalized video in political campaigns, asking how generative video shifts business communication. Joshua agrees, noting that AI video unlocks entirely new customer capabilities rather than just cost savings.
Generative Video as a Dynamic New Media Format 5500 Joshua posits that generative video represents a completely new, real-time dynamic format rather than static MP4 files. Sarah builds upon this thesis by citing educational research on personalized tutoring efficacy.
Evaluating Visual Aesthetics and Snapchat Lessons 4500 Elad asks about video model research challenges and Joshua's learnings from Snapchat. Joshua explains that optimizing mathematical loss functions fails to capture visual aesthetics, requiring in-product AB testing similar to mobile camera tuning.
HeyGen Traction, Customer Scale, and Hiring 3300 Sarah and Elad ask about HeyGen's team scale and hiring needs. Joshua shares that a 40-person team serves over 40,000 paying customers across diverse mainstream industries, and Elad highlights the strong efficiency ratio.

Statements from this episode (19)

Opinion
AI will replace physical cameras for visual content creation
“We wanted to replace the camera because we think AI can create the content, and AI could become the new camera, and that's how we get started with H&M, and our mission is to making visual storytelling accessible to all.”
Joshua Xu Jun 20, 2024 ▶ 2:01
Insight
Video editing is inexpensive and standardized, but camera production is costly
“Editing, we just learned from customer that editing is not that expensive because it's pretty standard service, but camera is super expensive.”
Joshua Xu Jun 20, 2024 ▶ 3:53
Prediction Not checkable as stated
Streaming generative video could replace real-time human conversations
“If we push forward into technology, making the performance much better, I think we will be able to create experience like generative video in a streaming way. And that actually will be potentially replace a lot of the, you know, real-time conversation we have …”
Joshua Xu Jun 20, 2024 ▶ 5:31
Assertion Supported
HeyGen localizes video into over 175 languages and dialects
“We can also Take existing video that localized that into a hundred, more than a 175 different languages and dialysis.”
Joshua Xu Jun 20, 2024 ▶ 6:34
Insight
Generative video below 90% quality threshold is unusable for enterprises
“There's an invisible line of a quality, you know, let's say that thresholds to 90 anything below 90 essentially is unusable for the customers because we cannot really replace the real life production process they have.”
Joshua Xu Jun 20, 2024 ▶ 8:03
Prediction Not checkable as stated
HeyGen streaming avatars could become the visualization layer for real-time AI
“The streaming avatar especially with the latest release on GPD for all, really, really helped to improve the performance of the real-time interaction with text and voice, and HN avatar could become a visualization layer for all those applications.”
Joshua Xu Jun 20, 2024 ▶ 9:22
Disclosure
HeyGen uses OpenAI and ElevenLabs alongside proprietary video models
“We work with OpenAI, ChatGPT on the text generation side. Obviously also serves like the brain of the orchestration engine that we build internally. And we work with you know, OpenAI and Event Lab on the voice engine, but we build the entire video stack in-hou…”
Joshua Xu Jun 20, 2024 ▶ 11:34
Prediction Not checkable as stated
Multimodal AI models will converge into unified single architectures
“So I think over time, I think the whole technology trend has been moving towards to a direction. A lot of all these things will be trained together. The multi-model model, multimedia, all get into one single model.”
Joshua Xu Jun 20, 2024 ▶ 11:57
Opinion
Modular AI video pipelines beat end-to-end models like Sora for enterprises
“And the other approach is that what we believe in at Hadrian is that we try to assemble the whole video into different components. Lastly, it will be A-Roll and B-Roll. B-Roll represents all different kinds of elements, like voiceover, music, transition, A-Rol…”
Joshua Xu Jun 20, 2024 ▶ 13:36
Disclosure
HeyGen views OpenAI's Sora as a component, not a direct competitor
“And in fact, we actually see Sora as our partner because we can, we are able to integrate that as one of the component, you know, generator, and then feed that into our acquisition engine for the business application.”
Joshua Xu Jun 20, 2024 ▶ 14:36
Disclosure
HeyGen video translation pipeline combines lip-sync, voice cloning, and ChatGPT
“One example would be when we look at, you know, the video translation technology is you know, it's a whole new way to translate a content compared to traditional dubbing. We preserve the user the natural voice and their facial expression. But if you look at re…”
Joshua Xu Jun 20, 2024 ▶ 15:57
Disclosure
HeyGen strictly prohibits all political and election content on its platform
“First of all, we do not allow any political or election content on our platform today.”
Joshua Xu Jun 20, 2024 ▶ 16:57
Disclosure
HeyGen uses live video consent and dynamic passcodes to secure avatars
“So we have our safety, you know, security safeguard include very advanced user verification, include, you know, live video consent, dynamic verbal passcode, and rapid human review in the back of all the other have been created on the platform.”
Joshua Xu Jun 20, 2024 ▶ 17:11
Prediction Not checkable as stated
Avatar generation becomes real-time in two years, full video in five
“Two years from now, it would not be crazy to look at a lot of avatar generation. Asynchronization pipeline will become real-time streaming capable. And I also see the world is moving towards a way that we can probably generate the entire video in real-time as …”
Joshua Xu Jun 20, 2024 ▶ 20:58
Insight
Generative video is a fundamentally new dynamic format, not just video
“I have an opinion like, you know, generative image is still image. But generative video is not a video. It is a new format.”
Joshua Xu Jun 20, 2024 ▶ 21:19
Assertion Supported
Publicis generated over 100,000 personalized HeyGen videos for global employees
“And one of the use case we have seen from customers that, you know, publicist group, they generate more than a 100,000 videos, a thank you video to send to all the employee globally and in localized into different languages personalized with a name and they ar…”
Joshua Xu Jun 20, 2024 ▶ 23:21
Insight
Lower AI loss functions do not guarantee better generative video quality
“Unlike a lot of other model, I think building video model you know, being able to integrate aesthetics into the AI model is pretty hard. So, you know, video generation is not only about solving a mathematical problem. It's actually about creating something the…”
Joshua Xu Jun 20, 2024 ▶ 24:10
Disclosure
HeyGen trains and evaluates video models using live A/B test data
“We have to rely on in-product signal, for example, AB test to know which model is actually better, because, you know, only the customer can be the judge for that. And this process generally is just not different from a mathematical standpoint. We kind of have …”
Joshua Xu Jun 20, 2024 ▶ 24:49
Assertion Not checkable as stated
HeyGen has just 40 employees serving over 40,000 paying customers
“We are a little bit over 40 people, but we are serving over 40, 40,000 paying customers on the platform today.”
Joshua Xu Jun 20, 2024 ▶ 26:15
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.