Jun 11, 2025 · 14m · a16z

The Ultimate AI Video Stack: Up-to-Date Best Tools to Make Content With AI

Justine Moore · 11m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this comprehensive AI content creator guide, venture capitalist and creator Justine Moore breaks down her ultimate video stack using cutting-edge tools like Google Veo 3, Kling AI, Hedra, Higgsfield AI, and Krea AI. She demonstrates practical workflows for text-to-video, image animation, character voice sync, visual effects, and model benchmarking to help creators produce high-quality AI media.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 0.0 Guest teaching 0.0 Guest disagreement 0.0 The host pushing back 0.0
05100:0010:000:46–4:04 · The host as informed peer 0/10 Text-to-Video Generation with Veo 3 on Google Labs This segment is a solo technical tutorial where Justine Moore demonstrates Google Labs' Veo 3 text-to-video capabilities. Because there is no active host interaction, all host and combativeness scores remain at zero.4:04–6:56 · The host as informed peer 0/10 Image-to-Video Animation and Sound FX in Kling AI Justine conducts a solo demonstration of Kling AI's image-to-video animation features and native sound effects generation without any host presence or dialogue.6:56–9:39 · The host as informed peer 0/10 Talking Character Creation and Voice Sync with Hedra Justine walks through Hedra for character animation and voice syncing using cloned audio. The host is absent, leaving the power dynamic completely neutral in monologue format.9:39–13:53 · The host as informed peer 0/10 Visual Effects and Presets with Higgsfield AI Justine demonstrates Higgsfield AI VFX presets and multi-model upscaling on Crea AI. Without host participation, no pushback or combativeness occurs.0:46–4:04 · Guest teaching 0/10 Text-to-Video Generation with Veo 3 on Google Labs This segment is a solo technical tutorial where Justine Moore demonstrates Google Labs' Veo 3 text-to-video capabilities. Because there is no active host interaction, all host and combativeness scores remain at zero.4:04–6:56 · Guest teaching 0/10 Image-to-Video Animation and Sound FX in Kling AI Justine conducts a solo demonstration of Kling AI's image-to-video animation features and native sound effects generation without any host presence or dialogue.6:56–9:39 · Guest teaching 0/10 Talking Character Creation and Voice Sync with Hedra Justine walks through Hedra for character animation and voice syncing using cloned audio. The host is absent, leaving the power dynamic completely neutral in monologue format.9:39–13:53 · Guest teaching 0/10 Visual Effects and Presets with Higgsfield AI Justine demonstrates Higgsfield AI VFX presets and multi-model upscaling on Crea AI. Without host participation, no pushback or combativeness occurs.0:46–4:04 · Guest disagreement 0/10 Text-to-Video Generation with Veo 3 on Google Labs This segment is a solo technical tutorial where Justine Moore demonstrates Google Labs' Veo 3 text-to-video capabilities. Because there is no active host interaction, all host and combativeness scores remain at zero.4:04–6:56 · Guest disagreement 0/10 Image-to-Video Animation and Sound FX in Kling AI Justine conducts a solo demonstration of Kling AI's image-to-video animation features and native sound effects generation without any host presence or dialogue.6:56–9:39 · Guest disagreement 0/10 Talking Character Creation and Voice Sync with Hedra Justine walks through Hedra for character animation and voice syncing using cloned audio. The host is absent, leaving the power dynamic completely neutral in monologue format.9:39–13:53 · Guest disagreement 0/10 Visual Effects and Presets with Higgsfield AI Justine demonstrates Higgsfield AI VFX presets and multi-model upscaling on Crea AI. Without host participation, no pushback or combativeness occurs.0:46–4:04 · The host pushing back 0/10 Text-to-Video Generation with Veo 3 on Google Labs This segment is a solo technical tutorial where Justine Moore demonstrates Google Labs' Veo 3 text-to-video capabilities. Because there is no active host interaction, all host and combativeness scores remain at zero.4:04–6:56 · The host pushing back 0/10 Image-to-Video Animation and Sound FX in Kling AI Justine conducts a solo demonstration of Kling AI's image-to-video animation features and native sound effects generation without any host presence or dialogue.6:56–9:39 · The host pushing back 0/10 Talking Character Creation and Voice Sync with Hedra Justine walks through Hedra for character animation and voice syncing using cloned audio. The host is absent, leaving the power dynamic completely neutral in monologue format.9:39–13:53 · The host pushing back 0/10 Visual Effects and Presets with Higgsfield AI Justine demonstrates Higgsfield AI VFX presets and multi-model upscaling on Crea AI. Without host participation, no pushback or combativeness occurs.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 2:15 Warning about tool interface model toggles

In a non-combative solo monologue, Justine playfully warns users that the interface might trick them by switching back to Veo 2.

Hardest push from the host ▶ 0:46 Absence of host pushback

The episode is entirely a solo software walkthrough by the guest, resulting in zero host pushback across the entire recording.

Biggest teaching moment ▶ 3:00 Tutorial on preventing weird AI filler words

Justine educates the audience on how providing insufficient prompt text for eight seconds of video forces the model to generate nonsensical filler words.

The host holds their own ▶ 0:46 Absence of host counter-expertise

Because the host does not engage during the walkthrough segments, there are no instances of host pushback or demonstrated host expertise.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Text-to-Video Generation with Veo 3 on Google Labs 0000 This segment is a solo technical tutorial where Justine Moore demonstrates Google Labs' Veo 3 text-to-video capabilities. Because there is no active host interaction, all host and combativeness scores remain at zero.
Image-to-Video Animation and Sound FX in Kling AI 0000 Justine conducts a solo demonstration of Kling AI's image-to-video animation features and native sound effects generation without any host presence or dialogue.
Talking Character Creation and Voice Sync with Hedra 0000 Justine walks through Hedra for character animation and voice syncing using cloned audio. The host is absent, leaving the power dynamic completely neutral in monologue format.
Visual Effects and Presets with Higgsfield AI 0000 Justine demonstrates Higgsfield AI VFX presets and multi-model upscaling on Crea AI. Without host participation, no pushback or combativeness occurs.

Statements from this episode (8)

Opinion
Moore: Veo 3 is currently the best text-to-video AI model
“So we're going to dive into my current AI video stack for consumer creators like me, starting with VO three, which I think is currently the best text to video model.”
Justine Moore Jun 11, 2025 ▶ 0:38
Insight
Insufficient dialogue prompts in Veo 3 trigger unwanted AI audio filler
“I've noticed that the model will often do something weird if you don't put in enough text to fill eight seconds of audio. For example if you make it sound like it's supposed to be a street interview, and it's an 8:02 video, and there's only two seconds of audi…”
Justine Moore Jun 11, 2025 ▶ 2:37
Opinion
Justine Moore considers Kling AI her favorite image-to-video model
“Up next is my favorite model for generating a video from an image. So this is where you're starting with a photo or any other sort of image, and you want to animate it. You want to make people move, maybe you want to make the background move, have things come …”
Justine Moore Jun 11, 2025 ▶ 4:05
Prediction Held up
Justine Moore predicts Kling AI will add multi-frame support to Kling 2.1
“Right now, Cling. 2.1 only supports a start frame, not a start and end frame, but I would imagine that they're going to add more frames soon as they've done for their other models.”
Justine Moore Jun 11, 2025 ▶ 4:36
Opinion
Justine Moore: Hedra is by far her favorite character audio-sync tool
“Alright, next we're going to be talking about how you make a character speak, and my favorite tool for this is by far, Hedra.”
Justine Moore Jun 11, 2025 ▶ 6:57
Insight
Justine Moore: Hedra character sync works best using neutral face inputs
“I've noticed that Hedra is better when you start with a neutral face from the character, like he's sort of smiling at the beginning while the speech that he's saying is not super happy, but that's just me being incredibly picky.”
Justine Moore Jun 11, 2025 ▶ 9:25
Opinion
Moore: Higgsfield AI stands out for playable community VFX presets
“Higgsfield, which is a very cool VFX platform. So what I really like about Higgsfield is you can sort of browse and see all of these really cool things that people are doing or things that they're making and then you can run them yourself.”
Justine Moore Jun 11, 2025 ▶ 9:43
Disclosure
Moore: Krea is her favorite tool for testing multi-model AI video
“Next is my favorite place to use open source models like Wan and Hunwan, or, ah, to test a bunch of different models in one place.”
Justine Moore Jun 11, 2025 ▶ 11:35
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.