Aug 1, 2025 · 38m · a16z

The Future of AI Video: How Fal.ai is Making AI Video Faster & Easier

Burkay Gur · 14m spoken Jennifer Li · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, Fal.ai co-founders Burkay Gur and Batuhan Taskaya discuss how their company built a dominant high-performance inference platform for generative AI video and media through first-principles GPU kernel optimization, multi-cloud infrastructure, and a developer-centric marketplace flywheel.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 4.9 Guest teaching 4.6 Guest disagreement 0.9 The host pushing back 0.9
05100:0010:0020:0030:001:06–4:53 · The host as informed peer 3/10 Company Origins and Early Talent Acquisition The host asks open-ended conversational questions about Fal's founding, early recruitment, and Batuhan's transition from Turkey and Poland. Batuhan shares his background as a core Python maintainer and compiler engineer, while the tone remains warm and collaborative.4:53–8:23 · The host as informed peer 4/10 Pivoting to Generative Media and Overcoming GPU Scarcity The host asks why the founders chose generative media over LLMs in 2021 when LLMs dominated attention. Burkay and Batuhan detail their early optimization work on Stable Diffusion 1.5 under severe GPU quota constraints.8:23–14:19 · The host as informed peer 4/10 First-Principles Kernel Optimization and Performance Engineering Batuhan provides detailed technical insight into GPU FLOP math, kernel sharding, and bottleneck profiling. The host tracks the broader industry timeline, referencing LLaMA-2 and Google Veo 3 capabilities.14:19–18:16 · The host as informed peer 7/10 Chained Workflows, ComfyUI, and Model Fine-Tuning The host demonstrates strong domain expertise by explaining chained workflows, latent upscaling, background removal, and ComfyUI dynamics compared to LLM zero-shot prompting. Batuhan expands on fine-tuning volume differences between text and image modalities.18:16–23:48 · The host as informed peer 5/10 Building a Two-Sided Marketplace for AI Models The host outlines how Fal evolved into a two-sided marketplace between developers and model builders. The guests detail rapid leapfrogging across foundation video models and their advisory role for developers.23:48–29:00 · The host as informed peer 6/10 Multi-Cloud Distributed Supercomputing Infrastructure Batuhan educates the host on why off-the-shelf Kubernetes failed for multi-cloud inference, requiring custom orchestration and NVMe memory caching. When the host asks if speed is a moat, Batuhan mildly counters that speed is not a static moat because open source eventually catches up.29:00–31:41 · The host as informed peer 4/10 Engineering Culture and Team Specialization The host asks how Fal structures its engineering and go-to-market teams to stay updated on new model releases. Burkay outlines their team composition, emphasizing that over half of their engineering staff consists of Applied ML specialists.31:41–35:52 · The host as informed peer 6/10 Technical Sales, Slack Connect, and Customer Obsession The host highlights how rare it is for deep infrastructure engineers to embrace technical sales and customer support. The guests explain their Slack Connect culture, while the host articulates modern dev-tool sales dynamics.1:06–4:53 · Guest teaching 3/10 Company Origins and Early Talent Acquisition The host asks open-ended conversational questions about Fal's founding, early recruitment, and Batuhan's transition from Turkey and Poland. Batuhan shares his background as a core Python maintainer and compiler engineer, while the tone remains warm and collaborative.4:53–8:23 · Guest teaching 4/10 Pivoting to Generative Media and Overcoming GPU Scarcity The host asks why the founders chose generative media over LLMs in 2021 when LLMs dominated attention. Burkay and Batuhan detail their early optimization work on Stable Diffusion 1.5 under severe GPU quota constraints.8:23–14:19 · Guest teaching 6/10 First-Principles Kernel Optimization and Performance Engineering Batuhan provides detailed technical insight into GPU FLOP math, kernel sharding, and bottleneck profiling. The host tracks the broader industry timeline, referencing LLaMA-2 and Google Veo 3 capabilities.14:19–18:16 · Guest teaching 5/10 Chained Workflows, ComfyUI, and Model Fine-Tuning The host demonstrates strong domain expertise by explaining chained workflows, latent upscaling, background removal, and ComfyUI dynamics compared to LLM zero-shot prompting. Batuhan expands on fine-tuning volume differences between text and image modalities.18:16–23:48 · Guest teaching 5/10 Building a Two-Sided Marketplace for AI Models The host outlines how Fal evolved into a two-sided marketplace between developers and model builders. The guests detail rapid leapfrogging across foundation video models and their advisory role for developers.23:48–29:00 · Guest teaching 7/10 Multi-Cloud Distributed Supercomputing Infrastructure Batuhan educates the host on why off-the-shelf Kubernetes failed for multi-cloud inference, requiring custom orchestration and NVMe memory caching. When the host asks if speed is a moat, Batuhan mildly counters that speed is not a static moat because open source eventually catches up.29:00–31:41 · Guest teaching 4/10 Engineering Culture and Team Specialization The host asks how Fal structures its engineering and go-to-market teams to stay updated on new model releases. Burkay outlines their team composition, emphasizing that over half of their engineering staff consists of Applied ML specialists.31:41–35:52 · Guest teaching 3/10 Technical Sales, Slack Connect, and Customer Obsession The host highlights how rare it is for deep infrastructure engineers to embrace technical sales and customer support. The guests explain their Slack Connect culture, while the host articulates modern dev-tool sales dynamics.1:06–4:53 · Guest disagreement 1/10 Company Origins and Early Talent Acquisition The host asks open-ended conversational questions about Fal's founding, early recruitment, and Batuhan's transition from Turkey and Poland. Batuhan shares his background as a core Python maintainer and compiler engineer, while the tone remains warm and collaborative.4:53–8:23 · Guest disagreement 1/10 Pivoting to Generative Media and Overcoming GPU Scarcity The host asks why the founders chose generative media over LLMs in 2021 when LLMs dominated attention. Burkay and Batuhan detail their early optimization work on Stable Diffusion 1.5 under severe GPU quota constraints.8:23–14:19 · Guest disagreement 1/10 First-Principles Kernel Optimization and Performance Engineering Batuhan provides detailed technical insight into GPU FLOP math, kernel sharding, and bottleneck profiling. The host tracks the broader industry timeline, referencing LLaMA-2 and Google Veo 3 capabilities.14:19–18:16 · Guest disagreement 1/10 Chained Workflows, ComfyUI, and Model Fine-Tuning The host demonstrates strong domain expertise by explaining chained workflows, latent upscaling, background removal, and ComfyUI dynamics compared to LLM zero-shot prompting. Batuhan expands on fine-tuning volume differences between text and image modalities.18:16–23:48 · Guest disagreement 1/10 Building a Two-Sided Marketplace for AI Models The host outlines how Fal evolved into a two-sided marketplace between developers and model builders. The guests detail rapid leapfrogging across foundation video models and their advisory role for developers.23:48–29:00 · Guest disagreement 2/10 Multi-Cloud Distributed Supercomputing Infrastructure Batuhan educates the host on why off-the-shelf Kubernetes failed for multi-cloud inference, requiring custom orchestration and NVMe memory caching. When the host asks if speed is a moat, Batuhan mildly counters that speed is not a static moat because open source eventually catches up.29:00–31:41 · Guest disagreement 0/10 Engineering Culture and Team Specialization The host asks how Fal structures its engineering and go-to-market teams to stay updated on new model releases. Burkay outlines their team composition, emphasizing that over half of their engineering staff consists of Applied ML specialists.31:41–35:52 · Guest disagreement 0/10 Technical Sales, Slack Connect, and Customer Obsession The host highlights how rare it is for deep infrastructure engineers to embrace technical sales and customer support. The guests explain their Slack Connect culture, while the host articulates modern dev-tool sales dynamics.1:06–4:53 · The host pushing back 0/10 Company Origins and Early Talent Acquisition The host asks open-ended conversational questions about Fal's founding, early recruitment, and Batuhan's transition from Turkey and Poland. Batuhan shares his background as a core Python maintainer and compiler engineer, while the tone remains warm and collaborative.4:53–8:23 · The host pushing back 1/10 Pivoting to Generative Media and Overcoming GPU Scarcity The host asks why the founders chose generative media over LLMs in 2021 when LLMs dominated attention. Burkay and Batuhan detail their early optimization work on Stable Diffusion 1.5 under severe GPU quota constraints.8:23–14:19 · The host pushing back 1/10 First-Principles Kernel Optimization and Performance Engineering Batuhan provides detailed technical insight into GPU FLOP math, kernel sharding, and bottleneck profiling. The host tracks the broader industry timeline, referencing LLaMA-2 and Google Veo 3 capabilities.14:19–18:16 · The host pushing back 2/10 Chained Workflows, ComfyUI, and Model Fine-Tuning The host demonstrates strong domain expertise by explaining chained workflows, latent upscaling, background removal, and ComfyUI dynamics compared to LLM zero-shot prompting. Batuhan expands on fine-tuning volume differences between text and image modalities.18:16–23:48 · The host pushing back 1/10 Building a Two-Sided Marketplace for AI Models The host outlines how Fal evolved into a two-sided marketplace between developers and model builders. The guests detail rapid leapfrogging across foundation video models and their advisory role for developers.23:48–29:00 · The host pushing back 2/10 Multi-Cloud Distributed Supercomputing Infrastructure Batuhan educates the host on why off-the-shelf Kubernetes failed for multi-cloud inference, requiring custom orchestration and NVMe memory caching. When the host asks if speed is a moat, Batuhan mildly counters that speed is not a static moat because open source eventually catches up.29:00–31:41 · The host pushing back 0/10 Engineering Culture and Team Specialization The host asks how Fal structures its engineering and go-to-market teams to stay updated on new model releases. Burkay outlines their team composition, emphasizing that over half of their engineering staff consists of Applied ML specialists.31:41–35:52 · The host pushing back 0/10 Technical Sales, Slack Connect, and Customer Obsession The host highlights how rare it is for deep infrastructure engineers to embrace technical sales and customer support. The guests explain their Slack Connect culture, while the host articulates modern dev-tool sales dynamics.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 28:17 Rejection of speed as a static moat

Batuhan directly pushes back against the host's framing that speed is a permanent moat, arguing that open-source alternatives inevitably catch up and require continuous innovation.

Hardest push from the host ▶ 4:53 Questioning choice of generative media over LLMs

The host challenges the founders on why they chose generative media back in 2021 when the entire industry consensus was focused on large language models.

Biggest teaching moment ▶ 25:40 Explaining Kubernetes cold-start failures

Batuhan educates the host on low-level infrastructure constraints, explaining how Kubernetes' five-second cold-start delays forced Fal to build proprietary multi-cloud orchestration.

The host holds their own ▶ 14:19 Demonstrating expertise in multi-step generative workflows

The host shows deep domain mastery by detailing how video/image pipelines require chaining, upscaling, and latent control, contrasting this with basic single-prompt LLM execution.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Company Origins and Early Talent Acquisition 3310 The host asks open-ended conversational questions about Fal's founding, early recruitment, and Batuhan's transition from Turkey and Poland. Batuhan shares his background as a core Python maintainer and compiler engineer, while the tone remains warm and collaborative.
Pivoting to Generative Media and Overcoming GPU Scarcity 4411 The host asks why the founders chose generative media over LLMs in 2021 when LLMs dominated attention. Burkay and Batuhan detail their early optimization work on Stable Diffusion 1.5 under severe GPU quota constraints.
First-Principles Kernel Optimization and Performance Engineering 4611 Batuhan provides detailed technical insight into GPU FLOP math, kernel sharding, and bottleneck profiling. The host tracks the broader industry timeline, referencing LLaMA-2 and Google Veo 3 capabilities.
Chained Workflows, ComfyUI, and Model Fine-Tuning 7512 The host demonstrates strong domain expertise by explaining chained workflows, latent upscaling, background removal, and ComfyUI dynamics compared to LLM zero-shot prompting. Batuhan expands on fine-tuning volume differences between text and image modalities.
Building a Two-Sided Marketplace for AI Models 5511 The host outlines how Fal evolved into a two-sided marketplace between developers and model builders. The guests detail rapid leapfrogging across foundation video models and their advisory role for developers.
Multi-Cloud Distributed Supercomputing Infrastructure 6722 Batuhan educates the host on why off-the-shelf Kubernetes failed for multi-cloud inference, requiring custom orchestration and NVMe memory caching. When the host asks if speed is a moat, Batuhan mildly counters that speed is not a static moat because open source eventually catches up.
Engineering Culture and Team Specialization 4400 The host asks how Fal structures its engineering and go-to-market teams to stay updated on new model releases. Burkay outlines their team composition, emphasizing that over half of their engineering staff consists of Applied ML specialists.
Technical Sales, Slack Connect, and Customer Obsession 6300 The host highlights how rare it is for deep infrastructure engineers to embrace technical sales and customer support. The guests explain their Slack Connect culture, while the host articulates modern dev-tool sales dynamics.

Statements from this episode (5)

Assertion Not checkable as stated
Gur: Model capability leaps trigger step-function adoption growth
“What we see internally is every time there's a big shift, In capabilities of models. The adoption and the use case is just, it's like a step function. It just keeps growing.”
Burkay Gur Aug 1, 2025 ▶ 13:34
Assertion Not checkable as stated
Gur: Imagen excels at character consistency while Flux dominates tooling workflows
“Just to give some examples, like Imogen three, Imogen four, very good at, like, character consistency. Flux, amazing at having an ecosystem of different tooling, you know, so that, like, you can just kind of use an existing workflow someone's built, you know, …”
Burkay Gur Aug 1, 2025 ▶ 18:38
Insight
Gur: AI video models are far from reaching marginal quality gains
“With video, we're earlier in the competition. There's still a lot of leapfrogging happening. There's just so much more to build, and there's just, like, you know, we haven't hit, like, a quality bar where there's just, like, marginal improvements.”
Burkay Gur Aug 1, 2025 ▶ 19:11
Assertion Not checkable as stated
Gur: Fal.ai's first 27 hires were all engineers
“Until we were like 28 people or so, we were all engineers, and I think like the 28th hire was a non engineer.”
Burkay Gur Aug 1, 2025 ▶ 29:36
Prediction Held up
Gur: Scaling AI models will enable interactive playable games
“I think what we all agree is that we're very excited about the net new use cases, recreational, you know, image generation, like people just like having fun with it, and that becoming like now video, that becoming maybe like short playable games or whatever, l…”
Burkay Gur Aug 1, 2025 ▶ 36:58
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.