Aug 1, 2025 · 38m · a16z
The Future of AI Video: How Fal.ai is Making AI Video Faster & Easier
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the a16z podcast, Fal.ai co-founders Burkay Gur and Batuhan Taskaya discuss how their company built a dominant high-performance inference platform for generative AI video and media through first-principles GPU kernel optimization, multi-cloud infrastructure, and a developer-centric marketplace flywheel.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the host, purple is the guest (3 minute bins)
Batuhan directly pushes back against the host's framing that speed is a permanent moat, arguing that open-source alternatives inevitably catch up and require continuous innovation.
Hardest push from the host ▶ 4:53 Questioning choice of generative media over LLMsThe host challenges the founders on why they chose generative media back in 2021 when the entire industry consensus was focused on large language models.
Biggest teaching moment ▶ 25:40 Explaining Kubernetes cold-start failuresBatuhan educates the host on low-level infrastructure constraints, explaining how Kubernetes' five-second cold-start delays forced Fal to build proprietary multi-cloud orchestration.
The host holds their own ▶ 14:19 Demonstrating expertise in multi-step generative workflowsThe host shows deep domain mastery by detailing how video/image pipelines require chaining, upscaling, and latent control, contrasting this with basic single-prompt LLM execution.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The host as informed peer | Guest teaching | Guest disagreement | The host pushing back | Why |
|---|---|---|---|---|---|---|
| Company Origins and Early Talent Acquisition | 3 | 3 | 1 | 0 | The host asks open-ended conversational questions about Fal's founding, early recruitment, and Batuhan's transition from Turkey and Poland. Batuhan shares his background as a core Python maintainer and compiler engineer, while the tone remains warm and collaborative. | |
| Pivoting to Generative Media and Overcoming GPU Scarcity | 4 | 4 | 1 | 1 | The host asks why the founders chose generative media over LLMs in 2021 when LLMs dominated attention. Burkay and Batuhan detail their early optimization work on Stable Diffusion 1.5 under severe GPU quota constraints. | |
| First-Principles Kernel Optimization and Performance Engineering | 4 | 6 | 1 | 1 | Batuhan provides detailed technical insight into GPU FLOP math, kernel sharding, and bottleneck profiling. The host tracks the broader industry timeline, referencing LLaMA-2 and Google Veo 3 capabilities. | |
| Chained Workflows, ComfyUI, and Model Fine-Tuning | 7 | 5 | 1 | 2 | The host demonstrates strong domain expertise by explaining chained workflows, latent upscaling, background removal, and ComfyUI dynamics compared to LLM zero-shot prompting. Batuhan expands on fine-tuning volume differences between text and image modalities. | |
| Building a Two-Sided Marketplace for AI Models | 5 | 5 | 1 | 1 | The host outlines how Fal evolved into a two-sided marketplace between developers and model builders. The guests detail rapid leapfrogging across foundation video models and their advisory role for developers. | |
| Multi-Cloud Distributed Supercomputing Infrastructure | 6 | 7 | 2 | 2 | Batuhan educates the host on why off-the-shelf Kubernetes failed for multi-cloud inference, requiring custom orchestration and NVMe memory caching. When the host asks if speed is a moat, Batuhan mildly counters that speed is not a static moat because open source eventually catches up. | |
| Engineering Culture and Team Specialization | 4 | 4 | 0 | 0 | The host asks how Fal structures its engineering and go-to-market teams to stay updated on new model releases. Burkay outlines their team composition, emphasizing that over half of their engineering staff consists of Applied ML specialists. | |
| Technical Sales, Slack Connect, and Customer Obsession | 6 | 3 | 0 | 0 | The host highlights how rare it is for deep infrastructure engineers to embrace technical sales and customer support. The guests explain their Slack Connect culture, while the host articulates modern dev-tool sales dynamics. |