Oct 22, 2025 · 59m · in-depth

The pivot that paid off: How fal found explosive growth | Gorkem Yurtseven (Co-founder and CTO)

Gorkem Yurtseven · 43m spoken Todd Jackson · 11m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In an in-depth conversation with First Round Capital's Todd Jackson, FAL co-founder and CTO Gorkem Yurtseven explains how the startup pivoted from data infrastructure to generative media inference, scaling from $2 million to over $100 million in ARR in just one year. Yurtseven shares technical insights on GPU optimization, bottom-up developer growth, enterprise expansion across Hollywood and creative tech, and unconventional organizational practices.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

Brett as informed peer 3.9 Guest teaching 4.1 Guest disagreement 0.4 Brett pushing back 0.4
05100:0015:0030:0045:001:36–3:43 · Brett as informed peer 4/10 What FAL Does and the $100M ARR Growth Surge Todd sets the context by detailing fal's rapid growth trajectory from $2M to $100M ARR over 18 months. Gorkem explains the timeline, noting the slow summer around SDXL before the Flux and AI video boom triggered massive scaling.3:43–6:58 · Brett as informed peer 5/10 Founding Origins and the Catalyst for Pivoting to Inference Todd asks about the company's founding story and original data infrastructure premise. Gorkem breaks down how the release of DALL-E, Stable Diffusion, and ChatGPT proved off-the-shelf pre-trained models made custom data preparation obsolete.6:58–9:38 · Brett as informed peer 6/10 Deciding to Walk Away from Paying Customers Todd highlights his role as seed investor and recalls advising them on the $1M vs $10M ARR framework. Gorkem notes that while their initial forecasts were technically incorrect, the strategic framework was decisive for executing the pivot.9:38–13:20 · Brett as informed peer 4/10 Focusing on Generative Media and Series A Fundraising Hurdles Todd probes why fal chose generative media over LLMs when competitors like Together AI and Baseten emerged. Gorkem educates him on the technical and architectural differences in image inference and the uphill battle of pitching Series A investors fatigued by generic inference pitches.13:20–16:40 · Brett as informed peer 4/10 Viral Real-Time Demos and First Product Architecture Todd asks about the viral George Clooney webcam demo and the architectural choices behind low-latency inference. Gorkem details writing Triton kernels, system optimizations, and deliberately choosing opinionated API endpoints over generic GPU orchestration.16:40–19:22 · Brett as informed peer 4/10 Early Developer Adoption and Day-Zero Flux Launch Todd questions whether early users were merely hobbyist toy builders. Gorkem counters by pointing out that developers were spending tens of thousands of dollars daily, demonstrating immediate commercial reality, and explains their Day-Zero Flux launch via early relationships.19:22–21:27 · Brett as informed peer 4/10 The AI Video Boom and Infrastructure Demands Todd inquires about the signals behind fal's early bet on AI video. Gorkem explains the sudden migration of elite diffusion researchers to video post-Sora and how multi-GPU cluster requirements made inference optimizations far higher-leverage.21:28–24:20 · Brett as informed peer 3/10 Operationalizing Agility and GPU Capacity Management Todd asks how a 45-person company operationalizes rapid model deployment. Gorkem describes their 15-person Applied ML team, speed-running model deployments publicly on live streams, and managing elastic GPU capacity.24:21–26:45 · Brett as informed peer 3/10 Hollywood Studios and the Future of AI Video Production Todd brings up generative AI in Hollywood feature films. Gorkem explains how studio sentiment shifted rapidly from hesitation to aggressive inbound exploration as creative teams realized AI enhanced workflows rather than purely replacing personnel.26:46–30:07 · Brett as informed peer 4/10 Generative Media as a Greenfield Market and Model Commoditization Todd explores why generative media represents a greenfield market. Gorkem delivers a detailed breakdown of why LLMs threaten search giants while media is net-new, and explains the structural mechanics of model commoditization across distillation and research leaks.30:07–34:10 · Brett as informed peer 4/10 Under the Hood: Multi-Model Orchestration and Caching Strategies Todd prompts Gorkem to share fal's backend engineering secrets. Gorkem details the technical complexity of serving 600 models across 28 data centers, cold start mitigation, in-memory caching strategies, and non-linear GPU scaling physics.34:11–37:24 · Brett as informed peer 3/10 Developer Obsession and Establishing Category Leadership Todd observes fal's obsessive responsiveness in developer channels. Gorkem explains monitoring response metrics across 500 enterprise Slack channels and how branding the company around 'generative media platform' created category authority.37:25–42:06 · Brett as informed peer 4/10 Navigating Enterprise Compliance and Tracking North Star Metrics Todd explores enterprise compliance and forecasting revenue growth. Gorkem explains navigating rigorous legal reviews, shifting from pay-as-you-go to annual commitments, and playfully dismisses non-revenue PM vanity metrics in favor of top-line revenue as the sole North Star.42:07–46:00 · Brett as informed peer 3/10 Engineering Recruitment and Identifying Non-Traditional Talent Todd asks how fal recruits elite ML and systems engineers without matching big tech salaries. Gorkem discusses hiring low-level systems engineers from databases, sourcing from their Turkish network, and evaluating genuine domain obsession over conventional pedigree.46:00–50:27 · Brett as informed peer 4/10 Product-Led Enterprise Sales and Internal AI Adoption Todd highlights fal reaching $100M ARR with fewer than 10 go-to-market personnel. Gorkem outlines their inbound-led enterprise sales model, where automated Salesforce triggers flag developers spending over $300/day to convert them into annual contracts.50:27–54:22 · Brett as informed peer 3/10 Authentic Developer Marketing and the 'GPU Poor' Hats Todd asks about fal's distinct marketing aesthetic and viral swag. Gorkem shares the backstory of partnering with designer Adam Ho and capitalizing on Dylan Patel's 'GPU poor' blog post to create viral conference merchandise.54:22–56:01 · Brett as informed peer 4/10 Eliminating Engineering Managers and Rethinking 1-on-1s Todd probes fal's operational quirks and asks when their management model without engineering managers will break. Gorkem defends the model, explaining that replacing traditional 1-on-1s with cross-functional 1-on-3/4 group sessions prevents meetings from degenerating into complaining sessions.1:36–3:43 · Guest teaching 3/10 What FAL Does and the $100M ARR Growth Surge Todd sets the context by detailing fal's rapid growth trajectory from $2M to $100M ARR over 18 months. Gorkem explains the timeline, noting the slow summer around SDXL before the Flux and AI video boom triggered massive scaling.3:43–6:58 · Guest teaching 4/10 Founding Origins and the Catalyst for Pivoting to Inference Todd asks about the company's founding story and original data infrastructure premise. Gorkem breaks down how the release of DALL-E, Stable Diffusion, and ChatGPT proved off-the-shelf pre-trained models made custom data preparation obsolete.6:58–9:38 · Guest teaching 2/10 Deciding to Walk Away from Paying Customers Todd highlights his role as seed investor and recalls advising them on the $1M vs $10M ARR framework. Gorkem notes that while their initial forecasts were technically incorrect, the strategic framework was decisive for executing the pivot.9:38–13:20 · Guest teaching 5/10 Focusing on Generative Media and Series A Fundraising Hurdles Todd probes why fal chose generative media over LLMs when competitors like Together AI and Baseten emerged. Gorkem educates him on the technical and architectural differences in image inference and the uphill battle of pitching Series A investors fatigued by generic inference pitches.13:20–16:40 · Guest teaching 5/10 Viral Real-Time Demos and First Product Architecture Todd asks about the viral George Clooney webcam demo and the architectural choices behind low-latency inference. Gorkem details writing Triton kernels, system optimizations, and deliberately choosing opinionated API endpoints over generic GPU orchestration.16:40–19:22 · Guest teaching 4/10 Early Developer Adoption and Day-Zero Flux Launch Todd questions whether early users were merely hobbyist toy builders. Gorkem counters by pointing out that developers were spending tens of thousands of dollars daily, demonstrating immediate commercial reality, and explains their Day-Zero Flux launch via early relationships.19:22–21:27 · Guest teaching 5/10 The AI Video Boom and Infrastructure Demands Todd inquires about the signals behind fal's early bet on AI video. Gorkem explains the sudden migration of elite diffusion researchers to video post-Sora and how multi-GPU cluster requirements made inference optimizations far higher-leverage.21:28–24:20 · Guest teaching 4/10 Operationalizing Agility and GPU Capacity Management Todd asks how a 45-person company operationalizes rapid model deployment. Gorkem describes their 15-person Applied ML team, speed-running model deployments publicly on live streams, and managing elastic GPU capacity.24:21–26:45 · Guest teaching 4/10 Hollywood Studios and the Future of AI Video Production Todd brings up generative AI in Hollywood feature films. Gorkem explains how studio sentiment shifted rapidly from hesitation to aggressive inbound exploration as creative teams realized AI enhanced workflows rather than purely replacing personnel.26:46–30:07 · Guest teaching 6/10 Generative Media as a Greenfield Market and Model Commoditization Todd explores why generative media represents a greenfield market. Gorkem delivers a detailed breakdown of why LLMs threaten search giants while media is net-new, and explains the structural mechanics of model commoditization across distillation and research leaks.30:07–34:10 · Guest teaching 6/10 Under the Hood: Multi-Model Orchestration and Caching Strategies Todd prompts Gorkem to share fal's backend engineering secrets. Gorkem details the technical complexity of serving 600 models across 28 data centers, cold start mitigation, in-memory caching strategies, and non-linear GPU scaling physics.34:11–37:24 · Guest teaching 3/10 Developer Obsession and Establishing Category Leadership Todd observes fal's obsessive responsiveness in developer channels. Gorkem explains monitoring response metrics across 500 enterprise Slack channels and how branding the company around 'generative media platform' created category authority.37:25–42:06 · Guest teaching 4/10 Navigating Enterprise Compliance and Tracking North Star Metrics Todd explores enterprise compliance and forecasting revenue growth. Gorkem explains navigating rigorous legal reviews, shifting from pay-as-you-go to annual commitments, and playfully dismisses non-revenue PM vanity metrics in favor of top-line revenue as the sole North Star.42:07–46:00 · Guest teaching 4/10 Engineering Recruitment and Identifying Non-Traditional Talent Todd asks how fal recruits elite ML and systems engineers without matching big tech salaries. Gorkem discusses hiring low-level systems engineers from databases, sourcing from their Turkish network, and evaluating genuine domain obsession over conventional pedigree.46:00–50:27 · Guest teaching 4/10 Product-Led Enterprise Sales and Internal AI Adoption Todd highlights fal reaching $100M ARR with fewer than 10 go-to-market personnel. Gorkem outlines their inbound-led enterprise sales model, where automated Salesforce triggers flag developers spending over $300/day to convert them into annual contracts.50:27–54:22 · Guest teaching 2/10 Authentic Developer Marketing and the 'GPU Poor' Hats Todd asks about fal's distinct marketing aesthetic and viral swag. Gorkem shares the backstory of partnering with designer Adam Ho and capitalizing on Dylan Patel's 'GPU poor' blog post to create viral conference merchandise.54:22–56:01 · Guest teaching 4/10 Eliminating Engineering Managers and Rethinking 1-on-1s Todd probes fal's operational quirks and asks when their management model without engineering managers will break. Gorkem defends the model, explaining that replacing traditional 1-on-1s with cross-functional 1-on-3/4 group sessions prevents meetings from degenerating into complaining sessions.1:36–3:43 · Guest disagreement 0/10 What FAL Does and the $100M ARR Growth Surge Todd sets the context by detailing fal's rapid growth trajectory from $2M to $100M ARR over 18 months. Gorkem explains the timeline, noting the slow summer around SDXL before the Flux and AI video boom triggered massive scaling.3:43–6:58 · Guest disagreement 0/10 Founding Origins and the Catalyst for Pivoting to Inference Todd asks about the company's founding story and original data infrastructure premise. Gorkem breaks down how the release of DALL-E, Stable Diffusion, and ChatGPT proved off-the-shelf pre-trained models made custom data preparation obsolete.6:58–9:38 · Guest disagreement 0/10 Deciding to Walk Away from Paying Customers Todd highlights his role as seed investor and recalls advising them on the $1M vs $10M ARR framework. Gorkem notes that while their initial forecasts were technically incorrect, the strategic framework was decisive for executing the pivot.9:38–13:20 · Guest disagreement 1/10 Focusing on Generative Media and Series A Fundraising Hurdles Todd probes why fal chose generative media over LLMs when competitors like Together AI and Baseten emerged. Gorkem educates him on the technical and architectural differences in image inference and the uphill battle of pitching Series A investors fatigued by generic inference pitches.13:20–16:40 · Guest disagreement 0/10 Viral Real-Time Demos and First Product Architecture Todd asks about the viral George Clooney webcam demo and the architectural choices behind low-latency inference. Gorkem details writing Triton kernels, system optimizations, and deliberately choosing opinionated API endpoints over generic GPU orchestration.16:40–19:22 · Guest disagreement 1/10 Early Developer Adoption and Day-Zero Flux Launch Todd questions whether early users were merely hobbyist toy builders. Gorkem counters by pointing out that developers were spending tens of thousands of dollars daily, demonstrating immediate commercial reality, and explains their Day-Zero Flux launch via early relationships.19:22–21:27 · Guest disagreement 0/10 The AI Video Boom and Infrastructure Demands Todd inquires about the signals behind fal's early bet on AI video. Gorkem explains the sudden migration of elite diffusion researchers to video post-Sora and how multi-GPU cluster requirements made inference optimizations far higher-leverage.21:28–24:20 · Guest disagreement 0/10 Operationalizing Agility and GPU Capacity Management Todd asks how a 45-person company operationalizes rapid model deployment. Gorkem describes their 15-person Applied ML team, speed-running model deployments publicly on live streams, and managing elastic GPU capacity.24:21–26:45 · Guest disagreement 0/10 Hollywood Studios and the Future of AI Video Production Todd brings up generative AI in Hollywood feature films. Gorkem explains how studio sentiment shifted rapidly from hesitation to aggressive inbound exploration as creative teams realized AI enhanced workflows rather than purely replacing personnel.26:46–30:07 · Guest disagreement 1/10 Generative Media as a Greenfield Market and Model Commoditization Todd explores why generative media represents a greenfield market. Gorkem delivers a detailed breakdown of why LLMs threaten search giants while media is net-new, and explains the structural mechanics of model commoditization across distillation and research leaks.30:07–34:10 · Guest disagreement 0/10 Under the Hood: Multi-Model Orchestration and Caching Strategies Todd prompts Gorkem to share fal's backend engineering secrets. Gorkem details the technical complexity of serving 600 models across 28 data centers, cold start mitigation, in-memory caching strategies, and non-linear GPU scaling physics.34:11–37:24 · Guest disagreement 0/10 Developer Obsession and Establishing Category Leadership Todd observes fal's obsessive responsiveness in developer channels. Gorkem explains monitoring response metrics across 500 enterprise Slack channels and how branding the company around 'generative media platform' created category authority.37:25–42:06 · Guest disagreement 2/10 Navigating Enterprise Compliance and Tracking North Star Metrics Todd explores enterprise compliance and forecasting revenue growth. Gorkem explains navigating rigorous legal reviews, shifting from pay-as-you-go to annual commitments, and playfully dismisses non-revenue PM vanity metrics in favor of top-line revenue as the sole North Star.42:07–46:00 · Guest disagreement 0/10 Engineering Recruitment and Identifying Non-Traditional Talent Todd asks how fal recruits elite ML and systems engineers without matching big tech salaries. Gorkem discusses hiring low-level systems engineers from databases, sourcing from their Turkish network, and evaluating genuine domain obsession over conventional pedigree.46:00–50:27 · Guest disagreement 0/10 Product-Led Enterprise Sales and Internal AI Adoption Todd highlights fal reaching $100M ARR with fewer than 10 go-to-market personnel. Gorkem outlines their inbound-led enterprise sales model, where automated Salesforce triggers flag developers spending over $300/day to convert them into annual contracts.50:27–54:22 · Guest disagreement 0/10 Authentic Developer Marketing and the 'GPU Poor' Hats Todd asks about fal's distinct marketing aesthetic and viral swag. Gorkem shares the backstory of partnering with designer Adam Ho and capitalizing on Dylan Patel's 'GPU poor' blog post to create viral conference merchandise.54:22–56:01 · Guest disagreement 2/10 Eliminating Engineering Managers and Rethinking 1-on-1s Todd probes fal's operational quirks and asks when their management model without engineering managers will break. Gorkem defends the model, explaining that replacing traditional 1-on-1s with cross-functional 1-on-3/4 group sessions prevents meetings from degenerating into complaining sessions.1:36–3:43 · Brett pushing back 0/10 What FAL Does and the $100M ARR Growth Surge Todd sets the context by detailing fal's rapid growth trajectory from $2M to $100M ARR over 18 months. Gorkem explains the timeline, noting the slow summer around SDXL before the Flux and AI video boom triggered massive scaling.3:43–6:58 · Brett pushing back 0/10 Founding Origins and the Catalyst for Pivoting to Inference Todd asks about the company's founding story and original data infrastructure premise. Gorkem breaks down how the release of DALL-E, Stable Diffusion, and ChatGPT proved off-the-shelf pre-trained models made custom data preparation obsolete.6:58–9:38 · Brett pushing back 1/10 Deciding to Walk Away from Paying Customers Todd highlights his role as seed investor and recalls advising them on the $1M vs $10M ARR framework. Gorkem notes that while their initial forecasts were technically incorrect, the strategic framework was decisive for executing the pivot.9:38–13:20 · Brett pushing back 1/10 Focusing on Generative Media and Series A Fundraising Hurdles Todd probes why fal chose generative media over LLMs when competitors like Together AI and Baseten emerged. Gorkem educates him on the technical and architectural differences in image inference and the uphill battle of pitching Series A investors fatigued by generic inference pitches.13:20–16:40 · Brett pushing back 0/10 Viral Real-Time Demos and First Product Architecture Todd asks about the viral George Clooney webcam demo and the architectural choices behind low-latency inference. Gorkem details writing Triton kernels, system optimizations, and deliberately choosing opinionated API endpoints over generic GPU orchestration.16:40–19:22 · Brett pushing back 2/10 Early Developer Adoption and Day-Zero Flux Launch Todd questions whether early users were merely hobbyist toy builders. Gorkem counters by pointing out that developers were spending tens of thousands of dollars daily, demonstrating immediate commercial reality, and explains their Day-Zero Flux launch via early relationships.19:22–21:27 · Brett pushing back 0/10 The AI Video Boom and Infrastructure Demands Todd inquires about the signals behind fal's early bet on AI video. Gorkem explains the sudden migration of elite diffusion researchers to video post-Sora and how multi-GPU cluster requirements made inference optimizations far higher-leverage.21:28–24:20 · Brett pushing back 0/10 Operationalizing Agility and GPU Capacity Management Todd asks how a 45-person company operationalizes rapid model deployment. Gorkem describes their 15-person Applied ML team, speed-running model deployments publicly on live streams, and managing elastic GPU capacity.24:21–26:45 · Brett pushing back 0/10 Hollywood Studios and the Future of AI Video Production Todd brings up generative AI in Hollywood feature films. Gorkem explains how studio sentiment shifted rapidly from hesitation to aggressive inbound exploration as creative teams realized AI enhanced workflows rather than purely replacing personnel.26:46–30:07 · Brett pushing back 0/10 Generative Media as a Greenfield Market and Model Commoditization Todd explores why generative media represents a greenfield market. Gorkem delivers a detailed breakdown of why LLMs threaten search giants while media is net-new, and explains the structural mechanics of model commoditization across distillation and research leaks.30:07–34:10 · Brett pushing back 0/10 Under the Hood: Multi-Model Orchestration and Caching Strategies Todd prompts Gorkem to share fal's backend engineering secrets. Gorkem details the technical complexity of serving 600 models across 28 data centers, cold start mitigation, in-memory caching strategies, and non-linear GPU scaling physics.34:11–37:24 · Brett pushing back 0/10 Developer Obsession and Establishing Category Leadership Todd observes fal's obsessive responsiveness in developer channels. Gorkem explains monitoring response metrics across 500 enterprise Slack channels and how branding the company around 'generative media platform' created category authority.37:25–42:06 · Brett pushing back 1/10 Navigating Enterprise Compliance and Tracking North Star Metrics Todd explores enterprise compliance and forecasting revenue growth. Gorkem explains navigating rigorous legal reviews, shifting from pay-as-you-go to annual commitments, and playfully dismisses non-revenue PM vanity metrics in favor of top-line revenue as the sole North Star.42:07–46:00 · Brett pushing back 0/10 Engineering Recruitment and Identifying Non-Traditional Talent Todd asks how fal recruits elite ML and systems engineers without matching big tech salaries. Gorkem discusses hiring low-level systems engineers from databases, sourcing from their Turkish network, and evaluating genuine domain obsession over conventional pedigree.46:00–50:27 · Brett pushing back 0/10 Product-Led Enterprise Sales and Internal AI Adoption Todd highlights fal reaching $100M ARR with fewer than 10 go-to-market personnel. Gorkem outlines their inbound-led enterprise sales model, where automated Salesforce triggers flag developers spending over $300/day to convert them into annual contracts.50:27–54:22 · Brett pushing back 0/10 Authentic Developer Marketing and the 'GPU Poor' Hats Todd asks about fal's distinct marketing aesthetic and viral swag. Gorkem shares the backstory of partnering with designer Adam Ho and capitalizing on Dylan Patel's 'GPU poor' blog post to create viral conference merchandise.54:22–56:01 · Brett pushing back 2/10 Eliminating Engineering Managers and Rethinking 1-on-1s Todd probes fal's operational quirks and asks when their management model without engineering managers will break. Gorkem defends the model, explaining that replacing traditional 1-on-1s with cross-functional 1-on-3/4 group sessions prevents meetings from degenerating into complaining sessions.

speaking balance: gold is Brett, purple is the guest (3 minute bins)

0:00 · Brett 0% · guest 100%0:00 · Brett 0% · guest 100%3:00 · Brett 0% · guest 100%3:00 · Brett 0% · guest 100%6:00 · Brett 0% · guest 100%6:00 · Brett 0% · guest 100%9:00 · Brett 0% · guest 100%9:00 · Brett 0% · guest 100%12:00 · Brett 0% · guest 100%12:00 · Brett 0% · guest 100%15:00 · Brett 0% · guest 100%15:00 · Brett 0% · guest 100%18:00 · Brett 0% · guest 100%18:00 · Brett 0% · guest 100%21:00 · Brett 0% · guest 100%21:00 · Brett 0% · guest 100%24:00 · Brett 0% · guest 100%24:00 · Brett 0% · guest 100%27:00 · Brett 0% · guest 100%27:00 · Brett 0% · guest 100%30:00 · Brett 0% · guest 100%30:00 · Brett 0% · guest 100%33:00 · Brett 0% · guest 100%33:00 · Brett 0% · guest 100%36:00 · Brett 0% · guest 100%36:00 · Brett 0% · guest 100%39:00 · Brett 0% · guest 100%39:00 · Brett 0% · guest 100%42:00 · Brett 0% · guest 100%42:00 · Brett 0% · guest 100%45:00 · Brett 0% · guest 100%45:00 · Brett 0% · guest 100%48:00 · Brett 0% · guest 100%48:00 · Brett 0% · guest 100%51:00 · Brett 0% · guest 100%51:00 · Brett 0% · guest 100%54:00 · Brett 0% · guest 100%54:00 · Brett 0% · guest 100%57:00 · Brett 0% · guest 100%57:00 · Brett 0% · guest 100%
Sharpest disagreement ▶ 55:13 Critique of traditional 1-on-1 management meetings

Gorkem forcefully critiques standard engineering manager 1-on-1s as artificial forums that prompt employees to complain, contrasting it with fal's collaborative small-group meetings.

Hardest push from Brett ▶ 55:12 Challenging the scalability of fal's no-manager model

Todd directly challenges Gorkem on operating without any engineering managers across 34 engineers, pressing him on when that organizational structure will inevitably break.

Biggest teaching moment ▶ 29:05 Masterclass on AI model commoditization

Gorkem lays out a detailed structural analysis of why frontier AI models cannot sustain quality moats due to research leaks, reproducible proof-of-concept, and rapid distillation.

Brett holds their own ▶ 9:02 Citing the seed-stage ARR milestone decision framework

Todd demonstrates his early investor insight by recalling the exact strategic framework ($1M vs $10M ARR velocity) he gave the founders to navigate their pivotal transition.

the scores for every segment, with the reasoning behind each
ChapterTopicBrett as informed peerGuest teachingGuest disagreementBrett pushing backWhy
What FAL Does and the $100M ARR Growth Surge 4300 Todd sets the context by detailing fal's rapid growth trajectory from $2M to $100M ARR over 18 months. Gorkem explains the timeline, noting the slow summer around SDXL before the Flux and AI video boom triggered massive scaling.
Founding Origins and the Catalyst for Pivoting to Inference 5400 Todd asks about the company's founding story and original data infrastructure premise. Gorkem breaks down how the release of DALL-E, Stable Diffusion, and ChatGPT proved off-the-shelf pre-trained models made custom data preparation obsolete.
Deciding to Walk Away from Paying Customers 6201 Todd highlights his role as seed investor and recalls advising them on the $1M vs $10M ARR framework. Gorkem notes that while their initial forecasts were technically incorrect, the strategic framework was decisive for executing the pivot.
Focusing on Generative Media and Series A Fundraising Hurdles 4511 Todd probes why fal chose generative media over LLMs when competitors like Together AI and Baseten emerged. Gorkem educates him on the technical and architectural differences in image inference and the uphill battle of pitching Series A investors fatigued by generic inference pitches.
Viral Real-Time Demos and First Product Architecture 4500 Todd asks about the viral George Clooney webcam demo and the architectural choices behind low-latency inference. Gorkem details writing Triton kernels, system optimizations, and deliberately choosing opinionated API endpoints over generic GPU orchestration.
Early Developer Adoption and Day-Zero Flux Launch 4412 Todd questions whether early users were merely hobbyist toy builders. Gorkem counters by pointing out that developers were spending tens of thousands of dollars daily, demonstrating immediate commercial reality, and explains their Day-Zero Flux launch via early relationships.
The AI Video Boom and Infrastructure Demands 4500 Todd inquires about the signals behind fal's early bet on AI video. Gorkem explains the sudden migration of elite diffusion researchers to video post-Sora and how multi-GPU cluster requirements made inference optimizations far higher-leverage.
Operationalizing Agility and GPU Capacity Management 3400 Todd asks how a 45-person company operationalizes rapid model deployment. Gorkem describes their 15-person Applied ML team, speed-running model deployments publicly on live streams, and managing elastic GPU capacity.
Hollywood Studios and the Future of AI Video Production 3400 Todd brings up generative AI in Hollywood feature films. Gorkem explains how studio sentiment shifted rapidly from hesitation to aggressive inbound exploration as creative teams realized AI enhanced workflows rather than purely replacing personnel.
Generative Media as a Greenfield Market and Model Commoditization 4610 Todd explores why generative media represents a greenfield market. Gorkem delivers a detailed breakdown of why LLMs threaten search giants while media is net-new, and explains the structural mechanics of model commoditization across distillation and research leaks.
Under the Hood: Multi-Model Orchestration and Caching Strategies 4600 Todd prompts Gorkem to share fal's backend engineering secrets. Gorkem details the technical complexity of serving 600 models across 28 data centers, cold start mitigation, in-memory caching strategies, and non-linear GPU scaling physics.
Developer Obsession and Establishing Category Leadership 3300 Todd observes fal's obsessive responsiveness in developer channels. Gorkem explains monitoring response metrics across 500 enterprise Slack channels and how branding the company around 'generative media platform' created category authority.
Navigating Enterprise Compliance and Tracking North Star Metrics 4421 Todd explores enterprise compliance and forecasting revenue growth. Gorkem explains navigating rigorous legal reviews, shifting from pay-as-you-go to annual commitments, and playfully dismisses non-revenue PM vanity metrics in favor of top-line revenue as the sole North Star.
Engineering Recruitment and Identifying Non-Traditional Talent 3400 Todd asks how fal recruits elite ML and systems engineers without matching big tech salaries. Gorkem discusses hiring low-level systems engineers from databases, sourcing from their Turkish network, and evaluating genuine domain obsession over conventional pedigree.
Product-Led Enterprise Sales and Internal AI Adoption 4400 Todd highlights fal reaching $100M ARR with fewer than 10 go-to-market personnel. Gorkem outlines their inbound-led enterprise sales model, where automated Salesforce triggers flag developers spending over $300/day to convert them into annual contracts.
Authentic Developer Marketing and the 'GPU Poor' Hats 3200 Todd asks about fal's distinct marketing aesthetic and viral swag. Gorkem shares the backstory of partnering with designer Adam Ho and capitalizing on Dylan Patel's 'GPU poor' blog post to create viral conference merchandise.
Eliminating Engineering Managers and Rethinking 1-on-1s 4422 Todd probes fal's operational quirks and asks when their management model without engineering managers will break. Gorkem defends the model, explaining that replacing traditional 1-on-1s with cross-functional 1-on-3/4 group sessions prevents meetings from degenerating into complaining sessions.

Statements from this episode (28)

Insight
Off-the-shelf models eliminate data prep for all but the largest enterprises
“If there's a ready-made model that changes everything, this whole like data preparation stage. Can be skipped and only like the biggest of the companies are going to do that. Everyone else, they'll just use something off the shelf.”
Gorkem Yurtseven Oct 22, 2025 ▶ 6:44
Insight
Running two different products simultaneously creates conflicting messaging for sales
“It's very hard when you are not screaming exactly what you are doing to your customers, to the potential customers, to people you work with. It's really hard to sell because, you know, they look at your website, they see something else, like So it is really ha…”
Gorkem Yurtseven Oct 22, 2025 ▶ 7:29
Disclosure
fal struggled to raise Series A because VCs doubted image inference
“And we tried to explain this to people because it was so new, like no one got it. Everyone thought, An inference platform is an inference platform. Doesn't matter what kind of model it is. There are other people who are more qualified or more prepared to do th…”
Gorkem Yurtseven Oct 22, 2025 ▶ 12:16
Insight
Raising alongside competitors puts startups at disadvantage due to investor fatigue
“And that's something we underestimated how, how disadvantaged of a situation it is to raise at the same time with seemingly all your competitors, because they all say the same story. Investors hear it over and over again, and there's some Fatigue of hearing th…”
Gorkem Yurtseven Oct 22, 2025 ▶ 13:04
Disclosure
fal never monetized its viral real-time image-to-image inference demo
“Still, we didn't make any money from that demo. It makes a really impressive, you know, technical demo for people to see how fast we can run inference, but we couldn't find any use case for that, like very fast image to image inference. Still to this day, it i…”
Gorkem Yurtseven Oct 22, 2025 ▶ 14:01
Disclosure
fal built managed APIs to control code over arbitrary GPU orchestration
“Instead of Focusing on like GPU orchestration and letting people deploy whatever they want. We decided to build, you know, APIs. Every single code that's deployed is owned by us and we control the whole process.”
Gorkem Yurtseven Oct 22, 2025 ▶ 15:49
Insight
Optimizing common AI workflows creates more value than arbitrary custom deployments
“Everyone wants to do the same thing over and over again. And therefore we thought there is value in actually optimizing the most common workflow.”
Gorkem Yurtseven Oct 22, 2025 ▶ 16:26
Disclosure
Early indie developer customers on fal spent tens of thousands daily
“Everyone was spending serious money on the platform, tens of thousands of dollars a day, all of a sudden.”
Gorkem Yurtseven Oct 22, 2025 ▶ 17:30
Assertion Not checkable as stated
Top diffusion researchers abandoned image for video generation chasing VC hype
“There were so much to do there, but the price for video was much shinier and there was a lot of VC money being poured into it. All the researchers left doing image research and then focused on video.”
Gorkem Yurtseven Oct 22, 2025 ▶ 20:30
Insight
Latency percentage gains matter substantially more on minute-long AI workloads
“If something takes 1:02, if you can shave off 20% of it, maybe not enough people care about it. But if something takes a minute and you can shave off a similar percentage, all of a sudden that's a lot more meaningful.”
Gorkem Yurtseven Oct 22, 2025 ▶ 21:15
Disclosure
fal maintains a 15-person applied ML team dedicated to model deployment
“We have a Applied ML team. It's around 15 people right now. And, you know, all they do every day is either deploy these models, optimize them, play with them, and they're obsessed with it.”
Gorkem Yurtseven Oct 22, 2025 ▶ 22:00
Assertion Not checkable as stated
AI research labs actively approach fal for day-zero model releases
“And now with FAL's position in the market, we also get some early information from the research labs. Everyone like tries to talk to us. And release their models on file on day zero.”
Gorkem Yurtseven Oct 22, 2025 ▶ 22:22
Assertion Not checkable as stated
Major film studios actively seek generative AI video tools from fal
“Something changed this summer and we are getting a ton of interest from basically all the studios in LA or elsewhere. Everyone is really interested to at least do something about it because now they understand this is good enough and they can actually save mon…”
Gorkem Yurtseven Oct 22, 2025 ▶ 25:43
Insight
Frontier AI model advantages last only three to four months
“In the whole AI market, it looks like it's very hard to differentiate with models. If you have a good model, you maybe have a three month lead, maybe four months people catch up.”
Gorkem Yurtseven Oct 22, 2025 ▶ 29:06
Prediction Not checkable as stated
Model commoditization will drive fragmentation, favoring multi-model platforms like fal
“So it looks like it's going to be really hard for someone to differentiate with the quality of the model. And we believe this is going to lead to even, even more fragmentation. So a company like Fall where you get to access many different models at the same ti…”
Gorkem Yurtseven Oct 22, 2025 ▶ 29:48
Assertion Supported
fal hosts around 600 models compared to under 10 at AI labs
“So the number of models they have to host is like less than 10. And for us, it's like around 600, which is, complicates things like a lot.”
Gorkem Yurtseven Oct 22, 2025 ▶ 31:08
Disclosure
fal operates generative media inference across 28 different data centers
“That's part of the strategy we are running in, I think, 28 different data centers.”
Gorkem Yurtseven Oct 22, 2025 ▶ 32:04
Disclosure
fal maintains roughly 500 shared Slack channels with customer engineering teams
“We have, I don't know, 500 different Slack channels with all the engineers from companies we work with. And the response rate of those Slack channels, you measure that daily and like we obsess over that.”
Gorkem Yurtseven Oct 22, 2025 ▶ 35:30
Assertion Not checkable as stated
fal surpassed $100M ARR, scaling from $2M the previous summer
“Now it's over a hundred.”
Gorkem Yurtseven Oct 22, 2025 ▶ 38:40
Disclosure
fal converts pay-as-you-go AI usage into multi-million dollar annual commitments
“We built a sales team early on, maybe earlier than some of our competitors. And we tried to get as much of this revenue in form of yearly commitments rather than pay as you go. To this day, I think we are doing an incredible job at that. And that protects the …”
Gorkem Yurtseven Oct 22, 2025 ▶ 40:15
Insight
Low-level systems engineers can master GPU optimization without prior GPU experience
“If they were that database company before, or if they did any like low level systems engineering, that is a big plus, even if they haven't worked with a GPU before. We believe they can learn very fast”
Gorkem Yurtseven Oct 22, 2025 ▶ 43:47
Disclosure
fal reached $100M ARR with only 6 to 10 go-to-market staff
“We have around six, maybe, maybe 10 if you include CSM and like all that.”
Gorkem Yurtseven Oct 22, 2025 ▶ 46:17
Insight
Overwhelming inbound demand turns AI software sales into a qualification challenge
“With AI, you have so much demand coming from the market. You have to qualify, like your problems are very different. Your problem is you have to qualify who to spend time with. You have to qualify who is going to have the most spend among these companies for y…”
Gorkem Yurtseven Oct 22, 2025 ▶ 46:48
Opinion
Cursor is better suited for product engineering than low-level ML optimization
“Our product team uses cursor or equivalent tools a lot. Like I see the monthly bill and it keeps increasing. Yeah, I think it's better suited for product engineering type work as opposed to some of the low level optimizations we are doing. On the ML side.”
Gorkem Yurtseven Oct 22, 2025 ▶ 49:44
Disclosure
fal hired specific employees simply because they had active X profiles
“We hired couple people just because they had an active X and They ended up being like really active members of the community as well.”
Gorkem Yurtseven Oct 22, 2025 ▶ 51:59
Assertion Not checkable as stated
fal employs over 30 engineers with zero dedicated engineering managers
“Yeah, we don't have engineering managers. We have around like. 30 to 34 engineers. We do have leads. Obviously we have leaders in the In the team, but you know, we don't have this engineering manager role. Everyone is always, like, contributing, writing code. …”
Gorkem Yurtseven Oct 22, 2025 ▶ 54:33
Insight
Small mixed-group feedback discussions are more constructive than traditional 1-on-1s
“Instead of one-on-ones, we try to do like smaller groups of discussions, like one-on-three or whatever, one-on-four. And we try to bring like people from within the team, but maybe someone who joined recently, someone who's been there for a while, someone who'…”
Gorkem Yurtseven Oct 22, 2025 ▶ 55:14
Disclosure
fal hired six account executives before hiring a head of sales
“We built a sales team before we hired the head of sales. I think this is number one question, like series A or series B companies ask themselves what comes first. We decided to hire I think like six AEs first, everyone reported to either me or Burkai, and then…”
Gorkem Yurtseven Oct 22, 2025 ▶ 56:44
Made with StarZero

Turn any episode into a week of clips.

This entire site, about 80 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.