Off-the-shelf models eliminate data prep for all but the largest enterprises
“If there's a ready-made model that changes everything, this whole like data preparation stage. Can be skipped and only like the biggest of the companies are going to do that. Everyone else, they'll just use something off the shelf.”
Running two different products simultaneously creates conflicting messaging for sales
“It's very hard when you are not screaming exactly what you are doing to your customers, to the potential customers, to people you work with. It's really hard to sell because, you know, they look at your website, they see something else, like So it is really ha…”
fal struggled to raise Series A because VCs doubted image inference
“And we tried to explain this to people because it was so new, like no one got it. Everyone thought, An inference platform is an inference platform. Doesn't matter what kind of model it is. There are other people who are more qualified or more prepared to do th…”
Raising alongside competitors puts startups at disadvantage due to investor fatigue
“And that's something we underestimated how, how disadvantaged of a situation it is to raise at the same time with seemingly all your competitors, because they all say the same story. Investors hear it over and over again, and there's some Fatigue of hearing th…”
fal never monetized its viral real-time image-to-image inference demo
“Still, we didn't make any money from that demo. It makes a really impressive, you know, technical demo for people to see how fast we can run inference, but we couldn't find any use case for that, like very fast image to image inference. Still to this day, it i…”
fal built managed APIs to control code over arbitrary GPU orchestration
“Instead of Focusing on like GPU orchestration and letting people deploy whatever they want. We decided to build, you know, APIs. Every single code that's deployed is owned by us and we control the whole process.”
Optimizing common AI workflows creates more value than arbitrary custom deployments
“Everyone wants to do the same thing over and over again. And therefore we thought there is value in actually optimizing the most common workflow.”
Early indie developer customers on fal spent tens of thousands daily
“Everyone was spending serious money on the platform, tens of thousands of dollars a day, all of a sudden.”
Latency percentage gains matter substantially more on minute-long AI workloads
“If something takes 1:02, if you can shave off 20% of it, maybe not enough people care about it. But if something takes a minute and you can shave off a similar percentage, all of a sudden that's a lot more meaningful.”
fal maintains a 15-person applied ML team dedicated to model deployment
“We have a Applied ML team. It's around 15 people right now. And, you know, all they do every day is either deploy these models, optimize them, play with them, and they're obsessed with it.”
AI research labs actively approach fal for day-zero model releases
“And now with FAL's position in the market, we also get some early information from the research labs. Everyone like tries to talk to us. And release their models on file on day zero.”
Major film studios actively seek generative AI video tools from fal
“Something changed this summer and we are getting a ton of interest from basically all the studios in LA or elsewhere. Everyone is really interested to at least do something about it because now they understand this is good enough and they can actually save mon…”
fal hosts around 600 models compared to under 10 at AI labs
“So the number of models they have to host is like less than 10. And for us, it's like around 600, which is, complicates things like a lot.”
fal operates generative media inference across 28 different data centers
“That's part of the strategy we are running in, I think, 28 different data centers.”
fal maintains roughly 500 shared Slack channels with customer engineering teams
“We have, I don't know, 500 different Slack channels with all the engineers from companies we work with. And the response rate of those Slack channels, you measure that daily and like we obsess over that.”
fal surpassed $100M ARR, scaling from $2M the previous summer
“Now it's over a hundred.”
fal converts pay-as-you-go AI usage into multi-million dollar annual commitments
“We built a sales team early on, maybe earlier than some of our competitors. And we tried to get as much of this revenue in form of yearly commitments rather than pay as you go. To this day, I think we are doing an incredible job at that. And that protects the …”
Low-level systems engineers can master GPU optimization without prior GPU experience
“If they were that database company before, or if they did any like low level systems engineering, that is a big plus, even if they haven't worked with a GPU before. We believe they can learn very fast”
fal reached $100M ARR with only 6 to 10 go-to-market staff
“We have around six, maybe, maybe 10 if you include CSM and like all that.”
Overwhelming inbound demand turns AI software sales into a qualification challenge
“With AI, you have so much demand coming from the market. You have to qualify, like your problems are very different. Your problem is you have to qualify who to spend time with. You have to qualify who is going to have the most spend among these companies for y…”
Cursor is better suited for product engineering than low-level ML optimization
“Our product team uses cursor or equivalent tools a lot. Like I see the monthly bill and it keeps increasing. Yeah, I think it's better suited for product engineering type work as opposed to some of the low level optimizations we are doing. On the ML side.”
fal hired specific employees simply because they had active X profiles
“We hired couple people just because they had an active X and They ended up being like really active members of the community as well.”
fal employs over 30 engineers with zero dedicated engineering managers
“Yeah, we don't have engineering managers. We have around like. 30 to 34 engineers. We do have leads. Obviously we have leaders in the In the team, but you know, we don't have this engineering manager role. Everyone is always, like, contributing, writing code. …”
Small mixed-group feedback discussions are more constructive than traditional 1-on-1s
“Instead of one-on-ones, we try to do like smaller groups of discussions, like one-on-three or whatever, one-on-four. And we try to bring like people from within the team, but maybe someone who joined recently, someone who's been there for a while, someone who'…”
fal hired six account executives before hiring a head of sales
“We built a sales team before we hired the head of sales. I think this is number one question, like series A or series B companies ask themselves what comes first. We decided to hire I think like six AEs first, everyone reported to either me or Burkai, and then…”
Yurtseven: Fal.ai has around two million developers on its platform
“Yeah, we have around two million developers on the platform.”
Yurtseven confirms Fal.ai has surpassed $100M in revenue
“And you guys are over a hundred million in revenue, right? Just, this is not, you know, just developers kind of kicking the tires. That's correct. Yeah.”
Taskaya: SDXL generated Fal's first $1 million in revenue
“And then SDXL came, which was like the first major model. That brought, like, our first million in revenue if we consider that”
Taskaya: Flux grew Fal's revenue from $2M to $20M in two months
“In the first month of flux models, we reached from, like, two million to ten million in revenue. That was, like, a big jump. Next month, we were at 20, like, it just started going from there”
Gorkem: Hosting LLMs is a bad business due to Google search competition
“Language models, hosting language models is not a good business. At the time we thought, okay, we are going to be competing against OpenAI and Anthropic and all these labs. Turned, turned out that it was even worse because the killer application of language mo…”
Yurtseven: Long-term inference differentiation relies on constant hardware and architecture churn
“I think it's very hard to create this differentiation over long term if there is no new architectures, if there's no new chips, but luckily there is all the time.”
Yurtseven: Customer A/B testing proved higher image latency directly reduces engagement
“One of our customers actually did a very extensive A-B test of, like, they purposely slowed down latency on file to see how it impacts, you know, their metrics, and it had a huge part in it, and it's almost like page load time.”
Taskaya: Fal optimizes inference for four major AI video companies
“We have like four different companies four, four major video companies that we are doing this with and one image company that I don't think we disclosed.”
Fal CEO: Newer Video Models Are Now Much Better Than OpenAI's Sora
“Now we have video models that are much better than Sora.”
Yurtseven: Video models now account for over 50% of Fal's revenue
“That was February, so now, now it's probably over 80. No, 50%. 50? Yeah, okay. It's like over 50. Yeah, yeah, a hundred percent.”
Taskaya: NSFW content makes up less than 1% of Fal traffic
“Moderation is optional to a level where, like, illegal content is moderated, and we also track, like, the non-illegal content NSFW moderation, and, like, we haven't seen, like, we haven't seen more than one percent.”
Consumer Demand for Top AI Models Has Increased Inference Costs 100X
“Instead of models getting cheaper, yes, maybe the running the same model got cheaper. But people trained much bigger models that are much more expensive to run now, and people expect to use the best model. So running inference in general, maybe a hundred X in …”
Fal: Company avoids sales quotas and pays full OTE
“One thing now we are experimenting, maybe we can do shorter term quotas, meaning maybe quarterly or monthly quotas rather than yearly, whereas it's more predictable, you can course correct if something changes, but right now we are also not, not doing quotas. …”
Yurtseven: Fal Recruits Researchers via Open Compute Grant Projects
“Most of the people we hired for that research team is through our research grants program. So we, it's open invitation. You can basically just send us an email with a project you have in mind, and we care about, let's call it efficient AI. It's either efficien…”
Yurtseven: Generative media AI represents a distinct market from general AI models
“Everyone pretended, oh, all AI models is the same market, you know, same companies will be running all the models, but also we identified early that the buyers of these models are going to be very different. Therefore, this is going to be a completely differen…”
AI-Native Startups Are Outspending Large Enterprises on AI Infrastructure
“One really interesting thing that's happening is lots of more AI native companies and newer companies are actually spending more than like bigger enterprises who are maybe not super sure about putting things into production, but we want more of them.”
Fanelli: Image and video generation is much more open-source driven than language
“FAL especially has done amazing because media is still very open source driven when it comes to image and video generation. In a way, the language is not as much.”
Gur: AI video models are far from reaching marginal quality gains
“With video, we're earlier in the competition. There's still a lot of leapfrogging happening. There's just so much more to build, and there's just, like, you know, we haven't hit, like, a quality bar where there's just, like, marginal improvements.”
Gur: Model capability leaps trigger step-function adoption growth
“What we see internally is every time there's a big shift, In capabilities of models. The adoption and the use case is just, it's like a step function. It just keeps growing.”
Gur: Fal.ai's first 27 hires were all engineers
“Until we were like 28 people or so, we were all engineers, and I think like the 28th hire was a non engineer.”
Justine Moore: Creators choose Fal and Replicate over Google's $125 Veo plan
“What we see is, like, a ton of creators are just going to the model enablement layer, whether it's a consumer-facing interface, like CREA, or more of a developer-facing interface, like a FALL or a Replicate, where you can generate Like, videos on a one-off bas…”
Justine Moore: Veo 3 API pricing is around 75 cents per second
“Some of the more developer-oriented, like, API platforms, like Fall or Replicate, are offering generations where you pay per video. It's priced around, like, 75 seconds. 75 cents per second today.”
Mike: Research API docs first to avoid AI coding implementation errors
“Cause again, you can't always assume the AI knows everything. And this is where I know we just jumped straight into bolt, but if I wanted to build an actual application, I would do research first on what I need, how it would work, how it works, what the setup …”
Ruiz: tldraw's real-time generative AI demo updates in 32 milliseconds
“No, I think it's now, like, to 32 milliseconds, basically as you go.”