Fireworks AI

also referred to as: fireworks

12 statements across 5 episodes · 6 bullish · 2 bearish · 5 people on the record · first statement Mar 6, 2024 by Soumith Chintala · said 68 times in 20 episodes since 2023 · across every show →

Mentions by year

brought up most by Shawn Wang (44), Alessio Fanelli (5), Lin Qiao (3), Mikhail Parakhin (2), Beyang Liu (2), Sarah Sachs (1)

tap a year for its mentions
002545082023202420252026episodesmentions
0482023202420252026episodes it came up in
0044882023202420252026episodesmentions per episode
2026 7 mentions in 5 episodes 1 per episode
2025 11 mentions in 7 episodes 2 per episode
2024 48 mentions in 7 episodes 7 per episode
2023 2 mentions in 1 episode

every mention, scene by scene, with the transcript →

Everything said about Fireworks AI, oldest first

Mar 6, 2024 bearish
Opinion
Chintala: Inference Moats from Fast CUDA Kernels Last Only Months
“I think, like, Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster, like, you know, hardware kernels in general. But those modes only last for a month or two. Like, these ideas quickly propagate.”
Soumith Chintala Mar 6, 2024 ▶ 28:36 Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI
Oct 11, 2024 negative
Opinion
Goyal: Open-source inference providers are far less reliable than OpenAI
“They are nowhere near as reliable as, I mean, every single time I use any of those products and run a benchmark, I find a bug, text the CEO, and they fix something. It's nowhere near where OpenAI is.”
Ankur Goyal Oct 11, 2024 ▶ 1:37:24 Production AI Engineering starts with Evals
Nov 25, 2024 positive
Assertion Not checkable as stated
Fireworks AI runs custom acceleration kernels for almost all served models
“For almost for all models, for all large language models, all your models.”
Lin Qiao Nov 25, 2024 ▶ 23:24 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 25, 2024
Assertion Supported
Fireworks AI operates with a team of only forty people
“No, but only 40 people.”
Lin Qiao Nov 25, 2024 ▶ 39:12 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 25, 2024
Assertion Supported
Fireworks AI's multi-LoRA system serves up to 1,000 adapters per base model
“One base model can sustain a hundred to a thousand LoRa adapters. And then basically all these different LoRa adapters can share the same, like direct the same traffic to the same base model where base model is dominating the cost.”
Lin Qiao Nov 25, 2024 ▶ 53:45 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 25, 2024 positive
Assertion Partly supported
Fireworks AI serves fine-tuned LoRA adapters at base model pricing
“We wrote multi LoRa last year, actually, and we actually have this function for a long time and many people have been using it, but it's not well known that, oh, if you find your model, you don't need to use on demand. If you find your model is LoRa. You can u…”
Lin Qiao Nov 25, 2024 ▶ 52:11 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 25, 2024 positive
Disclosure
Fireworks AI deployed a custom workload optimization stack specifically for Cursor
“We have a unique automation stack that is one size fits one. We actually deploy to cursor early on. Basically optimize for their specific workload, and that's a lot of juice to extract out of there, and we see success in, in that product is actually can be wid…”
Lin Qiao Nov 25, 2024 ▶ 43:26 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 25, 2024 bullish
Prediction Not checkable as stated
Specialized open-source expert models will outperform one-size-fits-all closed-source models
“And that's our prediction is With specialization, there will be a lot of expert models, really, really good, and even better than, like, one size fits all open source closed source model.”
Lin Qiao Nov 25, 2024 ▶ 33:45 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 25, 2024 neutral
Insight
Customer inference workloads rarely align with foundation model training distributions
“The data distribution in their inference workload doesn't align with the data distribution in the training data for the model, right? It's a given, actually. If you think about this, because researchers have to guesstimate what is important, what's not importa…”
Lin Qiao Nov 25, 2024 ▶ 15:29 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Nov 25, 2024 positive
Disclosure
Fireworks AI will release a reasoning model inspired by OpenAI's o1
“So another announcement is we will also announce a, our next Declarative system is going to be appear as a model that has extremely high quality, and this model is inspired by O-one announcement from OpenAI. You should see that by the time we announce this o…”
Lin Qiao Nov 25, 2024 ▶ 29:26 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Jan 1, 2025 bullish
Prediction Not checkable as stated
Swyx: Diff mode will become the norm for AI code tools in 2025
“Canvas has incorporated the diff mode that both Anthropic and OpenAI and Fireworks has now shipped that I think is going to be the norm for next year, that everyone Need some kind of diff mode code interpreter thing.”
Shawn Wang Jan 1, 2025 ▶ 59:25 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents
Apr 15, 2026
Disclosure
Notion Fine-Tuned Custom Function-Calling Models Before Standard Tool Calling Existed
“Before function calling came out, we were trying to fine tune with the frontier labs and with fireworks, like a function calling model on notion functions.”
Sarah Sachs Apr 15, 2026 ▶ 3:21 Notion’s Sarah Sachs & Simon Last on Custom Agents, Evals, and the Future of Work
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.