Insight certainty 4/5 debate potential 2/5

Customer inference workloads rarely align with foundation model training distributions

Lin Qiao · Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI · Nov 25, 2024 · at 15:29

Lin Qiao, CEO of Fireworks AI, explains why a standard one-size-fits-all distributed inference engine leaves performance and cost improvements uncaptured.

0:00 / 0:23exact quote · 23.5s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The data distribution in their inference workload doesn't align with the data distribution in the training data for the model, right? It's a given, actually. If you think about this, because researchers have to guesstimate what is important, what's not important during, like, You prepared it for training. So because of that misalignment, then we leave a lot of quality, latency, cost improvement on the table.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Lin Qiao

Prediction Not checkable as stated
Specialized open-source expert models will outperform one-size-fits-all closed-source models
“And that's our prediction is With specialization, there will be a lot of expert models, really, really good, and even better than, like, one size fits all open source closed source model.”
Lin Qiao Nov 25, 2024 ▶ 33:45 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Insight
AI inference matters more than training because it scales with global population
“Our prediction is for those kind of applications, the inference is much more important than training. Because inference scale is proportional to the upliminal world population. And training. Training scale is proportional to the number of researchers.”
Lin Qiao Nov 25, 2024 ▶ 8:59 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Insight
Compelling GenAI applications require compound AI systems spanning multiple modalities
“In order to really build a compelling application on top of JNI, we need a compound AI system. Compact AI system basically is going to have multiple models across modalities along with APIs, whether it's public APIs, internal proprietary APIs, storage systems,…”
Lin Qiao Nov 25, 2024 ▶ 19:41 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Disclosure
Fireworks AI will release a reasoning model inspired by OpenAI's o1
“So another announcement is we will also announce a, our next Declarative system is going to be appear as a model that has extremely high quality, and this model is inspired by O-one announcement from OpenAI. You should see that by the time we announce this o…”
Lin Qiao Nov 25, 2024 ▶ 29:26 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Opinion
Pre-training on human data is hitting limits; synthetic data is required
“So I think on the data side, we're approaching the limit and the only data to increase that is synthetic generated data.”
Lin Qiao Nov 25, 2024 ▶ 35:05 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Assertion Partly supported
Fireworks AI serves fine-tuned LoRA adapters at base model pricing
“We wrote multi LoRa last year, actually, and we actually have this function for a long time and many people have been using it, but it's not well known that, oh, if you find your model, you don't need to use on demand. If you find your model is LoRa. You can u…”
Lin Qiao Nov 25, 2024 ▶ 52:11 Why Compound AI + Open Source will beat Closed AI — with Lin Qiao, CEO of Fireworks AI
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.