The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Lin Qiao: AI future will feature millions of specialized models, not AGI
“I really believe the future will not be a few small number of AGI models dominating the world. I really believe the future will be, it may be scary, but I think that's true, it will be millions of specialized models, one per application, per use case.”
Specialized open-source expert models will outperform one-size-fits-all closed-source models
“And that's our prediction is With specialization, there will be a lot of expert models, really, really good, and even better than, like, one size fits all open source closed source model.”
Lin Qiao: AI token costs will drop 10x in three years, driving 100x usage
“I do think the cost of token will go down drastically. 10 X cost reduction in the next three years, and this 10 X cost reduction will drive a hundred X usage.”
Lin Qiao: The future of AI is private, specialized intelligence
“We believe the future of the frontier of the intelligence, actually private intelligence, our specialized intelligence.”
Lin Qiao: Fireworks runs distributed RL across 5-6 global data center regions
“We've designed a fully distributed system. We run across five, six data center regions globally, and tap into scattered GPUs, and they are able to run massive jobs, our jobs.”
Lin Qiao: All major AI coding companies use Fireworks infrastructure
“I think all major coding companies are on us.”
Qiao: Fireworks daily token count could grow 20x-100x by next year
“Anywhere ranging from 20 to a hundred X could be possible.”
Lin Qiao: Nations will build sovereign AI models like power grids
“I definitely see that possibility. I also see, if we think about the general intelligence model as the electricity layer, as a power line, every country should own their own power line, right?”
Qiao: AI market will shift from token maxing to ROI maxing
“In the next couple of years, as AI is getting more and more into production, there will be a lot of focus. In getting that clarity and getting that discipline out. The token maxing is just a thing in time, but we're quickly moving to ROI maxing, which is about…”
Qiao: Every company will own its custom AI intelligence within three years
“I really see people will own their, every single company will own their own intelligence as a must-have. It's not optional.”
Lin Qiao predicts a 10x AI cost reduction yields 100x more applications
“If this bar can be lowered by 10 times, you can imagine there's so many more, it will be hundred times more applications enter the, this arena to create a brand new experience to end consumers and prosumers. And by that, we'll see a much bigger consumption acr…”
Lin Qiao: The future of AI modeling belongs to open-source models
“The future of the future of modeling sits on open model side. And I believe that side is gonna be much more active in creating those hundreds or maybe thousands of expert models that is specialized delivering much better quality in certain domain.”
Lin Qiao: Enterprise data will never be shared with third parties
“It will never get shared with anyone else because this is company's proprietary IP.”
Lin Qiao: US can build its own open-source AI ecosystem if China restricts access
“I do believe, in terms of talent density and resources, I do believe U.S. Will be able to build that open system by ourselves, and we should.”
Lin Qiao: Fireworks processes over 40 trillion tokens daily
“Today we process more than 40 trillion tokens a day. So majority of those tokens are coming from a customized model, not from off-the-shelf models, are coming from a customized model.”
Qiao: Fireworks could optimize gross margins now but chooses rapid expansion
“If our focus is only optimized growth margin, we absolutely can do that, but we are sacrificing the speed of growth because we want to go everywhere.”
The AI Industry Lacks Systems Designed for 10-Trillion Parameter Models
“Great system designed for 10 trillion parameter models today.”
Lin Qiao: token costs will fall drastically and cheaper infrastructure will invite far more usage
“I do think the cost of token will go down drastically. High price will invite a lot of people coming in to solve the problem, and it will invite competition. Competition will bring down the cost, and then eventually will lead into a very economical solution, r…”
DeepSeek runs each single model replica across more than 300 GPUs
“DeepSeq actually that company itself was running and still running this model over more than 300 GPUs. So think about this deployment. One replica is 300 GPUs, and there are so many different, so many more replicas.”
Fireworks AI serves fine-tuned LoRA adapters at base model pricing
“We wrote multi LoRa last year, actually, and we actually have this function for a long time and many people have been using it, but it's not well known that, oh, if you find your model, you don't need to use on demand. If you find your model is LoRa. You can u…”
Fireworks AI's multi-LoRA system serves up to 1,000 adapters per base model
“One base model can sustain a hundred to a thousand LoRa adapters. And then basically all these different LoRa adapters can share the same, like direct the same traffic to the same base model where base model is dominating the cost.”
Lin Qiao: Fireworks AI will not enter the application layer
“We absolutely are not going to move into application layer. Very clear to us.”
Lin Qiao: Most world data is private enterprise data, not public internet
“If you think intelligence is a derivative of data, then majority of the data is actually not used for training a general intelligence model. The training data is coming from public internet and the label data. Public internet is very small. Corpus of data comp…”
Lin Qiao: Fireworks has achieved 5x to 10x inference cost reductions
“What we have seen in the past is five times to 10 times cost reduction.”
Lin Qiao: Almost all AI coding startups now fine-tune custom models
“In coding space, Cursor probably is one of the pioneers starting to tune their model, and now almost all coding companies tune their own models.”
Lin Qiao: Major AI base IQ leaps will occur every 9-12 months
“I see those as every year or every three quarters, there's a major leap.”
Qiao: AI model routing will evolve into fully automated self-learning systems
“We also think there's a space to build a automatic routing system that can learn by itself, and that compound with automatic tuning system eventually. We think it should all be automated, and then you can see a self-evolving system based on what flow through y…”
Fireworks achieves exact bit equivalence between AI training and inference
“Between the training system and the inference system, when models move over, We have bit equivalence. So, as in, the numerics are fully the same. We do not lose a bit of accuracy.”
Lin Qiao: Legal AI market consolidated from many startups to two
“Take legal, for example. I was on a dinner table, and interesting, it seems like there were a lot of those companies around two years ago, but now it's pretty much two.”
Fireworks AI improved speculative execution hit rates from 30% to 90%
“We have seen cases improving the prediction hit from 30% to 90%, and that's huge speed.”
Fireworks AI was first to enable function calling for DeepSeek models
“We have been working on function for calling for a long time, and we are the first one to enable function calling for deep seek models.”
Over 500 DeepSeek model variants hit Hugging Face within a month
“DeepSeq for example, just within one month of releasing their new models, There are, despite DeepSeq model, extremely hard to tune and optimize, extremely hard. There are 500, more than 500 variants published on Hugging Face, optimizing for local device, optim…”
PyTorch was originally built for researchers without considering production requirements
“PyTorch actually started as the framework for researchers. Don't care about production at all.”
Fireworks AI runs custom acceleration kernels for almost all served models
“For almost for all models, for all large language models, all your models.”
Lin Qiao: AI hardware vendors now release three SKUs per year
“And now within a year from one vendor alone, we have three SKUs.”
Meta Has Been Building Custom Silicon Chips for Over Five Years
“I think Meta has been building their chips for more than five years, way more than five years.”
Meta historically maintained three separate AI frameworks for mobile, research, and production
“Even within Mata, there are three different flavors. One for mobile, one for research, one for production.”
Meta spent five years rebuilding PyTorch's backend for internal scale
“It took us five years. Took us five years to get the stage supporting almost all internal needs using deep learning and mass and massive scale.”
Lin Qiao: OpenAI switched completely from TensorFlow to PyTorch
“OpenAI switched to use PyTorch fully.”
Lin Qiao: Meta had hundreds of engineers building PyTorch and its infrastructure
“We have hundreds of engineers building PyTorch and infrastructure around PyTorch, but at the same time, I believe PyTorch within Meta probably has thousands of users.”