The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Lin Qiao argument clarity score 4.1/5 from 12 exchanges on raw tape · average scores: directness 3.9 · coherence 4.3 · precision 4 · compression 3.5 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
12exchanges match
12on raw tape
2redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And the GPUs themselves, they sit on some cloud or like, how does that, um, how does that work? Like, do you run on like all sorts of NVIDIAs and MDs and like, are you agnostic that way? Or like, how does, how does that work?

A Yeah, similar to PyTorch IDEA, we want to provide, we want to build on top of the best hardware possible, and the different hardware skill has different strengths, so we build on top of NVIDIA GPUs, um, from the, um, many, many generations to the most recent one, B-TU-Hundred. Uh, we also build on top of AMD's hardware, uh, and we optimize on top of ROCCM. Is there, um, CUDA equivalent layer? Um, so yeah, so we have been deployed to many, more than 15 regions globally, uh, currently on top of five different clouds. We, we plan to expand to possibly 10 different clouds and, uh, a lot more regions, uh, globally. Um, and in terms of deployment, we have our hosted, managed, fully managed single tenant API. Um, we can also deploy into your VPC, uh, and into your cloud, uh, for privacy, um, concerns. Or we can connect VPC to VPC through private link or VPC peering. So all these possible way of deployment we have enabled for various different enterprises and startups.

AI assessment note: “currently on top of five different clouds”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And then any thoughts on, uh, the future of, uh, the industry at large? Uh, like the next 12 to 18 months, is that the year of agent that everybody's talking about? Is that the year of adoption? What, what's your, uh, what's your prediction? What are your predictions?

A Uh, so there's no doubt. 2025 is the year of agents. Literally, no matter where you go, people are building all kind of agents. So, uh, I think there will be thousands of agents being built, focus on solving specific problems. Um, at the same time, I also think 2025 is the year of open models. And it's clear that Deep Seek has, has created a big dent there, but the meaning is not about Deep Seek itself by itself. It's the precedent it has set in the world about setting a huge, like a very solid baseline for open models, for open model providers, um, that it's just Need to be better before anyone can open source new models. Um, it also set a presence for all model providers to be better. Um, I'm very bullish on open models because I have seen the power of open source that PyTorch has been able to leverage. Um, DeepSeq, um, for example, just within one month of releasing their new models, There are, despite DeepSeq model, extremely hard to tune and optimize, extremely hard. There are 500, more than 500 variants published on Hugging Face, optimizing for, um, for local device, optimizing for cloud infrastructure. People tune and DeepSeq model, people distill all in DeepSeq model into various different kind of models. Companies like Perplexity also tune DeepSeq model Uh, for.

AI assessment note: “2025 is the year of agents. Literally, no matter where you go”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q From a business standpoint, there is this, uh, wide expectation across industry that the cost of inference needs to keep going down and will keep going down. What does that mean for you in terms of, uh, business as an inference infrastructure provider? Is there an expectation that you're going to be cheaper and cheaper and cheaper, and if so, how do you build your business over time?

A Right, so I believe, I believe this overall AI infrastructure would go down by order of magnitude in terms of cost. And, and it should. Where it is today, um, is, I don't think it's a stable state to be where it is today. And this whole trend of infrastructure should become a utility is really, really good for the industry. Uh, because today we see a lot of case that, um, I think today there, the infrastructure costs basically start to separate out two different concepts. One is product market fit. As in, is this product useful? Does this product add value? It is a different concern to, can you build a sustainable business? The fundamental reason is, even if people are, can't have a product benefit, they get really good feedback from their users. They cannot scale it to the maximum degree, because this is so expensive, and we literally have seen cases that a startup can scale to bankruptcy. And for, for big companies, for, for very big companies, uh, like, the GPU bill surfaced to their CEO and they're, like, asking questions, what the heck is going on here? Why are we burning so much money? And we should really think about ROI here. So because of that high cost of infrastructure, AI infrastructure for GNI, it basically creates a high bar of what kind of application can be a viable business. If this bar can be lowered by 10 times, you can imagine there's so many more, it will b…

AI assessment note: “I don't think I worry about my business. I think actually I want”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q And now you have this concept of foundation models that you can build on top of. Therefore, the whole complexity has moved to how you deploy the models, which is What's going to lead us into fireworks? Uh, but just, uh, just to, um, just to play it back, because it's such an important concept that, um, you know, people may or may not have internalized how that changes everything.

A That's exactly right. So, so that's why we see an explosion of adoption of this technology, uh, from the inception of our company. And, uh, there are AI-native startups. Um, they just, there's no box around what's possible, and they create Completely new user experience that we will never imagine, but also the enterprises, whether it's digital native or traditional enterprise, they are all joining, uh, this transformation and really think, uh, rethink their existing product experience or rethink how to power their internal productivity or customer service, uh, many different surface areas. So I always say across the board, what, It's most surprising to me, uh, from the beginning of when I started this company, I had this sequential go-to-market motion in my head, is to go to native startups first because they are most technical advanced, and then digital native because they are more, like, they're embracing, uh, new technology in open-minded way, or, and then traditional enterprise because they want to prove, I see enough proven point before enter. Now it's happening across the board. Everywhere. That's, that partly I will attribute to Like, Gen AI models are so easy to use and so accessible.

AI assessment note: “That's exactly right. So, so that's why we see an explosion of adoption”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Let's cover some of this quickly. This is fascinating. So what does that mean? Uh, the different ways of partitioning the model, you said, like what works?

A Right, right. Because we, we look at the model in, let's take a large language model, for example. Um, uh, some parts of the model, uh, is doing Prompt, uh, like processing the prompt, and some part of the model is predicting what is, uh, the possibility of next token, and, um, and, and when, and during runtime, um, they actually bottleneck by different things. You know, I'm talking about that in extremely simplified way, uh, is prompt processing is bottlenecked by computation, and, uh, generating next, predicting next token is bottlenecked by memory bandwidth. Right. So if you think about the model in one chunk, um, and scale them, uh, all together, and then you always get stuck in one bottleneck versus the other. So the way we, we, we have seen similar kind of infrastructure problem in other domains, like databases, um, and the way to kind of solve that problem is, um, is pull these, those Parts that bottleneck by different system resource to scale them completely differently. So, uh, so if it's compute bound, then we should just add more GPUs to it, right? Uh, if it's memory bottleneck bound, then there's different way to scale it. Uh, if it's computation bound, then we, like, we make computation more efficient. So by scaling those different parts of model, uh, execution, Um, independently, then we remove all possible bottlenecks, and that's the most efficient way to run tho…

AI assessment note: “scaling those different parts of model execution independently then we remove all possible bottlenecks”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q It seems that another important part of the platform is, uh, function calling. Do you want to talk about, um, Uh, what it is, how you do it, and why it's important?

A Right. Talking about that, we need to kind of talk about agent development. Um, so on Gen AI, I think the, the first wave of application development of Gen AI has been focusing on, um, building user experience for humans to consume. Because the nature of the content, the results being generated are human consumable content, right? So, um, and the model was trained to generate Um, optimized for human consumable content. So building agents is, uh, is more complicated than being direct human facing because those models need to speak to each other or connect with external APIs to enrich the results. So the fundamental need here is to make the model output feeding to some kind of programmable framework. That's a fundamental change in this new agent development. So first of all, we need to teach model to generate programmable output. It's called constraint output generation, and to follow certain schema, and usually it's JSON format for most of API integration, but we have seen many other customers, they want to generate XML, they want to generate HTML, those are all schematized formats, so we support a new mode called grammar mode, that is any BNF Format we can support. It's very generic. Um, so, so that empower developers to extend beyond JSON. That's the base foundation. And on top of that foundation, then there are two different ways to build, to like, satisfy agent development n…

AI assessment note: “the fundamental need here is to make the model output feeding to some kind of programmable framework.”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q So maybe to close, uh, some, uh, forward-looking, future-guessing, uh, kind of, uh, kind of thoughts. Uh, so one, uh, for Fireworks specifically, in the next 12 to 18 months, what's on the roadmap? What are you excited to build that you can talk about today? Uh, and then for the broader industry, any, uh, any thoughts for, you know, where things may be going from your perspective?

A Yeah, I'm super excited. Um, we're doubling down on the file optimizer. Uh, it's our customization engine across quality, speed, and cost. Um, we have made a tremendous amount of progress in the past one year, and, uh, we are getting that to massive production with many customers. They now have a viable path Path to have partner market fit, both partner market fit and durability. Um, and we are, we are gonna heavily focus a lot more for file optimizer on the quality side, because we're already doing really well on the speed and cost, and we want to maximize the quality gain that through customization, the quality should be much better than they build on top of any Generic model. Any generic model API. Because it shouldn't, right? So if you fuse, um, use case specific data to it, it should just shine. And, and for application default, I think, fundamentally, what is that mode, right? Their mode is probably not the user experience, but because it's very easy to copy. Anyone can study the product and copy. Their mode is data. It's the more, the more is the data they curate through, or the inside of the usage, like usage pattern, uh, and, and, and kind of create this flying wheel. And this flying wheel should feed back into the JNI model they use. And we want to kind of really double down to make customization heavily push on the next level quality. So I'm super excited about that …

AI assessment note: “we're doubling down on the file optimizer. Uh, it's our customization engine”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q sort of a low level and high level stuff. So, so it does anything from whatever the research scientists need to do in terms of, I don't know, like a back prop, whatever on one hand, but on the other hand, it's going to do, uh, GPU usage optimization. Is that fair? So it spans the whole, all the things that you need to do to develop faster and simpler.

A Right. So, uh, I think the API is, um, it, it kind of adopt Python as the programming language. That's why it's called PyTorch. Uh, the idea is from Torch to lower Torch time. Uh, and, uh, um, and, uh, and Python is a very simple, uh, easy to use program language for many researchers. Um, and they, uh, think about, uh, deep learning neural network, uh, in code. Right? You don't need to think about that as graphs, uh, as nodes and edges and how to tie them together. Uh, when this graph becomes, um, like 100,000 of nodes, then it's how you even kind of manage that thinking process. But represent that in code, in Python code, is very easy to manage. And then you, it, you become software engineer, uh, to think about a neural network. Uh, and it's very easy to debug because all Python tooling Um, the tool chains or can all work. Uh, so, uh, so yeah, so that, uh, actually unblock, uh, a lot of progress made in the modern innovation. Um, and it's also, um, of course it can run on GPU. It can run on different hardware skills. We have many different hardware vendor plugin, uh, to, uh, to a lower level to support PyTorch. And then that, Including TPUs. Uh, so yeah, so, um, and because like Google has a compiler that can directly connect with PyTorch and the lower, uh, lower the program into, uh, into kind of the TPU runtime. Uh, so, so that enable, uh, enable significant advancement, uh,…

AI assessment note: “Right. So, uh, I think the API is, um, it, it kind of adopt Python”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q great. Maybe a few words on your story, your background, in particular, uh, you were one of the key people, uh, on PyTorch, uh, at, uh, Meta, or Facebook at the time? Meta, yeah, Facebook at the time. I guess, what is PyTorch for anybody that doesn't follow those things, uh, you know, in great detail, and then, um, how did it all come about? What was your role there?

A Yeah, that, that was a, that was an amazing journey. When PyTorch happened, what was the kind of industry context here? So I joined Meta at, at that time, Meta, it was called Facebook at the time, was going through the transition of finishing mobile first to starting AI first. Like, there are fundamental reasons why it's sequenced that way, because after, um, Metatransition is application from desktop to mobile first, and it's widely accessible everywhere. Using your phone, you can connect with people, and that drive a lot of user engagement and a lot of data being created, and we all know data-fueled AI. So, so quickly after mobile first transition, we started AI first. At that time, there's no software, no hardware, there's no GPU, no people building all this infrastructure. So we bootstrapped the whole thing from ground up. Uh, and it was interesting time. Feels like right now, uh, there's so many companies building their own AI framework. Um, there's a gazillion number of AI frameworks, even within Mata, there are three different flavors. Uh, one for mobile, one for research, one for production. So, um, it's very confusing, um, and we decided to unify all of that and have one framework, um, Aiming to be the best tool for researchers to create innovative model architecture, but also one framework to deploy and deliver those innovative model into production to power all the p…

AI assessment note: “we decided to unify all of that and have one framework”

Partly raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q How did your experience at, uh, Meta and then PyTorch influence the idea for fireworks in the first place?

A I think one of the biggest success we saw from the PyTorch experience is simplicity scales. And this simplicity is user experience simplicity. Um, because we have a very strong engineering team, we can solve any infrastructure, uh, challenges and complexities. But when it comes to adoption, people don't want to spend time figuring out how to make things work. They just want it to work. And that's kind of where PyTorch shine. And that's why in the heavy competition or very noisy market of If every company is building their, was building their own AI framework, I think there were like tens of those AI frameworks at that time, but none of those were focusing on the simplicity part. Everyone was focused on production, um, and, uh, all the nitty gritty details to get into production, but to the point it is very hard for researchers to adopt, um, and a kind of interesting funnel we saw From the PyTorch journey is the researcher adoption is the top of the funnel because they create new model architecture which get open source and they get adopted by industry practitioners for experimentation. Maybe one or 10 of those experiments will be successful and then it goes down to the next phase of the funnel for production. Um, and once it goes to production, people start to transition from PyTorch into other production, more production focused to optimize the framework, but that transition i…

AI assessment note: “one of the biggest success we saw from the PyTorch experience is simplicity scales.”

Not addressed raw tape D 1 · C 4 · P 3 · Cm 3 2.70

Q Let's get into the weeds of the product, uh, itself. So, starting with models, uh, so you sort of offer power hundreds of models. What is the latest number of Models that you, um, that you offer?

A Yeah, so our product has multiple layers. Uh, at the lowest layer, we start, we have been working on distributed inference engine from the beginning. Um, and that's basically a customized PyTorch inference infrastructure we set up for JNS specifically. Um, and that distributed inference, uh, infrastructure was designed Very composable and building block based. The fundamental reason is we want to launch any state-of-the-art model very quickly, as, um, as we just discussed. We have the track record of, um, new open model get released in the same day we launch it, right? So, uh, that's because our, um, we, we kind of took the PyTorch idea, um, and built our inference infrastructure In a composable way. And then we can pick and choose what is the best options and setup for that particular model. And, uh, but the, the search space is very big. So we think about, first of all, our design principle is we don't believe in one size fits all. We don't believe in one size fits all for quality, as in, we don't believe one model will solve all the problems in GNI in the best way. That setup doesn't exist. We also don't believe there's one inference configuration that's best for, um, for all different use cases. Instead, we believe in one size fits one. We believe in a platform which can automatically do customization across three dimensions. As I mentioned, application developer, they usua…

AI assessment note: “Yeah, so our product has multiple layers. Uh, at the lowest layer”

Not addressed raw tape D 1 · C 3 · P 3 · Cm 2 2.25

Q That search problem, you solve it on a per-task basis, so sort of a priori before the, for the specific use case, or are you going toward, or maybe you do it today, towards real-time optimization?

A Right. So, uh, we have two modes, uh, similar to agentic development. I think agent is a very, very hot topic these days. Yeah, I'm hearing, yes. Everybody's talking about agents. Uh, and there are basically two extremes of, uh, thinking about agent development, right? So one extreme is autonomous, right? Uh, the goal here is to replace human, um, and generate The same quality or even better results than what human can do. And they are tapping into human resource budgets of, of different professions and so on. So that's one extreme. The other extreme is, uh, human assist and to make the profession, the professionals become much more productive, right? Uh, so similarly, when we think about, Uh, this optimization, we also have two modes. One is, uh, human in the loop. Uh, the second is more automated. Uh, so human in the loop, uh, that we have, uh, take DeepSeq for example. Uh, DeepSeq is very, very complicated model architecture because it has more than two 50 experts. Um, and DeepSeq actually, um, that company itself was running and still running Uh, this model over more than 300 GPUs. Uh, so think about this deployment. One replica is 300 GPUs, and there are so many different, so many more replicas. It's a very big distributed system, uh, problem to solve. So that's, that's why the, the, the model itself is very hard to tune. So we enabled, um, supervised fine tuning, uh, for …

AI assessment note: “when we think about, Uh, this optimization, we also have two modes.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.