The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Lin Qiao no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q model, or whether it's like a chain of models, and Gnome and basically everyone on the Strawberry team was very insistent that what they did for reinforcement learning, train of thought, cannot be replicated by a whole bunch of open source model calls. Do you think that that is, they're wrong? Have you done the same amount of work on RL as they have, or was it a different direction?

A I think they take a very specific approach where I do, the caliber of team is very high, right? So I do think they are the domain expert in doing the things they are doing, but I, I don't think there's only one way to achieve the same goal. We are on the same direction in the sense that the quality scaling law is shifting from training to inference. We are definitely honest. For that, I fully agree with them, but we are taking a completely different approach to the problem. All of that is because, of course, we didn't train the model from scratch. All of that is because we built on the show of giants, right? So the current model available we have access to is getting better and better. The future trend is the gap between the open source model, the open source model, it's just going to shrink to the point there's not much difference. And then we are on the same level field. That's why I kind of, I think, our early investment in inference and all the work we do around balancing across quality, latency, and cost pay off because we have accumulated a lot of experience there, and that empowers us to, to build and release this new model that is approaching OpenAI's quality.

AI assessment note: “I fully agree with them, but we are taking a completely different approach to the”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q think about, you know, people are like, hey, Lama, 2.2 is X on MMLU, but maybe, you know, using speculative decoding, you go down a different path. Maybe some providers run a quantized model. How should people think about How much they should care about how you're actually running the model? You know, like what's the delta between all the magic that you do and like what a raw model.

A Okay. So there are two big development cycle. One is experimentation where they need fast situation. They don't want to think about quality and they just kind of want to experiment with product experience and so on. Right. So that's one. And then it looks good and they want to kind of postpartum market by scaling. And the quality is really important and latency and all the other things are becoming important. During the experimentation phase is just pick a good model. Don't worry about anything else. Make sure even like JNI is the right solution to your product. And that's the focus. And then postmodern market fit, then that's kind of the three dimensional optimization curve start to kick in. Across quality, latency, cost, where you should land. And to me, it's a purely a product decision. To many products, if you choose a lower quality, but better speed and lower cost, but it doesn't make a difference to the product experience, then you should do it. So that's why I think inference is part of the validation. The validation doesn't stop at offline well. The validation is kind of, will go through A-B testing through inference. And that's why we kind of offer various different configurations for you to test which is the best setting. So this is the, like, traditional product evaluation. So product evaluation should also include your new model versions, um, and different model set…

AI assessment note: “to me, it's a purely a product decision. To many products, if you choose”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q much about PyTorch at all. You know, they're just trying to go to my model from a product perspective, like what were some of the decisions early on? Like right in October, November, you were just like, Hey, Most people just care about the model, not about the framework. We're going to make it super easy, or was it more a gradual transition to the model library you have today?

A Yeah. So our product decision all based on who is our ICP. And, uh, one thing we want to acknowledge here is the generic technology is disruptive. It's very different from AI before GNI. So it's a clear Leap forward. Because before Gen AI, the companies that want to invest in AI, they have to train from scratch. There's no other way. There's no foundation model. It doesn't exist. So that means they need to start a team. First hire a team who is capable of crunch data. There's a lot of data to crunch, right? Because training from scratch, you know, you have to prepare a lot of data. And then they need to have, uh, then you have GPUs to train, uh, and then you start to manage GPUs. So then it becomes a very complex, Complex project. Uh, it takes a long time and not, not many company can afford it actually. And the JNI is a very different game right now because it is a foundation model. So you don't have to train anymore that make AI much more accessible as a technology as an app developer or product manager, even not a developer, they can interact with JNI models directly. So in our goal is make it accessible to all App developers and product engineers. That's our goal. So then getting them into the building model doesn't make any sense anymore with this new technology and then building easy, accessible API is the most important. Our, uh, early on when we got started, we decided …

AI assessment note: “early on when we got started, we decided we're going to be OpenAI compatible”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q What do people not know about the work that you do? I guess like people are like, okay, fireworks, you run model very quickly. You have the functional model. Is there any kind of like underrated part of fireworks that more people should try?

A Yeah, actually, one user posted on x.com, he mentioned, oh, actually, firewalls can allow me to upload the LoRa adapter to the service model. With at the same cost and use it at the same cost. Nobody has provided that. That's because we have a very special, like we, we wrote multi LoRa last year, actually, and we actually have this function for a long time and many people have been using it, but it's not well known that, oh, if you find your model, you don't need to use on demand. If you find your model is LoRa. You can upload your LAR adapter and redeploy it as if it's a new model. And then you use, you get your, uh, endpoint and you can use that directly. But at the same cost as the base model. So I'm happy that user is marketing it for us. He discovered that feature, but we have that for last year. Uh, so I think we have, I think to, um, feedback to me is we have a lot of very, very good features as, um, Sean just mentioned.

AI assessment note: “many people have been using it, but it's not well known”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Cursor. We were the first podcast to have Cursor on, and then obviously since then they have blown up. Cause and effect are not related. But you guys especially worked on a fast apply model Where you were one of the first people to, to work on a speculative decoding in, in a production setting. Maybe just talk about like, what was the behind the scenes of working with Cursor?

A Right. I will say Cursor is a very, very unique team. I think the unique part is the team has very high technical caliber, right? There's no question about it. But they have decided, although like they, many companies building coding, they will say, I'm going to build a whole entire stack because I can. And they are unique in the sense They seek partnership, not because they cannot, they're fully capable, but they know where to focus. That to me is amazing. And, uh, of course they want to find a best partner. So we spent some time working together. They are pushing us very aggressively, uh, because for them to deliver high caliber product experience, they need the latency. They need the interactive, but also high quality at the same time. So actually we expanded our product future quite a lot. As we supporting Cursor, and they are growing so fast, and we massively scaled quickly across multiple regions, and we developed Pretty high intense inference stack, almost like similar to what we do for Meta. That, I think that's a very, very, uh, interesting engagement. And through that, there are a lot of trust being built as in they realize, Hey, this is a team they can really partner with and they can go big with. That comes back to, Hey, we are really customer obsessed and all the engineers working with them. There's just enormous amount of time syncing together with them. And discu…

AI assessment note: “we massively scaled quickly across multiple regions, and we developed Pretty high intense inference stack”

Redirected raw tape D 3 · C 3 · P 3 · Cm 2 2.85

Q I think the, the, the typical challenge for people is understanding, like, that has value, uh, and then, like, there are other people who are also offering open source models, right? Like, your moat is, is your ability to offer, like, a good experience for all these customers, but if your existence is entirely reliant on people releasing nice open source models, other people can also do the same thing.

A Right, yeah. So I would say we build on top of open source model foundation. So that, that's the kind of foundation we build on top of. But we look at our, the value prop from the lens of application developers and product engineers. So they want to create new UX. So what's happening in the industry right now is people are thinking about completely new way of designing products. And I'm talking to so many founders. It's just mind blowing. They help me understand existing way of doing PowerPoint, existing way of coding, existing way of managing customer service. It's actually putting a box in our head. For example, PowerPoint, right? So PowerPoint generation is we always need to think about how to fit into my storytelling into this format of slide one after another. And I, I'm going to juggle through like design together with, you know, what story to tell, but the most important thing is what's your story telling lies, right? And why don't we create a space that is not limited to any format and those kinds of new product UX design combined with automated content generation through JNI is the new thing that many founders are doing. What are the challenges they're facing? Let's go from there. One is, again, because a lot of products built on top of JNI, they are consumer, consumer, developer facing, and they require interactive experience. It's just a kind of product experience we…

AI assessment note: “we look at our, the value prop from the lens of application developers”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.