The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Clement Delangue argument clarity score 4.3/5 from 14 exchanges on raw tape · average scores: directness 4.6 · coherence 4.6 · precision 4.1 · compression 3.8 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
14exchanges match
14on raw tape
1redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q And so you're fully distributed, but, uh, you have a large office in Paris these days. Is that, is that fair?

A Yes. So we're a hybrid company, uh, so people can work from anywhere in the world. Usually when there are a couple of people, three, four, five people, they take an office in the city that they are at. So I think we have maybe, uh, 1213, uh, small offices all over the world. Uh, the biggest one being Paris, where we have, I think around 20, 25% of the team based, based here. The three founders of Hugging Face are, are French. So we have quite, quite a lot of, uh, French roots, uh, but we consider ourselves like an international company. So not an American company, not a French company, but an international company with people from all over the world.

AI assessment note: “Yes. So we're a hybrid company, uh, so people can work from anywhere”

Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q When was that moment? Because having a, you know, a highly successful library is one thing, but then turning yourself into a platform to host a bunch of different libraries and models and all the things, when, when did that happen?

A It started quite, uh, quite early. Even in, kind of, like, the first days of, of Thomas releasing, kind of, like, the, uh, first port of, uh, BERT, I think, uh, contributors, open source contributors started to solve bugs, right, and kind of, like, improve some small things. Um, and then progressively as we added more models, more contributors started to, to add, uh, more, more models. Uh, so progressively the community, Contributed more and more, and we felt this, uh, this movement where the more we contributed to the community, uh, the more open source we, we did, uh, the more the community was giving us back. Um, so it, it validated us into, into this approach, uh, to today, as I mentioned, where we have, uh, five million AI builders using, using our platform, who collaboratively shared, uh, Uh, one million public models. Half of them have been downloaded in the past 30 days. Uh, so most of them are actually very useful and active for the community. They also contributed, I think it's, uh, more than 200,000 data sets to the platform. Um, so it's open data sets that anyone can go and use to fine tune, to customize their models for specific language, specific domains, Specific, ah, use cases. And collectively they built over 300,000 spaces which are the apps, ah, on, on the Hugging Face platform.

AI assessment note: “It started quite, uh, quite early. Even in, kind of, like, the first days”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q You've become one of the leading voices in open source, uh, AI. What, uh, would you say is the current state of open source AI?

A Three, four years ago, um, AI was extremely open and collaborative. Uh, all the research, uh, that was created was published. Most of the models were open sourced, right? So at the time, as you mentioned, It was GPT, GPT-II, there was, there was BERT, there was ExcelNet, um, obviously Google, uh, kicked off this wave with attention is all you need, which is like the seminal paper for, for Transformers, which is the T in, in ChatGPT. Um, over the following years, uh, the field had become, uh, became a bit more closed, A bit more commercial, commercially driven, right? Some of the big players started to, uh, share a little bit less of their research, uh, open source models a bit, a bit more, a bit less actually, uh, and put them behind, behind APIs. Um, and that's, that's where we are today. I think the field is much more, uh, closed, much less collaborative, Uh, much less science-driven than it, uh, than it used to be. Uh, and I regret it, um, because for me, uh, the progress of AI depends on being more open, more, more collaborative. As a matter of fact, I think we got where we are today thanks to this era of being, uh, open in terms of science and in terms of, like, uh, models. So I hope we can, uh, go back To fostering more open science and open source in the, in the future, I think it would be, uh, the right direction for, for the development of this technology.

AI assessment note: “I think the field is much more, uh, closed, much less collaborative”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Lama, um, as, you know, the, perhaps there is an altruistic aspect to it somewhere, but fundamentally, I guess it's a question of, um, you know, Meta not wanting platform dependency for AI the way that platform dependency Uh, for, for Apple, um, so do you think in a world where there's so much money at stake, uh, people have any kind of incentive to do open source these days?

A I think so, because I think, uh, compared to software, AI is much more science-driven, and I think if you look at science and how it's been done, uh, forever, it's more for altruistic motive, and, and it's been more in a way that researchers are, are publishing, sharing their research, and building on top of each other. Uh, so I think there is, there is this kind of like incentive that is very strong, uh, for AI. As I mentioned, there are very important, uh, business opportunities in open source AI. Uh, sometimes people simplistically, uh, oppose kind of like open source with, with monetization, but actually, as we've seen in software, there are really great open source companies, open source commercial companies that are very successful. And I would expect there to be even more in AI. Um, we'll have kind of like the, the MongoDB, the Elastic, the Red Hat of, of AI, which are going to be like open source AI, AI companies. So there are a lot of commercial opportunities with, with open source. Um, and finally, I think in AI, um, it's such a, an important technology that I believe the, the public I believe the policy makers, uh, will require it to, uh, be shared across as some sort of kind of like a fundamental infrastructure for, for all. So they might be in the future, hopefully more and more, um, policy incentives, maybe, uh, for more open source AI as a way to, uh, give access…

AI assessment note: “I think so, because I think, uh, compared to software, AI is much more science-driven”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q By the way, what is, uh, open source AI? Um, may have asked that at the beginning of, of this conversation about, uh, open AI, but is that, is that, is that open sourcing the model? Is that open sourcing the data set? There was like this whole controversy about weights versus not weights. What, what, where are we on all of that?

A So it's a gradient, right? Openness is a gradient from kind of like full open source AI, which would be considered kind of like a practice where you share the weights, you share the data sets that you use to train the model. You usually share the training scripts for reproducibility, for, to give everyone the ability to retrain the model themselves. That's kind of like the most extreme Uh, sides of the, uh, of the gradient. Um, and then, you know, like, uh, uh, the more you go towards, uh, more secrets, for example, you open source the weights, you don't open source the data sets, which in my opinion are still a good step, and a step in the right direction, and still kind of like an impactful thing to do. Um, depending on your company constraints, of course, like not everyone can share everything that they're doing. But in general, the more we can push towards the, uh, more open side of the gradients, uh, the better for, for the world because, uh, then you give more power for everyone to, to build with, with AI. You, uh, foster the progress of the field because everyone can learn from your mistake or for what works to, uh, uh, reproduce that. Um, and the more kind of, like, uh, transparent it is, so it helps with, like, biases. For example, you can see that some bias are contained in, in the dataset. It helps with education, because you can understand why the model is not answe…

AI assessment note: “Openness is a gradient from kind of like full open source AI”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What, what do you look for, uh, in, in general, what's interesting to you for, like, anybody listening to this that, uh, feels like working with Hugging Face would be a lot more fun than, uh, being by themselves. Like, what, what are you looking for?

A We do it, uh, quite similarly to how we hire in the sense that we don't really take kind of like a top-down strategic approach to it where we're like, okay, there are these topics, these topics, and we need to acquire companies in this topic. Instead, we're trying to find teams that share our culture of, uh, doing things, our company culture, uh, share our mission, Of democratizing AI, thanks to open source. Um, and then we assume that if we bring in, uh, teams like that, they'll manage to find their impact. Uh, and then we adapt our roadmap and our evolution based, based on them. Um, so what works for us is really, uh, teams that are, uh, impact action driven. And I like to work in a very decentralized approach. Uh, with like, uh, openness, uh, and a strong drive to contribute to, to the community. Um, and then we can acquire different stages. Uh, we, as I said, we're like lucky to be, to be profitable. Uh, we raised a bit less than five hundred million so far. Uh, most of it is still, uh, in the bank. Uh, so we have, uh, the chance to kind of like do a lot of like different operations. Depending on what we are excited about.

AI assessment note: “we're trying to find teams that share our culture of, uh, doing things”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Uh, was that a, like a one exception that's just not part of the model, or do you view yourself as developing some models, or maybe you do actively?

A Yeah, we do, uh, we do develop some models. Um, not so much as our core kind of like a business model than like some other companies, um, but more as a way to, uh, show what's possible, kind of like lead by example, uh, and create value for, for the community. So when we do, we release everything in, in open source, not only the, the models, but usually the data sets and the training scripts for it. Um, like we did with, uh, like we did with, with Bloom. Um, it's an interesting example because at the time, uh, there wasn't, uh, many open source large language models. Uh, people were saying that it was too dangerous to open source large language models. Um, and at the time, uh, there was kind of like, uh, several teams building Large language models. There was a big science with, with Bloom, and Meta also was, was doing a foray with something called OPT at the time, and actually some of the OPT researchers told us after the fact that the fact that we released Bloom in open source really helped them on their internal case to, uh, to release OPT. They were like, uh, internally saying if, uh, if Bloom Uh, could be open source. We could open source OPT, and they ended up actually open sourcing OPT, which in some ways kind of like, uh, created the open sourcing dynamics, uh, at, at Meta. So in a similar way, when we're seeing gaps in the, in the market, in a way, in the fields, for, …

AI assessment note: “Yeah, we do, uh, we do develop some models.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q uh, by the way, as a, as a, as a, you know, French-born person, but it must have been quite an experience to, like, uh, show up and testify in front of, uh, the U.S. Congress. Anything that you have seen people, you know, do or talk about from a policy standpoint that would be, uh, that you'd be in favor of, uh, uh, from an open source AI perspective?

A Yeah, this week is a good, uh, good example in terms of, uh, in terms of regulation if you take the specific case of California, um, because yesterday, uh, Governor Newsom, uh, vetoed and accepted two, uh, different bills. Um, one, uh, which was, um, really kind of like, uh, more creating, um, more constraints for Um, small tech, uh, and everyone to, to build AI and, and was really having the risk of, of concentrating power. Um, and another one that, um, that is more creating, I think, better incentive for, for AI, um, which actually mandates, uh, AI builders to share the data sets that their models have been trained on, which is an opportunity to, Create more inclusivity. Understand the bias, the limitations of, of your models. So, like on any technology topics, uh, there's going to be some, uh, good regulation, some, some bad regulation. In general, for me, I'm, uh, in favor of the regulations that are creating more competition, more opportunities for everyone to, to build with AI, and tackling some of the, um, end user, uh, risks of AI today, like the risk of misinformation, the risk of, of biases. Instead of focusing on more sci-fi driven long-term kind of like, uh, uh, hypothesis like, uh, the doomsday scenario where AI is gonna take control and, and kill us all. Uh, I think, uh, policymakers have made a lot of progress in the past few years in their understanding of, of A…

AI assessment note: “in favor of the regulations that are creating more competition, more opportunities for everyone”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Thank you. Very nicely done. So, um, Why do you think your, your open source project has been so successful? Um, so clearly there's the, the, the timing and NLP becoming a really important thing. Any other sort of like lessons or tricks that you may want to share with the audience?

A Yeah, I mean, I think in the current world, there's like a part missing between science and engineering, um, especially at, uh, big companies where you have these huge, uh, science departments that are doing really, really cool stuff, uh, and on the other side, you have these, uh, huge engineering departments that are doing really, really cool stuff, uh, but they don't talk to each other, right? Um, so we're filling this kind of, like, gap. In between, where, ah, we're releasing state-of-the-art models produced by the science, ah, but that are, ah, engineers, ah, friendly. Um, so I, I think that's, that's, that's the main gap, ah, that, ah, we're filling, ah, with obviously a strategy, ah, which is open research, open source by default, ah, that is setting us apart. From, uh, most of the people who are either like doing closed source or not real open source, which is more kind of like raising something and forgetting about it, uh, which is not what most of the companies, uh, need when they need to use open source.

AI assessment note: “we're filling this kind of, like, gap. In between, where, ah, we're releasing state-of-the-art models”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q And then speaking of modality, um, another interesting topic that, um, I've heard you talk about is large models versus smaller models and specialized models. Uh, what, what, what do you, what's your current thinking on this of, uh, which one matters more, which one lends itself to open source more?

A So we just crossed one million public models on Hugging Face. Uh, so it's a very big counterpoint to the one model to rule them all. Uh, I call it a fallacy. I think, um, I think big models, big generalist models are, uh, good for some use cases, especially when you have a generalist, uh, uh, use case. So if you're doing search, And you want your search, or if you're doing ChatGPT, and you want it to be able to talk about everything, a big model makes sense. Now, when you want to do kind of like, uh, let's say banking customer support chatbots, you don't really need it to tell you about the meaning of life, right? You can use a much smaller, more specialized model that is going to be cheaper to run, faster to run, easier to retrain, more controllable, that makes more sense, and so I think that's what Companies are, are realizing, uh, sometimes using Canflag the big models as the first experiment, first Canflag proof of concepts, and then when they want to go faster, save money, increase accuracy on their specific, uh, use case, then they switch to a smaller, more customized model. So I think, uh, we're gonna end up in a world where there are, uh, both Right? On both sides of the spectrum, some extremely large, powerful, generalist, costly models for some use cases, all the way down to a very specialized, simple, optimized models, and depending on, on your use case, you're gonna…

AI assessment note: “we're gonna end up in a world where there are, uh, both”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q 18, that, that general period of time. There were a number of companies that were sort of talking about, about doing this, and none of them really ever took off. And, you know, in some ways, GitHub never became the GitHub. Why do you think that is, and why do you think you were able to break through? Was it perfect timing? Was it Transformers? What, what were the reasons?

A Luck. I was, uh, I think it's, it's a big aspect in, in anyone's success. Sometimes we forget to, to talk about it, and we kind of like, uh, look back at things, and we're like, okay, the perfect strategy, the perfect approach to, to things. Um, the truth is that we, we got lucky in many ways. Uh, we got lucky in terms of, uh, timing, as, as you mentioned. I think we were at the right time, at the right place. Uh, we were lucky in terms of, like, uh, who we were as founders, um, compared to, uh, in alignment with what we were trying to, to build. We were, um, before we had some experience around consumer, and so I think it helped us to have a very community driven approach instead of maybe a more, like, enterprise driven approach. Uh, or more traditional kind of like B to B approach. And then we got lucky that, uh, yeah, the community adopted us and, and started to, to contribute to our platform and that we could collaborate with a lot of the, uh, products, uh, that were around at the time. Right from the beginning, actually, we took much more of a collaborative approach with, uh, with others and a competitive approach. And so it helped us kind of like, uh, to focus on where we were adding value in a way to the community, where we were building kind of like something useful for the community, and usually integrating with other offering and other technology products when it was …

AI assessment note: “The truth is that we, we got lucky in many ways.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q dot o being, uh, you know, the, the way that software has been built for many years and software two dot o being, Uh, building, uh, technology and software, uh, with AI, which is actually, as much as people are excited about AI, not necessarily a widely held view that all technology is going to be built with AI, but is that, is that, is that, is that your view?

A So AI is a terrible name for, for the current technology, because artificial intelligence is kind of like calling these concepts of, uh, independence, this concept of, like, uh, sentiency, sentience, the, this, this concept that, uh, you know, it's like a foreign kind of, like, uh, intelligence, obviously, calling from a lot of sci-fi around, around the topic. The way we see it is, Just kind of like the new paradigm to, to build old software in the way that before the way you were doing technology was by writing millions lines of code and kind of like, you know, really rule-based approach to building products. Now you do it differently. You start from the data, from the data sets, and you train a model, and you optimize a model, and then that's what kind of like, ah, um, kind of like, ah, is the foundation for, for your technology, your technology product. Um, and I think that's the right way to approach this, ah, this new technology, and, and realize that, you know, companies are building this technology, controlling this technology, And, and iterating on it to progressively do the current use cases of technology better, and also creating new capabilities. So, you know, we started by making search better with, uh, with, with AI, uh, and also we created new use cases like ChatGPT, which is like a new way, new way of conversing, uh, with a technology, technology system. And hope…

AI assessment note: “The way we see it is, Just kind of like the new paradigm to, to build old software”

Partly raw tape D 3 · C 4 · P 4 · Cm 4 3.70

Q things. So there is a nine dollars, 20 dollars, and then the enterprise tier. Maybe talk about the enterprise tier, which sounds like it's the bulk of the, the business. And at some point, um, you know, it sounded like you had a significant consulting business around enterprise customers, but it's Feels like it's evolved towards being more of that sort of collaboration software platform. Is that, is that fair?

A Yeah, I mean, the, the way we approach things is that, um, we're a platform for AI builders, right? Um, we think, uh, most of our usage and users are always going to stay open source and free, uh, and we need to build some sort of a freemium model. Right? Where a small percentage of the usage and the users are going to be paid and kind of, like, fund the rest of the, of the platform. And when you think about the, uh, difference between open source free and, and premium, uh, we identified three different kind of, like, variables. Um, it's like premium support and premium features. And usually these appeal the most to enterprise. We need kind of like, uh, advanced features and more enterprise features, especially around user management, around security and things like that. And then you have premium compute, premium infrastructure, which is very important for, for AI. And we do this with, uh, collaborations with the cloud providers, uh, and with our own cloud offering with, like, inference endpoints and, and spaces, GPU. So that's kind of like how we, how we approach things.

AI assessment note: “we identified three different kind of, like, variables. Um, it's like premium support”

Redirected raw tape D 3 · C 3 · P 3 · Cm 2 2.85

Q What a, what a, what a concept. You meaning you, you don't burn five million a year?

A Yeah, in AI, in AI, it's, uh, people don't understand sometimes. Like, you mean profitable just on the inference, right? Not on the training. Uh, meaning profitable based on the billions of dollars that you raised. Um, yeah, it, it's quite, uh, quite unusual for, for AI startups. Um, but I think we'll see it more and more in the current cycle of, uh, VC funding, too, where, uh, You think that's gonna dry up? I, I think, um, you know, as I was mentioning, there are a lot of investors who caught up and, and became obsessed about AI leading to some kind of, like, irrational decisions. Um, and that changed a little bit in the past year, uh, 18 months, maybe. So it's definitely, Uh, proved to be a bit, uh, harder, I think, for most startups, uh, in AI these days to raise than it used to be. So we think, for example, a lot of, uh, M&A activity, uh, these days, both for, kind of, like, big organizations, right, the, the depth inflection of, of the world, uh, but also for, for smaller organizations. For example, I Receive tons of inbounds these days of like early stage startups, uh, uh, wanting or getting excited about being acquired by Hugging Face. We actually acquired, acquired two in the past, in the past four months.

AI assessment note: “it's quite unusual for AI startups”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.