The Exchanges, every show

Every argument clarity score on this site is built from rows on this page, here across all 44 shows. Each question and answer was assessed with names hidden, the hosts' own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

shows every show 44 of 44
every show
Omar Sanseviero no published score: a fair score needs 8 or more exchanges on raw tape on one show record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score rests on one show's raw tape, the show with the most assessed exchanges, and shrinks small samples toward that show's cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match on 44 shows
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q What goes into shipping a mainline model like this? Like, what's the behind the scenes?

A It's complex. The Gemma team is actually relatively small. We have, like, two or three PMs. We have one marketing person, and then there are, like, engineers and researchers working on shipping this. Of course, there's, like, default training part. How do we do the post-training, distillation, post-training techniques, and so on. What is quite exciting is that once we have the model, then we collaborate with a bunch of open source partners, right? So for example, we work with Lama CPP, Olama, MLX, Hogan Faces, BLM, NVIDIA, AMD. So we have almost 50 external partners for every, well, for the Gemma for launch, which has been the most complex launch. And also internally, we collaborate with a bunch of different teams. So think of Google Cloud, Vertex, Vertex Models as a Service, ADK, uh, and then Android as well, right? So we work, for example, with the Android team, and with the launch of Gemma IV, we released an integration with Android Studio. So in Android Studio, there is this agent mode where you can have a model helping you buy code and do things within Android Studio. And they should say integration with offline models using Lama CPP or BLM or any OpenAI compatible endpoint. So now you can use Gemma IV to also buy code Android applications in Android Studio.

AI assessment note: “The Gemma team is actually relatively small. We have, like, two or three PMs.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Do you see a future where, you know, small models get good enough? Like, does it cannibalize? It's an interesting position. You have big Gemini, you have Gemma, both get exponentially better over time. Like, current Gemma is much better than what we had closed source a few years ago.

A Yeah, for me, it's quite exciting. I mean, if you look at Gemma, you compare to how we were one year ago, I would say Gemma four is matching state of the art from one, one and a half years ago for most. Things. With local models or models that you can run in your own hardware, you can get capabilities, so you can get agent capabilities, function calling, system instructions, like conversational, and that kind of stuff. Knowledge is much trickier, so for knowledge, you do need a larger model, right? That's why if you compare Gemini to Gemma, Gemini has much better knowledge understanding of the world, right? Like facts, information, and so on. So it really depends. I do think We are heading towards a future in one, two years where imagine like you can run a Gemini three pro powerful model directly in your phone, right? And I think once we get there, things will be quite exciting, uh, from our product integration, from which experiences we can, uh, enable the users. Uh, I wouldn't say it cannibalizes, uh, it's still like two very different things. Like if you want that flagship capabilities, like this super complex, long running task, you would use Gemini if you need factuality and so on. But I don't think for many of these things, we'll get to a point in which we can do very powerful things directly on device.

AI assessment note: “I wouldn't say it cannibalizes, uh, it's still like two very different things.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Are people fine tuning? Outside of, you know, we see a few big companies do, okay, like Cursor has a really good consistent, there's a few that have done fine tuning, but it seems like it's not picking up as, you know.

A Yeah, so there was this period, 20, 24, in which there was like this, maybe 20, 23, like there were all of these fine tuning communities, and I think it's been changing quite a bit over the last two years, because models are getting very good out of the box. So as I was saying, like for Gemma four, we had 50 To 60 partners. And some of them were like, oh yeah, we're going to try and fine tune the 27 B model for this vision task. And they were like, oh, actually the model works too well out of the box. We don't need to fine tune it. Yeah. We saw lots of those things. So I'm seeing this excitement around fine tuning nowadays as general conversational models. Yeah. There is still quite a bit of excitement around fine tuning for specific domains like finance, healthcare, Specific types of data that the model didn't see, but as general conversational, like just changing how the model behaves , you can do most, most of that via prompting nowadays, and in terms of capabilities, the models are very good out of the box. So it's been changing quite a bit. There's still like the onslaught people. I don't know if you know, uh, Daniel Han.

AI assessment note: “it's been changing quite a bit over the last two years, because models are getting very good”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q What are some, so I didn't read that part, what are some insights on the tokenization?

A And this comes from Gemma III, like this has been done already for over a year, but the tokenizer is pretty much the same as Gemini, which means that the tokenizer lends itself to capture the right tokens for different languages. It's like a very good multilingual tokenizer, which means that if you compare Gemma III, so I'm going to the previous generation, if you compare Gemma III to other models from back then, maybe the other models were better than Gemma III like as general model, But if you train all of these models for, I don't know, a specific Southeast Asian language, I don't know, Vietnamese, let's say, Gemma would yield better results even if the other base models were potentially better.

AI assessment note: “the tokenizer is pretty much the same as Gemini, which means that the tokenizer lends itself”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q And when you're constrained or running on device, small efficient models, you guys did an offload, so you're like caring about efficiency. Um, but you know, do you see a world of multi LoRa as per task? Should people be fine tuning the small one?

A I think this is a big challenge in general in the whole developer ecosystem, because let's say that you want to have 20 apps in your phone, right? And let's say that each of those apps comes with its own LoRa, right? What happens when you update the base model? You also need to update all of these LoRa. So from a developer point of view, I think it will be very tricky because one, you don't want to have 20 different base models in the phone of the users. The battery will just die. You also don't want to have to update 20 LoRa every time you update the base model, right? So the release cycles in the Android world are, and in the iOS world, are very different. So yeah, I think it's more of a general industry challenge that we need to figure out how we think that people should build Email, like, on-device, a phone, a power, like, AI experiences.

AI assessment note: “I think it will be very tricky because one, you don't want to have”

Answered raw tape D 4 · C 3 · P 3 · Cm 3 3.30

Q Okay, we've got to wrap up soon. Uh, I just wanted to end a little bit on your, your growth in your team. Um, and, you know, Paige is here. Uh, Logan is over in SF. Uh, and you've been hiring all my friends, uh, Thor and Ivan and all these, um, what does the team look like? Where are you looking to grow?

A It's been quite exciting. We are hiring lots of very high agency people. I think maybe three, four years ago, we, we did a, like a nice interview about how I was growing, like, uh, at FoggingFace and how we were thinking, like, DevRel should look like. DevRel is also interesting. It's redefining what DevRel should be in an AI, very AI-centric organization at the frontier. It's a research lab at the end of the day, and we are in this AI era, so it's also rethinking what DevRel should look like in 20, 26. We are, yeah, we're having pretty much, like, high agency people excited to build things, to engage with the community, and so on. Right now, we are growing in Singapore, so we are looking to hire someone in Singapore.

AI assessment note: “Right now, we are growing in Singapore, so we are looking to hire someone”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.