The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Oriol Vinyals no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q I mean, I'm thinking of things like the older, like MedPalm II models and things like that, where they outperformed human physicians in terms of output. Relative to the physician expert panels. And then obviously you could then do some post training with physician experts as the, as the key, but at some point the machine will be better than that. And so how do you keep scaling reward functions?

A Traditionally, right? You, you just get it's supervised learning, right? Reward function means good or bad. So we can scale that process as much as we, we have so far. Um, I mean, obviously many, many players are, are realizing the power of Human annotation. And in fact, deep learning comes thanks to amazing like Fei Fei Li and lab effort to, to label a data set of a million examples, right? So, so that way of scaling is one. Uh, but then I strongly believe that there might be a bootstrapping effect of the models that become better at judging their own outputs, right? And so maybe, and that really is probably maybe even the main hypothesis of reinforcement learning as I see it. I mean, I'm not Huge expert in RL, but If checking that something is correct is easier than creating the solution, then we're in business because the language models will be able to evaluate their own samples more accurately than to generate them. And then we have a sort of reinforcement learning loop because we can reinforce the ones that seem more promising and then the model gets better, right? So that using the model itself as a reward, um, which Incidentally uses language, which is already fuzzy, is one area that, I mean, I'm excited about. There's a leaderboard of reward models. Um, some of them are, I think the name that they use is maybe generative reward model. I think that area goes beyond this…

AI assessment note: “a bootstrapping effect of the models that become better at judging their own outputs”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q games or any other constrained domain like this that, uh, you know, it's, it's, and, and tell me if you think this is like a really niche view, but that it's a dead end versus like general reasoning advancement. Um, even though it has all these attractive attributes, like the, you know, ability to generate and, uh, self validate in many ways, like how do you react to that criticism?

A There's a validity that, that again, going back to the reward function sort of question, right? You want to create the most general model or, or agent or intelligence. Um, turns out that reward is never perfectly defined by the environment. Um, even if you go on or surviving is, look, it is, it is extremely complicated to compute a reward. So, You could argue like, look, when you do get these rewards, um, from certainly somewhat artificial, you know, interesting, but artificial domains, that might not generalize to the real problem. And in that sense, it could be a dead end. Then how you do the research is important, right? Let's use math as an example, right? We do have access to the reward, but even there, it's not that simple, right? You know, if, if, I mean, if I'm doing a simple calculation, okay, yeah, four plus four equals eight. That I can check. Simple. But now you start thinking, well, prove this theorem. Um, the proof is either correct or not, but that starts to be more complex to get a crisp reward signal sometimes, unless you can formalize it, and then there might be bugs in the formalization process or in the, um, maybe in the engine that checks the math underneath. And even then, I mean, if I say four plus four, and instead of saying eight, I say four plus four equals eight. Is that correct or not, right? How do you check correctness? I mean, it's, you start to i…

AI assessment note: “There's a validity that... in that sense, it could be a dead end.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q chat bubble could pop up or something else, it could just happen organically based on the type of intent that the user is exhibiting. And so, um, I think you're raising really interesting points about the capabilities of, of Google and, um, and all the rest of it, I guess, related to that, what has been the most surprising thing about how users or companies have been interacting with Gemini?

A Yeah, there's quite a few, right? I guess, um, Maybe. The ones that surprised me the most, because initially I even thought this was just a number that you could report is, is the fact that, you know, infinite context length is coming eventually. Um, so I thought, look, I mean, this is interesting, right? We come from a world where we had recurrent neural networks and LSTMs that actually had infinite memory, although it was not very capable, right? You, the models in, in practice, they never remember more than a few hundred words or so. So that was kind of first that we could make the context length so long and then seeing the use cases just emerge even internally when we were first just trying the model that, I don't know, that, that seems very trivial now in hindsight, but, you know, putting a whole one hour video and just, just ask anything and it feels superhuman, right? You just literally put the video in and after 10 seconds, 30 seconds, I mean, it does take some time to process the context, but You just can't ask anything, right? And I mean, thinking of computer vision as a field or video question answering some of these data sets that we come from, I mean, they all seem very dwarf then compared to the capability that was in our hands. And then we put it in the hands of developers and we saw like, I mean, amazing, obviously demos and things that people could do, um, even…

AI assessment note: “putting a whole one hour video and just, just ask anything and it feels superhuman”

Partly raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q How do you interact with the rest of the company and like Google as a business? And I'm like, I feel like I have to ask you, does AI replace traditional search?

A So even running that from a research standpoint, um, is super interesting, right? There's, there's, um, two major centers, one in California, one in London, given the organizations that we come from. So that in itself is very interesting. In a way, we, we have the project running 24 seven, which is helpful when you train these large models. And then you have to do a few things, right? One of the things we do, of course, is trying to build state of the art technology, showing from sort of a research, knowing where the field is coming from and when it's going to, trying to really, um, showcase from our, our own sort of intuitions and ambition What might come next, right? So a prime example of this was, for example, the long context that we released earlier in the year, right? Millions, millions of, uh, tokens now are being able to be processed by, by our models. But then of course, we also, um, sort of take into consideration all the different needs, right? From the different products that we work with. Google has a lot of product areas. So we try to focus, of course, initially, especially to form the project, we try to focus on critical projects, and you see that very much, um, by how Gemini is first surfaced to, to users or to enterprises, right? So obviously cloud, um, and enterprise is very important. Developers as well. Uh, super cool to put these models in the hands of crea…

AI assessment note: “Google has a lot of product areas. So we try to focus”

Answered raw tape D 3 · C 3 · P 3 · Cm 2 2.85

Q Maybe one last one for you would just be, uh, do you, it sounds, I'm guessing no, because it sounds like the mission continues quite a bit beyond that, but do you live your life any differently, believing in twenty-twenty-eight?

A I think I was, I was reflecting on, on like cellular, like, I guess smartphones, right? So, so yes, with kids, like, how do you, how do they, do you present with the option of smartphones? And I mean, I, we don't have that many samples or data points, but then, I mean, obviously like that's obsolete, right? There's, there's these technologies. Um, I have, my kids are young, young, so I don't get to the, The kind of, oh, you can try like this Gemini, chat GPT, whatnot. But I think that Is worrying in the more human being at that sense. I think, I think you, you adapt as well, sort of how you do with like scaling up, right? You know, scaling up is very important as you progress in your career, you go from individual contributor writing codes to helping others sort of figure out what their, their, you know, their path is. So I think the scaling up, thanks to this technology is pretty, there's a huge opportunity. So I've been, of course, trying to, Use these to, you know, figure out what has happened in the endless chat rooms that I'm in. I'm in London. So, so when I wake up, I mean, California has like given me like a lot of tokens and long context to process. So, so I think there's, of course, personally, you also try to figure out how to best use the technology to scale yourself. And I talked to quite a few people that are not that much into technology. And all I say is, look, T…

AI assessment note: “personally, you also try to figure out how to best use the technology to scale yourself”

Partly raw tape D 2 · C 3 · P 3 · Cm 2 2.55

Q How do you contextualize this moment in time just in terms of, you know, what the biggest limitations are for current state of the art LLMs and like what's worth working on?

A I mean, there's one reflection that even many years ago with friends who were kind of early quote unquote in the game, they said, well, Get ready. Lots of brilliant people as this gets mainstream will enter the field. And you certainly see this, right? With open sourcing and a bit of a random search event, right? It's not like, I mean, you just selection bias, like someone does something random, but people actually want that. And then that becomes sort of viral in a way. So I think there's the, the sheer size of the field that is one aspect that I think we were sort of anticipating, but To me, that's one of the biggest changes that I've seen, that there's more brains, more different backgrounds coming into the field, and that is combined. I usually tend to assign credit, uh, what's with what has happened to, of course, the scale of data and compute algorithmic advances that you, you can simplify, but there's certainly been some that have been important in the last 1020 years or so. And then actually the accessibility, right? The software, the open sourcing efforts, um, those have been quite critical to then create these sort of exponentials or linear trends in log scale that we're seeing. Now, a bit more into sort of How I see the field from maybe, like, I tend to call maybe the 2000, let's say 10 to 20, like, deep learning era. So what that era did, right, is it took a set of …

AI assessment note: “I tend to call maybe the 2000, let's say 10 to 20, like, deep learning era.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.