Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. Can, so can you talk about what precisely it does different than traditional reasoning models?
A Like the current, um, reasoning thinking models, most of the time, at least I can talk from, from our research point of view, builds a single chain of thought, right? And then as you build a single chain of thought, and as the model continues to attend to its chain of thought, it builds a better understanding of what response it wants to give you. It can alternate between different hypotheses, reflect on what it has done before. Now, Of course, like one, if you think about it just also in a visual kind of space, one kind of scalability that you can bring onto the table is, can you have multiple parallel chains of thoughts so that you can, you can actually, um, analyze different hypotheses in parallel, and then you will have more capacity exploring different kinds of hypotheses, and then you can look at, you can compare those, and then you can eliminate the ones, Or you can, you can, you can continue pursuing, and you can sort of expand on particular ones. It's a very intuitive process in a way, but of course it is more involved.
AI assessment note: “can you have multiple parallel chains of thoughts so that you can”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q of Google, are struggling with what you get from when you make the models bigger. And so I just want to ask you about that. I mean, it seems like it's nice that there are all these techniques. But again, thinking about this one technique that was supposed to have limitless potential, is that a disappointment for the generative AI field overall, if that's not going to be the case?
A Yeah, um, I, I really, I really don't think about it that way because we have been able to, um, push the capabilities of the models quite effectively, right? I think in a way, um, the whole scale discussion starts from the scaling laws, right? Like scaling laws, uh, explain the performance of the models under both data and compute and number of parameters, right? And Like, researching all three in combination is the important thing, and when, when I look at the kind of, um, progress that we are getting from that general technology, I think it is, it is still improving. Um, what I, what I think is important is to make sure that there is a broad spectrum of research that is going on across the board, and, like, Rather than thinking about scaling only in one dimension, there's actually many different ways to think about it, and investing in those, and we can see the returns that I think across the field, really, not just, um, not just here at Google, but across the field, many different models are improving with quite significant steps, right? Um, so I think as a field, the progress Has been quite stellar. I think it's very exciting, and in Google we are very excited about the progress that we have been having with Gemini models. Going from 1.5 to two to 2.5, I think we had a very steady progress, very steady improvement in the capabilities of models, both in the spectrum of the c…
AI assessment note: “I really don't think about it that way because we have been able to”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q an example. The progress, this is something that everybody who comes on the show says the progress of going from like GPT three to GPT four was undeniable. GPT four to 4.5 less of a leap. So I want to ask you just in terms of the velocity of improvement, if that's the right way to put it. Are we coming back down to earth a little bit right now?
A Again, when I look at our model family, right, going from Gemini one to 1.5 to two to now to 2.5, I'm very excited about the pace that we have. When I look at the capabilities that we keep adding, right, like we have always designed Gemini models to be multimodal from the beginning, right? Like that was our ambition because we want to build AGI. We want to make sure that we have models that can fulfill the capabilities that we expect. From, from, from a general intelligence. So multimodality was key from the beginning, and we have been, as the versions have been progressing, we have been adding that natural multimodality more and more and more, and when, when I look at the pace of improvement in our reasoning capabilities, like lately we have added the thinking capabilities, and I think with 2.5 pro, um, we wanted to make a big leap in our reasoning capabilities, our coding capabilities, And I think one of the critical things is we are bringing all these together in one single model family, and that is actually one of the catalyzers of, of, of improvement and improvement at pace as well. It's harder, but we find that Creating a single model that can understand the world, and then you can ask questions about, oh, can you code me, um, this, this, this sort of, like, a simulation of a tree growing, and then it can do it.
AI assessment note: “going from Gemini one to 1.5 to two to now to 2.5, I'm very excited about the pace”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q now Google has a tremendous amount of compute. At your disposal. And so you basically have the option. Is it scale that you want to throw at these models? Or is it new techniques? So let me just ask it to you as plainly as I can, is scale the star right now? Or is it a supporting actor in terms of trying to get models to the next step?
A It's a, um, it's a good question. I think also the way you framed it, uh, because, uh, it is definitely an important, definitely an important factor, right? Like, because I think the way I like to think about this is it's rare that in any research problem you would have a dimension that Pretty confident it would give you improvements, right? Like with, of course, like with maybe diminishing returns, but most of the time with research, it's always like that. So like when we think about our research right now, in the case of generative AI models, right? Scale is definitely one of those, but it's one of those things that are equally important with other things. When we are thinking about our architectures, Like the architectural elements, the algorithms that we put in there that come, that make up the model, right? They are as important as the scale. We, of course, analyze and understand as with scale, how do these different architectures, different algorithms become more and more effective? That's an important part because you know that you are putting more computational capacity and like you want to make sure that you research the kinds of architectures and algorithms that pay off The best under that kind of scaling property, right? But as I said, that's not the only one. Data is really important. I think it is as critical as any other thing. The algorithms, architectures, modul…
AI assessment note: “it's one of those things that are equally important with other things”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q started your remarks today saying that the goal is AGI and there's progress that needs to happen before AGI, you just said. We had Jan Lacoon on the show a couple of weeks ago. You worked in Jan's lab. Jan emphatically stated, there is no way the AI industry is going to reach human level intelligence, which is his term for AGI, just by scaling up LLMs. Do you agree?
A Well, I mean, I think, um, that's a hypothesis, right? That might turn out to be true or not. But also, I don't think that there's any research lab that is trying to only do scaling up the LLM. So, like, I don't know if anyone is actually trying to negate that hypothesis or not. I mean, we are not. From, from my point of view, we are investing in such a broad spectrum of research. That I think that is what is necessary. And clearly I think like many of the researchers that I talked to and me myself, I think that, um, there is a lot more, um, critical elements that needs to be invented, right? So there is critical innovations on our path to AGI that we need to, uh, we need to get through. That's why we are still looking at this as a very ambitious research problem. And I think it is important to keep that kind of critical thinking in mind. With any research problem, you always try to look at multiple different hypotheses, try to look at many different solutions. A research problem this ambitious, like probably the most important problem that we are working in our lifetimes, right? It is the hardest problem, maybe we are working. As, um, as a problem, as a research, um, problem in our, in our, um, in our work. I think, like, um, like having that really ambitious research agenda and portfolio and making investments in many different directions is the important thing. From my point…
AI assessment note: “I don't think that there's any research lab that is trying to only do scaling”
Answered raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q you could argue. Uh, so this, this person wanted to know, and I think it's a really good question. Um, is there a coordination or, um, or, or possible between open source and Proprietary. I mean, we see OpenAI doing the, you know, their new open source model or teasing it, or should each sort of side try to get its own part of the market? What do you think?
A Um, I think, like, I want to say a couple of things, right? Like, um, First and foremost, again, like, take a step back. There's a lot of research that went into building this technology, right? Like, of course, like, in the last, like, um, two, three years, I think it became so accessible and so general that people are using in their daily lives, but there's a long history of research that built up to this point, right? So, like, as a research lab, Google, and, like, before, of course, like, there was DeepMind and Google Brain, two separate labs, That are working in tandem, um, in different aspects. And many of the technologies that we see today has been built as research prototypes, right? As research ideas and have been published in papers, as you said, transformers, the most critical technology that is underlying things. And then, uh, and then models like AlphaGo, right? AlphaFault, all of these kinds of things, all these research ideas have been evolving into building the knowledge space that we have right now. All that research, I think publications and, and, and, and open sourcing all those have been a critical element because we were, ah, we were, we were really in the exploratory space at those times. Nowadays, I think, like, the other thing that we always, like, need to remember is, actually we have, at Google, we have our Gemma models, right? That, that are, that are…
AI assessment note: “So I feel like it's not an either [...]”