Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q What, what else has been the most interesting in your, your work at Google so far?
A Yeah, I started working out, um, in representation learning and federated learning. Uh, so this is kind of the technology, representation learning in particular is kind of the technology underlying a lot of the, you know, deep neural networks of today, including GPT-III, GPT-IV, and so on. And so this is largely about, um, learning, uh, representations of text, of images, of other modalities, such that you can efficiently encode them, you can learn from them in the future, you can generalize the new, uh, text and images and so on. And so work for this really started, you know, back in the beginnings of the deep learning era, like in 2013 with convolutional neural networks and scaling those up and work to VEC around 20 15 and glove and all these things. And I think since then, you know, we've been working on technologies around self supervised learning, um, around, you know, doing that in a privacy preserving way. And so, you know, after a couple of years of working on that at Google, I had the opportunity to kind of quickly grow and start to lead a team. Um, I kind of got to the point where I was thinking like, okay, I've upscaled in a lot of ways. I've, I've, I've gotten to the point where I can mentor many other researchers in, in a lot of ways. And now it's a great time to be thinking about my next thing and, you know, going for something ambitious in terms of shaping the tr…
AI assessment note: “I started working out, um, in representation learning and federated learning.”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q We're starting to get really good performance here. And so Can we do something that's in the medical domain specifically? And that was MedPalm. And so how did you do that alignment that you mentioned? Was it some form of RLHF? Was it some other form of fine tuning? Was it how you train the model to begin with? Like, what was the difference in terms of MedPalm versus Palm?
A Yeah, absolutely. I mean, when, when we tried evaluating POM in the medical setting, we noticed that our box on multiple choice questions performing pretty well, and when we took a variation of POM, the FlanPOM model, which was, again, work from Jason Wei and team, um, you know, this is an instruction to a model, a model that's been trained to follow instructions better. Um, you know, again, it was able to perform quite well out of the box, and this was the first model that was able to perform Above the pass mark on the med QA set of USMLE style exam, uh, questions. But then what we noticed is that when we evaluated on long form medical question answering, like actually getting the model to generate response, there was a lot of limitations. And when we compare, compared that to clinician performance, it actually didn't do super well. And so really that was the motivation for that med POM, uh, specific alignment. And so what we did there was really thinking about instruction prompt tuning, which was this technique Which we explored in that, in that MedPom paper, which is kind of a data efficient technique, and a technique that doesn't require too much data to work because, you know, getting labels from doctors is expensive, which took a bunch of expert demonstrations of good behavior from doctors, and then used that to tune the parameters of the model, and do that in a way that'…
AI assessment note: “what we did there was really thinking about instruction prompt tuning”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q is, you know, more clinical decision-making around who gets a certain pharmaceutical. Do you view this as a technology that's initially a physician's assistant? Do you view it as something that helps with adjudication of medical claims and billing? Like there's so many places where this can sort of insert I'm just sort of curious, like, you know, where do you think you'll, you'll see this technology popping up first?
A Yeah, I think we're already starting to see it in some clinical workflows when it comes to documentation and building. I think there, there are a lot of, um, companies and people thinking about Taking models like GPT-IV and applying them in that setting. And I think that, that is definitely going to be something. I think that is also going to be something where players like Epic are going to be able to partner with existing models and I think potentially deliver real value there. Um, and I think, I think that's, that's, that's very exciting. I think that's something that also, you know, general domain models will be potentially quite good at as well. Um, I think where there might be more of a need for specialized models Is when it comes down to, um, kind of higher stakes workflows, and I think that might look in the short term more like a physician's assistant. And so imagine, for example, an agent that can work with a radiologist, help them interpret a scan, and leverage the benefits of AI to kind of help contextualize, um, you know, what a patient's medical record or any previous scans or different angles of scans that, that a patient has had, uh, to help a radiologist write a more accurate report. I think that's something, that's the kind of thing which I think, you know, is in the sweet spot of, of both feasible today, you know, leverages the benefits of AI in terms of taki…
AI assessment note: “we're already starting to see it in some clinical workflows when it comes to documentation”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q mission critical AI applications in the past. Maybe on that note, like one last ask for you in terms of encouraging some optimism, you know, you're working on the state of the art in this field and thinking about the, um, the barriers to, uh, the applied use, like five years from now, like how do you hope we are using large language models in the, in the medical field?
A Yeah, I guess, I guess I think about this in, in two broad buckets. I think there, there are two broad types of things that we can do for large language models in the medical field. I think the first is increasing the standard of care very broadly. And so that looks a lot like, you know, increasing access to health information, uh, providing assistance to physicians. So the radiology example I gave earlier, um, potentially clinical decision support, like double checking a doctor's decision or quality assurance for a radiologist report. So if, if, you know, a radiologist is dictating a report, they say no plural effusion scene, but then it's written down as plural effusion scene, then maybe an AI double checks that and just make sure that that's, that's what was intended. I think augmenting telemedicine, I think is, is, is kind of a short-term opportunity that I think in the next five years is, is very achievable. I think the other big bucket of things that is very much achievable, um, is augmenting scientific workflows. And I think this could be a longer term thing than five years, but I think there's also short term things that we can do as well. So thinking about looking at, you know, correlations across modalities and existing data to find novel biomarkers for existing diseases that we know about, or, um, kind of using large language models as research assistants. Um, so I t…
AI assessment note: “I think there are two broad types of things that we can do”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah, going back to perhaps more promising near-term areas of research, you've had this idea of building a medical assistant as a sort of laboratory for safety and alignment research. Can you talk about that?
A Yeah, absolutely. I mean, this is a lot of what got me thinking about the setting, especially coming into the setting as somebody who, you know, didn't have much of a medical background in terms of expertise. I was really thinking about, you know, what are the, what are the big things that I could do to help shape the trajectory of AI or nudge it in a more beneficial direction? And thinking about, um, AI safety, um, seriously in terms of both short-term and longer-term risks, I think was important to me. And so, you know, one thing I've, I've become more convinced of, uh, about over time is this idea that, you know, many organizations right now, um, Google, DeepMind, uh, Anthropic, OpenAI are right now looking at the idea of a general chat assistant and kind of instead of like doing alignment research in a vacuum, are looking at that setting as a way in which we can think about kind of better refining these models and better aligning them to human values. I think there's a good chance that this setting, this medical setting, uh, for example, medical question answering, Or maybe more broadly, I think ends up being a better scenario to study, uh, concerns about technical safety and to mitigate concerns like, um, misaligned with human values or hallucinations or things like that. And so, I mean, I think this comes down to things like making sure the incentives are aligned with res…
AI assessment note: “this medical setting... ends up being a better scenario to study, uh, concerns about technical safety”
Answered raw tape
D 4 · C 5 · P 4 · Cm 3 4.15
Q So I knew he was looking at my kid's symptoms. He had no idea. Right. And so there's this bar from, Hey, it needs to be incredibly accurate and correct on through to, well, the state of the art actually isn't that amazing in many circumstances. And so how do you think about the right quality bar for these sorts of things in terms of real use application or practice?
A That's, that's an amazing, great question. I think, I think, as you said, there's two competing forces here, right? Obviously, the stakes are high in the medical setting, and counterfactually, you want to make sure that the information you provide versus the information they would have otherwise gotten is actually high quality, and so that's, you know, very, very careful as you think about, you know, any informational use case for these models. At the same time, I, I think it's useful to recognize that people are searching, uh, for health information online, and Indecision is a decision as well. And so, you know, a large percentage, roughly 10% of, of searches on the internet are for health information. And some of these are coming from physicians themselves, as you, as you mentioned a lot. And so, you know, I think there is a responsibility to think about how to shepherd this technology carefully and safely towards, um, that real world impact for patient health information. And I think that, you know, is crucial as well. And I think one thing that has been missing from our work so far is really Grounded evaluations in a specific use case in a workflow to show that there is a benefit, uh, both in terms of safety, um, in the short term and in terms of kind of long-term patient outcomes as well. And so, you know, I think that could be a health informational use case. It could be …
AI assessment note: “Grounded evaluations in a specific use case in a workflow to show that there is a benefit”