Google Research engineer Karan Singhal discusses the data foundations of Med-PaLM and Med-PaLM 2 with Sarah Guo and Elad Gil.
Assertion Supported
Singhal: Evaluators can barely distinguish Med-PaLM 2 answers from human physicians
“One thing we're seeing with MedPOM-II as we get closer to physician-level performance on medical question answering is that it's hard to tell the difference anymore. It's hard to tell the difference between different models. It's hard to tell the difference be…”
Assertion Supported
Singhal: Flan-PaLM was the first AI model to pass the USMLE
“When we took a variation of POM, the FlanPOM model, which was, again, work from Jason Wei and team you know, this is an instruction to a model that's been trained to follow instructions better. You know, again, it was able to perform quite well out of the box,…”
Insight
Singhal: Fine-tuning outperforms prompt tuning when providing over 100 examples
“If you have three to five examples, let's say, then I would prompt it. If you have maybe 10 or 50 examples, it would either be prompt tuning or fine tuning. I think generally in that realm, prompt tuning and fine tuning perform similarly, and I would prefer pr…”
Disclosure
Singhal: Medical AI still lacks grounded evaluations within specific clinical workflows
“One thing that has been missing from our work so far is really Grounded evaluations in a specific use case in a workflow to show that there is a benefit both in terms of safety in the short term and in terms of kind of long-term patient outcomes as well.”
Prediction Not checkable as stated
Singhal: Specialized AI models will assist radiologists in the near term
“I think where there might be more of a need for specialized models Is when it comes down to kind of higher stakes workflows, and I think that might look in the short term more like a physician's assistant. And so imagine, for example, an agent that can work wi…”
Prediction Not checkable as stated
Singhal: Federated learning won't drive biggest near-term healthcare AI advances
“But I think there are like real world obstacles to doing federally learning on health data, which actually kind of increased activation energy to the point where in the next few years, I doubt that like the biggest advances are going to come. From using federa…”