Google Research software engineer Karan Singhal discusses the challenge of scalable oversight as Med-PaLM 2 approaches physician-level performance.
Assertion Supported
Singhal: Flan-PaLM was the first AI model to pass the USMLE
“When we took a variation of POM, the FlanPOM model, which was, again, work from Jason Wei and team you know, this is an instruction to a model that's been trained to follow instructions better. You know, again, it was able to perform quite well out of the box,…”
Insight
Singhal: Fine-tuning outperforms prompt tuning when providing over 100 examples
“If you have three to five examples, let's say, then I would prompt it. If you have maybe 10 or 50 examples, it would either be prompt tuning or fine tuning. I think generally in that realm, prompt tuning and fine tuning perform similarly, and I would prefer pr…”
Disclosure
Singhal: Medical AI still lacks grounded evaluations within specific clinical workflows
“One thing that has been missing from our work so far is really Grounded evaluations in a specific use case in a workflow to show that there is a benefit both in terms of safety in the short term and in terms of kind of long-term patient outcomes as well.”
Prediction Not checkable as stated
Singhal: Specialized AI models will assist radiologists in the near term
“I think where there might be more of a need for specialized models Is when it comes down to kind of higher stakes workflows, and I think that might look in the short term more like a physician's assistant. And so imagine, for example, an agent that can work wi…”
Assertion Supported
Singhal: Med-PaLM and Med-PaLM 2 were trained without patient health information
“Like for example, MedPOM and MedPOM-II are trained without any patient health information. They, they're just kind of taking all the knowledge of POM and POM-II and then just kind of Aligning them and making them behave in a certain way.”
Prediction Not checkable as stated
Singhal: Federated learning won't drive biggest near-term healthcare AI advances
“But I think there are like real world obstacles to doing federally learning on health data, which actually kind of increased activation energy to the point where in the next few years, I doubt that like the biggest advances are going to come. From using federa…”