Assertion Supported
Singhal: Evaluators can barely distinguish Med-PaLM 2 answers from human physicians
“One thing we're seeing with MedPOM-II as we get closer to physician-level performance on medical question answering is that it's hard to tell the difference anymore. It's hard to tell the difference between different models. It's hard to tell the difference be…”
Assertion Supported
Singhal: Flan-PaLM was the first AI model to pass the USMLE
“When we took a variation of POM, the FlanPOM model, which was, again, work from Jason Wei and team you know, this is an instruction to a model that's been trained to follow instructions better. You know, again, it was able to perform quite well out of the box,…”
Insight
Singhal: Fine-tuning outperforms prompt tuning when providing over 100 examples
“If you have three to five examples, let's say, then I would prompt it. If you have maybe 10 or 50 examples, it would either be prompt tuning or fine tuning. I think generally in that realm, prompt tuning and fine tuning perform similarly, and I would prefer pr…”
Disclosure
Singhal: Medical AI still lacks grounded evaluations within specific clinical workflows
“One thing that has been missing from our work so far is really Grounded evaluations in a specific use case in a workflow to show that there is a benefit both in terms of safety in the short term and in terms of kind of long-term patient outcomes as well.”
Prediction Not checkable as stated
Singhal: Specialized AI models will assist radiologists in the near term
“I think where there might be more of a need for specialized models Is when it comes down to kind of higher stakes workflows, and I think that might look in the short term more like a physician's assistant. And so imagine, for example, an agent that can work wi…”
Assertion Supported
Singhal: Med-PaLM and Med-PaLM 2 were trained without patient health information
“Like for example, MedPOM and MedPOM-II are trained without any patient health information. They, they're just kind of taking all the knowledge of POM and POM-II and then just kind of Aligning them and making them behave in a certain way.”
Prediction Not checkable as stated
Singhal: Federated learning won't drive biggest near-term healthcare AI advances
“But I think there are like real world obstacles to doing federally learning on health data, which actually kind of increased activation energy to the point where in the next few years, I doubt that like the biggest advances are going to come. From using federa…”
Opinion
Singhal: Medical AI is an ideal testbed for safety and alignment
“I think there's a good chance that this setting, this medical setting for example, medical question answering,
Or maybe more broadly, I think ends up being a better scenario to study concerns about technical safety and to mitigate concerns like misaligned with…”
Assertion Not checkable as stated
Singhal: Prior biomedical LLMs lacked systematic benchmarking and human evaluation
“There was a bit of a shortage of kind of a systematic way of doing evaluation of these models. And so it didn't feel like there was a systematic way to think about automated evaluation of the clinical knowledge of these models. So for example, via multiple cho…”
Prediction Held up
Singhal: Epic will likely partner with foundation models for clinical documentation
“I think that is also going to be something where players like Epic are going to be able to partner with existing models and I think potentially deliver real value there.”
Assertion Supported
Singhal: Google's 540B PaLM was the largest densely activated model in 2022
“The first Palm model was released in twenty-twenty-two, which was kind of this 540 B decoder only transformer model at the time, the largest densely activated model.”
Assertion Supported
Singhal: Roughly 10% of all internet searches seek health information
“Roughly 10% of searches on the internet are for health information.”
Prediction Not checkable as stated
Singhal: Augmenting telemedicine with LLMs is highly achievable within five years
“I think augmenting telemedicine, I think is, is, is kind of a short-term opportunity that I think in the next five years is, is very achievable.”
Assertion Not checkable as stated
Singhal: Google presents health info by attributing it to authoritative sources
“Where, for example, Google is doing that with health information is largely because it can attribute things to the Mayo Clinic and other organizations.”