Google Research
part of Google
3 statements across 1 episodes · 1 bullish · 1 bearish · 1 people on the record · first statement May 18, 2023 by Karan Singhal · across every show →
Everything said about Google Research, oldest first
May 18, 2023 neutral
Singhal: Medical AI still lacks grounded evaluations within specific clinical workflows
“One thing that has been missing from our work so far is really Grounded evaluations in a specific use case in a workflow to show that there is a benefit both in terms of safety in the short term and in terms of kind of long-term patient outcomes as well.”
May 18, 2023 negative
Singhal: Prior biomedical LLMs lacked systematic benchmarking and human evaluation
“There was a bit of a shortage of kind of a systematic way of doing evaluation of these models. And so it didn't feel like there was a systematic way to think about automated evaluation of the clinical knowledge of these models. So for example, via multiple cho…”
May 18, 2023 positive
Singhal: Evaluators can barely distinguish Med-PaLM 2 answers from human physicians
“One thing we're seeing with MedPOM-II as we get closer to physician-level performance on medical question answering is that it's hard to tell the difference anymore. It's hard to tell the difference between different models. It's hard to tell the difference be…”