Assertion certainty 4/5 debate potential 2/5

Singhal: Prior biomedical LLMs lacked systematic benchmarking and human evaluation

Karan Singhal · No Priors Ep. 17 | With Karan Singhal · May 18, 2023 · at 12:35

Google Research's Karan Singhal explains why earlier scientific and biomedical language models suffered from inadequate clinical evaluation standards prior to Med-PaLM.

0:00 / 0:35exact quote · 35.9s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“There was a bit of a shortage of kind of a systematic way of doing evaluation of these models. And so it didn't feel like there was a systematic way to think about automated evaluation of the clinical knowledge of these models. So for example, via multiple choice benchmarks there were a few popular benchmarks like the MedQA benchmark. But, you know, it varied across paper what benchmarks they were studying. In some cases, we felt like these benchmarks were not high quality. And so that was one thing that we saw. I mean, another thing that we saw, which was more acute, I think, was kind of a lack of detailed human evaluation across many of these works.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Karan Singhal

Assertion Supported
Singhal: Evaluators can barely distinguish Med-PaLM 2 answers from human physicians
“One thing we're seeing with MedPOM-II as we get closer to physician-level performance on medical question answering is that it's hard to tell the difference anymore. It's hard to tell the difference between different models. It's hard to tell the difference be…”
Karan Singhal May 18, 2023 ▶ 30:09 No Priors Ep. 17 | With Karan Singhal
Assertion Supported
Singhal: Flan-PaLM was the first AI model to pass the USMLE
“When we took a variation of POM, the FlanPOM model, which was, again, work from Jason Wei and team you know, this is an instruction to a model that's been trained to follow instructions better. You know, again, it was able to perform quite well out of the box,…”
Karan Singhal May 18, 2023 ▶ 6:42 No Priors Ep. 17 | With Karan Singhal
Insight
Singhal: Fine-tuning outperforms prompt tuning when providing over 100 examples
“If you have three to five examples, let's say, then I would prompt it. If you have maybe 10 or 50 examples, it would either be prompt tuning or fine tuning. I think generally in that realm, prompt tuning and fine tuning perform similarly, and I would prefer pr…”
Karan Singhal May 18, 2023 ▶ 11:12 No Priors Ep. 17 | With Karan Singhal
Disclosure
Singhal: Medical AI still lacks grounded evaluations within specific clinical workflows
“One thing that has been missing from our work so far is really Grounded evaluations in a specific use case in a workflow to show that there is a benefit both in terms of safety in the short term and in terms of kind of long-term patient outcomes as well.”
Karan Singhal May 18, 2023 ▶ 16:00 No Priors Ep. 17 | With Karan Singhal
Prediction Not checkable as stated
Singhal: Specialized AI models will assist radiologists in the near term
“I think where there might be more of a need for specialized models Is when it comes down to kind of higher stakes workflows, and I think that might look in the short term more like a physician's assistant. And so imagine, for example, an agent that can work wi…”
Karan Singhal May 18, 2023 ▶ 20:44 No Priors Ep. 17 | With Karan Singhal
Assertion Supported
Singhal: Med-PaLM and Med-PaLM 2 were trained without patient health information
“Like for example, MedPOM and MedPOM-II are trained without any patient health information. They, they're just kind of taking all the knowledge of POM and POM-II and then just kind of Aligning them and making them behave in a certain way.”
Karan Singhal May 18, 2023 ▶ 25:58 No Priors Ep. 17 | With Karan Singhal
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.