May 18, 2023 · 42m · no-priors

No Priors Ep. 17 | With Karan Singhal

Karan Singhal · 27m spoken Elad Gil · 7m spoken Sarah Guo · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Google researcher Karan Singhal joins Sarah Guo and Elad Gil to discuss the architecture, clinical evaluation, safety alignment, and real-world deployment of Google's Med-PaLM and Med-PaLM 2 models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.3% of the talking time here. How this is scored →

The hosts as informed peer 5.9 Guest teaching 4.4 Guest disagreement 0.1 The hosts pushing back 0.2
05100:0015:0030:000:02–4:01 · The hosts as informed peer 2/10 Karan Singhal's Background and the Origins of Medical AI at Google Sarah sets the stage warmly and invites Karan to detail his background transitioning from early misinformation detection into representation learning and pitching the Brain Moonshot at Google.4:01–9:28 · The hosts as informed peer 6/10 Technical Architecture and Evolution from PaLM to Med-PaLM 2 Elad asks precise technical questions regarding architectural differences between PaLM and Med-PaLM, prompting Karan to explain Flan-PaLM instruction prompt tuning and UL2 mixture of objectives alongside Chinchilla scaling laws.9:29–13:48 · The hosts as informed peer 5/10 Frameworks for Domain Alignment and Medical Evaluation Standards Sarah prompts a framework breakdown for aligning models across different data scale tiers. Karan outlines concrete heuristics from few-shot prompting to full fine-tuning and reviews gaps in human evaluation benchmarks like MedQA.13:49–17:22 · The hosts as informed peer 7/10 Defining the Quality Bar for Medical AI and Health Information Elad draws on his operating experience at Color and a personal emergency room anecdote to highlight the discrepancy between the perceived medical quality bar and real-world physician search practices. Karan acknowledges that 10% of internet searches are health-related and emphasizes grounded clinical workflow evaluations.17:22–21:37 · The hosts as informed peer 7/10 Commercial Workflows and Near-Term Clinical Applications Elad breaks down healthcare GDP allocations between pharmaceuticals and clinical decision workflows. Karan explains why drug discovery has proven an easier early commercial playbook compared to high-stakes physician assistant workflows like radiology reporting.21:38–27:44 · The hosts as informed peer 7/10 Navigating Patient Privacy, HIPAA Regulations, and Federated Learning Elad critiques legacy HIPAA restrictions with an MIT glioblastoma trial example, while Sarah asks about federated learning adoption. Karan explains why centralized foundation models currently outperform federated setups and outlines intermediate trusted execution environments.27:46–33:52 · The hosts as informed peer 6/10 Medical AI as a Testbed for Alignment and Scalable Oversight Sarah and Karan explore medical AI as an ideal crucible for technical alignment and scalable oversight. Karan outlines self-critique, AI debate protocols, and constitutional AI when model competence reaches or surpasses physician evaluation baselines.33:53–36:50 · The hosts as informed peer 8/10 Historical Lessons on Medical Adoption and Physician Perspectives Elad demonstrates deep historical perspective by citing Stanford's 1970s MYCIN expert system and its adoption roadblocks. Karan contextualizes current physician sentiment between fast-moving inflection point anxiety and practical excitement.36:52–42:10 · The hosts as informed peer 5/10 Future Outlook: Multimodality, Grounding, and 5-Year Vision Sarah frames the societal need for pragmatic safety bars rather than impossible standards. Karan outlines his five-year outlook covering multimodality, Toolformer-style grounding in authoritative medical literature, and refined human feedback methods.0:02–4:01 · Guest teaching 3/10 Karan Singhal's Background and the Origins of Medical AI at Google Sarah sets the stage warmly and invites Karan to detail his background transitioning from early misinformation detection into representation learning and pitching the Brain Moonshot at Google.4:01–9:28 · Guest teaching 6/10 Technical Architecture and Evolution from PaLM to Med-PaLM 2 Elad asks precise technical questions regarding architectural differences between PaLM and Med-PaLM, prompting Karan to explain Flan-PaLM instruction prompt tuning and UL2 mixture of objectives alongside Chinchilla scaling laws.9:29–13:48 · Guest teaching 5/10 Frameworks for Domain Alignment and Medical Evaluation Standards Sarah prompts a framework breakdown for aligning models across different data scale tiers. Karan outlines concrete heuristics from few-shot prompting to full fine-tuning and reviews gaps in human evaluation benchmarks like MedQA.13:49–17:22 · Guest teaching 4/10 Defining the Quality Bar for Medical AI and Health Information Elad draws on his operating experience at Color and a personal emergency room anecdote to highlight the discrepancy between the perceived medical quality bar and real-world physician search practices. Karan acknowledges that 10% of internet searches are health-related and emphasizes grounded clinical workflow evaluations.17:22–21:37 · Guest teaching 4/10 Commercial Workflows and Near-Term Clinical Applications Elad breaks down healthcare GDP allocations between pharmaceuticals and clinical decision workflows. Karan explains why drug discovery has proven an easier early commercial playbook compared to high-stakes physician assistant workflows like radiology reporting.21:38–27:44 · Guest teaching 5/10 Navigating Patient Privacy, HIPAA Regulations, and Federated Learning Elad critiques legacy HIPAA restrictions with an MIT glioblastoma trial example, while Sarah asks about federated learning adoption. Karan explains why centralized foundation models currently outperform federated setups and outlines intermediate trusted execution environments.27:46–33:52 · Guest teaching 6/10 Medical AI as a Testbed for Alignment and Scalable Oversight Sarah and Karan explore medical AI as an ideal crucible for technical alignment and scalable oversight. Karan outlines self-critique, AI debate protocols, and constitutional AI when model competence reaches or surpasses physician evaluation baselines.33:53–36:50 · Guest teaching 3/10 Historical Lessons on Medical Adoption and Physician Perspectives Elad demonstrates deep historical perspective by citing Stanford's 1970s MYCIN expert system and its adoption roadblocks. Karan contextualizes current physician sentiment between fast-moving inflection point anxiety and practical excitement.36:52–42:10 · Guest teaching 4/10 Future Outlook: Multimodality, Grounding, and 5-Year Vision Sarah frames the societal need for pragmatic safety bars rather than impossible standards. Karan outlines his five-year outlook covering multimodality, Toolformer-style grounding in authoritative medical literature, and refined human feedback methods.0:02–4:01 · Guest disagreement 0/10 Karan Singhal's Background and the Origins of Medical AI at Google Sarah sets the stage warmly and invites Karan to detail his background transitioning from early misinformation detection into representation learning and pitching the Brain Moonshot at Google.4:01–9:28 · Guest disagreement 0/10 Technical Architecture and Evolution from PaLM to Med-PaLM 2 Elad asks precise technical questions regarding architectural differences between PaLM and Med-PaLM, prompting Karan to explain Flan-PaLM instruction prompt tuning and UL2 mixture of objectives alongside Chinchilla scaling laws.9:29–13:48 · Guest disagreement 0/10 Frameworks for Domain Alignment and Medical Evaluation Standards Sarah prompts a framework breakdown for aligning models across different data scale tiers. Karan outlines concrete heuristics from few-shot prompting to full fine-tuning and reviews gaps in human evaluation benchmarks like MedQA.13:49–17:22 · Guest disagreement 1/10 Defining the Quality Bar for Medical AI and Health Information Elad draws on his operating experience at Color and a personal emergency room anecdote to highlight the discrepancy between the perceived medical quality bar and real-world physician search practices. Karan acknowledges that 10% of internet searches are health-related and emphasizes grounded clinical workflow evaluations.17:22–21:37 · Guest disagreement 0/10 Commercial Workflows and Near-Term Clinical Applications Elad breaks down healthcare GDP allocations between pharmaceuticals and clinical decision workflows. Karan explains why drug discovery has proven an easier early commercial playbook compared to high-stakes physician assistant workflows like radiology reporting.21:38–27:44 · Guest disagreement 0/10 Navigating Patient Privacy, HIPAA Regulations, and Federated Learning Elad critiques legacy HIPAA restrictions with an MIT glioblastoma trial example, while Sarah asks about federated learning adoption. Karan explains why centralized foundation models currently outperform federated setups and outlines intermediate trusted execution environments.27:46–33:52 · Guest disagreement 0/10 Medical AI as a Testbed for Alignment and Scalable Oversight Sarah and Karan explore medical AI as an ideal crucible for technical alignment and scalable oversight. Karan outlines self-critique, AI debate protocols, and constitutional AI when model competence reaches or surpasses physician evaluation baselines.33:53–36:50 · Guest disagreement 0/10 Historical Lessons on Medical Adoption and Physician Perspectives Elad demonstrates deep historical perspective by citing Stanford's 1970s MYCIN expert system and its adoption roadblocks. Karan contextualizes current physician sentiment between fast-moving inflection point anxiety and practical excitement.36:52–42:10 · Guest disagreement 0/10 Future Outlook: Multimodality, Grounding, and 5-Year Vision Sarah frames the societal need for pragmatic safety bars rather than impossible standards. Karan outlines his five-year outlook covering multimodality, Toolformer-style grounding in authoritative medical literature, and refined human feedback methods.0:02–4:01 · The hosts pushing back 0/10 Karan Singhal's Background and the Origins of Medical AI at Google Sarah sets the stage warmly and invites Karan to detail his background transitioning from early misinformation detection into representation learning and pitching the Brain Moonshot at Google.4:01–9:28 · The hosts pushing back 0/10 Technical Architecture and Evolution from PaLM to Med-PaLM 2 Elad asks precise technical questions regarding architectural differences between PaLM and Med-PaLM, prompting Karan to explain Flan-PaLM instruction prompt tuning and UL2 mixture of objectives alongside Chinchilla scaling laws.9:29–13:48 · The hosts pushing back 0/10 Frameworks for Domain Alignment and Medical Evaluation Standards Sarah prompts a framework breakdown for aligning models across different data scale tiers. Karan outlines concrete heuristics from few-shot prompting to full fine-tuning and reviews gaps in human evaluation benchmarks like MedQA.13:49–17:22 · The hosts pushing back 1/10 Defining the Quality Bar for Medical AI and Health Information Elad draws on his operating experience at Color and a personal emergency room anecdote to highlight the discrepancy between the perceived medical quality bar and real-world physician search practices. Karan acknowledges that 10% of internet searches are health-related and emphasizes grounded clinical workflow evaluations.17:22–21:37 · The hosts pushing back 0/10 Commercial Workflows and Near-Term Clinical Applications Elad breaks down healthcare GDP allocations between pharmaceuticals and clinical decision workflows. Karan explains why drug discovery has proven an easier early commercial playbook compared to high-stakes physician assistant workflows like radiology reporting.21:38–27:44 · The hosts pushing back 1/10 Navigating Patient Privacy, HIPAA Regulations, and Federated Learning Elad critiques legacy HIPAA restrictions with an MIT glioblastoma trial example, while Sarah asks about federated learning adoption. Karan explains why centralized foundation models currently outperform federated setups and outlines intermediate trusted execution environments.27:46–33:52 · The hosts pushing back 0/10 Medical AI as a Testbed for Alignment and Scalable Oversight Sarah and Karan explore medical AI as an ideal crucible for technical alignment and scalable oversight. Karan outlines self-critique, AI debate protocols, and constitutional AI when model competence reaches or surpasses physician evaluation baselines.33:53–36:50 · The hosts pushing back 0/10 Historical Lessons on Medical Adoption and Physician Perspectives Elad demonstrates deep historical perspective by citing Stanford's 1970s MYCIN expert system and its adoption roadblocks. Karan contextualizes current physician sentiment between fast-moving inflection point anxiety and practical excitement.36:52–42:10 · The hosts pushing back 0/10 Future Outlook: Multimodality, Grounding, and 5-Year Vision Sarah frames the societal need for pragmatic safety bars rather than impossible standards. Karan outlines his five-year outlook covering multimodality, Toolformer-style grounding in authoritative medical literature, and refined human feedback methods.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 28.6% · guest 71.4%0:00 · the hosts 28.6% · guest 71.4%3:00 · the hosts 3.5% · guest 96.5%3:00 · the hosts 3.5% · guest 96.5%6:00 · the hosts 26.2% · guest 73.8%6:00 · the hosts 26.2% · guest 73.8%9:00 · the hosts 22.2% · guest 77.8%9:00 · the hosts 22.2% · guest 77.8%12:00 · the hosts 41.5% · guest 58.5%12:00 · the hosts 41.5% · guest 58.5%15:00 · the hosts 50.7% · guest 49.3%15:00 · the hosts 50.7% · guest 49.3%18:00 · the hosts 21.9% · guest 78.1%18:00 · the hosts 21.9% · guest 78.1%21:00 · the hosts 22.7% · guest 77.3%21:00 · the hosts 22.7% · guest 77.3%24:00 · the hosts 41.1% · guest 58.9%24:00 · the hosts 41.1% · guest 58.9%27:00 · the hosts 7.7% · guest 92.3%27:00 · the hosts 7.7% · guest 92.3%30:00 · the hosts 21.7% · guest 78.3%30:00 · the hosts 21.7% · guest 78.3%33:00 · the hosts 51.4% · guest 48.6%33:00 · the hosts 51.4% · guest 48.6%36:00 · the hosts 55.1% · guest 44.9%36:00 · the hosts 55.1% · guest 44.9%39:00 · the hosts 1.5% · guest 98.5%39:00 · the hosts 1.5% · guest 98.5%42:00 · the hosts 34.5% · guest 65.5%42:00 · the hosts 34.5% · guest 65.5%
Sharpest disagreement ▶ 25:40 Gentle divergence on federated learning utility

In a very collegial discussion, Karan softly tempers enthusiasm around federated learning in health data, noting activation energy and compute realities favor centralized foundation models in the near term.

Hardest push from the hosts ▶ 13:48 Elad challenging idealized safety baselines

Elad challenges the unrealistic perfection standard demanded of AI by recounting an ER physician Googling symptoms in a cubicle, arguing the real baseline is often imperfect.

Biggest teaching moment ▶ 10:25 Karan's concrete fine-tuning decision framework

Karan provides a structured heuristic detailing exactly when to prompt (3-5 examples), prompt tune (10-50 examples), or full fine-tune (>100 examples) based on compute constraints.

The host holds their own ▶ 34:05 Elad citing the 1970s Stanford MYCIN system

Elad demonstrates specialized historical expertise by comparing modern clinical AI adoption barriers to Stanford's MYCIN expert system from four decades earlier.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Karan Singhal's Background and the Origins of Medical AI at Google 2300 Sarah sets the stage warmly and invites Karan to detail his background transitioning from early misinformation detection into representation learning and pitching the Brain Moonshot at Google.
Technical Architecture and Evolution from PaLM to Med-PaLM 2 6600 Elad asks precise technical questions regarding architectural differences between PaLM and Med-PaLM, prompting Karan to explain Flan-PaLM instruction prompt tuning and UL2 mixture of objectives alongside Chinchilla scaling laws.
Frameworks for Domain Alignment and Medical Evaluation Standards 5500 Sarah prompts a framework breakdown for aligning models across different data scale tiers. Karan outlines concrete heuristics from few-shot prompting to full fine-tuning and reviews gaps in human evaluation benchmarks like MedQA.
Defining the Quality Bar for Medical AI and Health Information 7411 Elad draws on his operating experience at Color and a personal emergency room anecdote to highlight the discrepancy between the perceived medical quality bar and real-world physician search practices. Karan acknowledges that 10% of internet searches are health-related and emphasizes grounded clinical workflow evaluations.
Commercial Workflows and Near-Term Clinical Applications 7400 Elad breaks down healthcare GDP allocations between pharmaceuticals and clinical decision workflows. Karan explains why drug discovery has proven an easier early commercial playbook compared to high-stakes physician assistant workflows like radiology reporting.
Navigating Patient Privacy, HIPAA Regulations, and Federated Learning 7501 Elad critiques legacy HIPAA restrictions with an MIT glioblastoma trial example, while Sarah asks about federated learning adoption. Karan explains why centralized foundation models currently outperform federated setups and outlines intermediate trusted execution environments.
Medical AI as a Testbed for Alignment and Scalable Oversight 6600 Sarah and Karan explore medical AI as an ideal crucible for technical alignment and scalable oversight. Karan outlines self-critique, AI debate protocols, and constitutional AI when model competence reaches or surpasses physician evaluation baselines.
Historical Lessons on Medical Adoption and Physician Perspectives 8300 Elad demonstrates deep historical perspective by citing Stanford's 1970s MYCIN expert system and its adoption roadblocks. Karan contextualizes current physician sentiment between fast-moving inflection point anxiety and practical excitement.
Future Outlook: Multimodality, Grounding, and 5-Year Vision 5400 Sarah frames the societal need for pragmatic safety bars rather than impossible standards. Karan outlines his five-year outlook covering multimodality, Toolformer-style grounding in authoritative medical literature, and refined human feedback methods.

Statements from this episode (18)

Assertion Supported
Singhal: Google's 540B PaLM was the largest densely activated model in 2022
“The first Palm model was released in twenty-twenty-two, which was kind of this 540 B decoder only transformer model at the time, the largest densely activated model.”
Karan Singhal May 18, 2023 ▶ 4:24
Assertion Supported
Singhal: Flan-PaLM was the first AI model to pass the USMLE
“When we took a variation of POM, the FlanPOM model, which was, again, work from Jason Wei and team you know, this is an instruction to a model that's been trained to follow instructions better. You know, again, it was able to perform quite well out of the box,…”
Karan Singhal May 18, 2023 ▶ 6:42
Insight
Singhal: Fine-tuning outperforms prompt tuning when providing over 100 examples
“If you have three to five examples, let's say, then I would prompt it. If you have maybe 10 or 50 examples, it would either be prompt tuning or fine tuning. I think generally in that realm, prompt tuning and fine tuning perform similarly, and I would prefer pr…”
Karan Singhal May 18, 2023 ▶ 11:12
Assertion Not checkable as stated
Singhal: Prior biomedical LLMs lacked systematic benchmarking and human evaluation
“There was a bit of a shortage of kind of a systematic way of doing evaluation of these models. And so it didn't feel like there was a systematic way to think about automated evaluation of the clinical knowledge of these models. So for example, via multiple cho…”
Karan Singhal May 18, 2023 ▶ 12:35
Assertion Supported
Singhal: Roughly 10% of all internet searches seek health information
“Roughly 10% of searches on the internet are for health information.”
Karan Singhal May 18, 2023 ▶ 15:37
Disclosure
Singhal: Medical AI still lacks grounded evaluations within specific clinical workflows
“One thing that has been missing from our work so far is really Grounded evaluations in a specific use case in a workflow to show that there is a benefit both in terms of safety in the short term and in terms of kind of long-term patient outcomes as well.”
Karan Singhal May 18, 2023 ▶ 16:00
Assertion Partly supported
Gil: Healthcare is 20% of GDP; pharmaceuticals are 20% of that
“If you look at healthcare, it's 20% of GDP. Pharmaceuticals are about 20% of that. And then drug development is a fraction of that, right? So really what you folks are focused on in terms of the types of models that you're building is at least, you know, 16% o…”
Elad Gil May 18, 2023 ▶ 19:30
Prediction Held up
Singhal: Epic will likely partner with foundation models for clinical documentation
“I think that is also going to be something where players like Epic are going to be able to partner with existing models and I think potentially deliver real value there.”
Karan Singhal May 18, 2023 ▶ 20:27
Prediction Not checkable as stated
Singhal: Specialized AI models will assist radiologists in the near term
“I think where there might be more of a need for specialized models Is when it comes down to kind of higher stakes workflows, and I think that might look in the short term more like a physician's assistant. And so imagine, for example, an agent that can work wi…”
Karan Singhal May 18, 2023 ▶ 20:44
Opinion
Gil: HIPAA legislation has backfired in many ways for patient good
“HIPAA is kind of interesting from the context of it was an incredibly well intentioned piece of legislation, but the flip side of it is it's really backfired in all sorts of ways in terms of actual patient good.”
Elad Gil May 18, 2023 ▶ 23:41
Assertion Supported
Singhal: Med-PaLM and Med-PaLM 2 were trained without patient health information
“Like for example, MedPOM and MedPOM-II are trained without any patient health information. They, they're just kind of taking all the knowledge of POM and POM-II and then just kind of Aligning them and making them behave in a certain way.”
Karan Singhal May 18, 2023 ▶ 25:58
Prediction Not checkable as stated
Singhal: Federated learning won't drive biggest near-term healthcare AI advances
“But I think there are like real world obstacles to doing federally learning on health data, which actually kind of increased activation energy to the point where in the next few years, I doubt that like the biggest advances are going to come. From using federa…”
Karan Singhal May 18, 2023 ▶ 26:57
Opinion
Singhal: Medical AI is an ideal testbed for safety and alignment
“I think there's a good chance that this setting, this medical setting for example, medical question answering, Or maybe more broadly, I think ends up being a better scenario to study concerns about technical safety and to mitigate concerns like misaligned with…”
Karan Singhal May 18, 2023 ▶ 28:52
Assertion Supported
Singhal: Evaluators can barely distinguish Med-PaLM 2 answers from human physicians
“One thing we're seeing with MedPOM-II as we get closer to physician-level performance on medical question answering is that it's hard to tell the difference anymore. It's hard to tell the difference between different models. It's hard to tell the difference be…”
Karan Singhal May 18, 2023 ▶ 30:09
Prediction Not checkable as stated
Gil: AI physician assistants will create unique closed-loop clinical datasets
“One thing that I feel like would also sort of be generated as a side effect of all this is just you end up with these really interesting closed loop data sets over time that may be unique outside of an EMR or somewhere else or a really robust medical record sy…”
Elad Gil May 18, 2023 ▶ 33:15
Assertion Supported
Gil: Stanford's 1970s MYCIN AI outperformed doctors but was never clinically adopted
“In the 19 seventies, there was something known as the Mycene project at Stanford, where they built an expert system... They had a expert system that outperformed all of Stanford's medical staff on the prediction of the infectious disease that somebody had. So …”
Elad Gil May 18, 2023 ▶ 33:58
Prediction Not checkable as stated
Singhal: Augmenting telemedicine with LLMs is highly achievable within five years
“I think augmenting telemedicine, I think is, is, is kind of a short-term opportunity that I think in the next five years is, is very achievable.”
Karan Singhal May 18, 2023 ▶ 39:12
Assertion Not checkable as stated
Singhal: Google presents health info by attributing it to authoritative sources
“Where, for example, Google is doing that with health information is largely because it can attribute things to the Mayo Clinic and other organizations.”
Karan Singhal May 18, 2023 ▶ 41:12
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.