May 18, 2023 · 42m · no-priors
No Priors Ep. 17 | With Karan Singhal
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of No Priors, Google researcher Karan Singhal joins Sarah Guo and Elad Gil to discuss the architecture, clinical evaluation, safety alignment, and real-world deployment of Google's Med-PaLM and Med-PaLM 2 models.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 28.3% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
In a very collegial discussion, Karan softly tempers enthusiasm around federated learning in health data, noting activation energy and compute realities favor centralized foundation models in the near term.
Hardest push from the hosts ▶ 13:48 Elad challenging idealized safety baselinesElad challenges the unrealistic perfection standard demanded of AI by recounting an ER physician Googling symptoms in a cubicle, arguing the real baseline is often imperfect.
Biggest teaching moment ▶ 10:25 Karan's concrete fine-tuning decision frameworkKaran provides a structured heuristic detailing exactly when to prompt (3-5 examples), prompt tune (10-50 examples), or full fine-tune (>100 examples) based on compute constraints.
The host holds their own ▶ 34:05 Elad citing the 1970s Stanford MYCIN systemElad demonstrates specialized historical expertise by comparing modern clinical AI adoption barriers to Stanford's MYCIN expert system from four decades earlier.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Karan Singhal's Background and the Origins of Medical AI at Google | 2 | 3 | 0 | 0 | Sarah sets the stage warmly and invites Karan to detail his background transitioning from early misinformation detection into representation learning and pitching the Brain Moonshot at Google. | |
| Technical Architecture and Evolution from PaLM to Med-PaLM 2 | 6 | 6 | 0 | 0 | Elad asks precise technical questions regarding architectural differences between PaLM and Med-PaLM, prompting Karan to explain Flan-PaLM instruction prompt tuning and UL2 mixture of objectives alongside Chinchilla scaling laws. | |
| Frameworks for Domain Alignment and Medical Evaluation Standards | 5 | 5 | 0 | 0 | Sarah prompts a framework breakdown for aligning models across different data scale tiers. Karan outlines concrete heuristics from few-shot prompting to full fine-tuning and reviews gaps in human evaluation benchmarks like MedQA. | |
| Defining the Quality Bar for Medical AI and Health Information | 7 | 4 | 1 | 1 | Elad draws on his operating experience at Color and a personal emergency room anecdote to highlight the discrepancy between the perceived medical quality bar and real-world physician search practices. Karan acknowledges that 10% of internet searches are health-related and emphasizes grounded clinical workflow evaluations. | |
| Commercial Workflows and Near-Term Clinical Applications | 7 | 4 | 0 | 0 | Elad breaks down healthcare GDP allocations between pharmaceuticals and clinical decision workflows. Karan explains why drug discovery has proven an easier early commercial playbook compared to high-stakes physician assistant workflows like radiology reporting. | |
| Navigating Patient Privacy, HIPAA Regulations, and Federated Learning | 7 | 5 | 0 | 1 | Elad critiques legacy HIPAA restrictions with an MIT glioblastoma trial example, while Sarah asks about federated learning adoption. Karan explains why centralized foundation models currently outperform federated setups and outlines intermediate trusted execution environments. | |
| Medical AI as a Testbed for Alignment and Scalable Oversight | 6 | 6 | 0 | 0 | Sarah and Karan explore medical AI as an ideal crucible for technical alignment and scalable oversight. Karan outlines self-critique, AI debate protocols, and constitutional AI when model competence reaches or surpasses physician evaluation baselines. | |
| Historical Lessons on Medical Adoption and Physician Perspectives | 8 | 3 | 0 | 0 | Elad demonstrates deep historical perspective by citing Stanford's 1970s MYCIN expert system and its adoption roadblocks. Karan contextualizes current physician sentiment between fast-moving inflection point anxiety and practical excitement. | |
| Future Outlook: Multimodality, Grounding, and 5-Year Vision | 5 | 4 | 0 | 0 | Sarah frames the societal need for pragmatic safety bars rather than impossible standards. Karan outlines his five-year outlook covering multimodality, Toolformer-style grounding in authoritative medical literature, and refined human feedback methods. |