Aug 23, 2023 · 47m · mad
Hippocratic AI’s Munjal Shah: Building the First Safety-First LLM for Healthcare
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of The MAD Podcast, host Matt Turck interviews Munjal Shah, Co-Founder and CEO of Hippocratic AI, about developing the first safety-focused large language model specifically designed for healthcare. Shah discusses the strategic decision to focus on non-diagnostic patient adherence and super staffing, detailing the technical, clinical, and regulatory frameworks required to deploy voice-enabled AI agents safely.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.7% of the talking time here. How this is scored →
speaking balance: gold is Matt, purple is the guest (3 minute bins)
Munjal forcefully rejects the premise of building LLMs for medical diagnosis, calling competing founders naive, comparing them to college dropouts building bar-line apps, and claiming they will kill somebody.
Hardest push from Matt ▶ 4:01 Questioning the Avoidance of Medical DiagnosisMatt directly challenges Munjal's strategic positioning, asking why a technology as transformative as generative AI shouldn't be applied to medical diagnosis.
Biggest teaching moment ▶ 5:20 Healthcare Market Spend BreakdownMunjal educates the host on healthcare macro-economics, breaking down the 4.2 trillion dollar market to demonstrate that diagnosis represents only 600 billion while non-diagnostic care represents 3.6 trillion dollars.
Matt holds his own ▶ 20:03 Host Citing Benchmark StatisticsMatt displays thorough research by citing precise performance metrics, noting that Hippocratic AI outperformed GPT-4 on 105 out of 114 medical board exams.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Matt as informed peer | Guest teaching | Guest disagreement | Matt pushing back | Why |
|---|---|---|---|---|---|---|
| Munjal Shah's Background and Origin Story | 2 | 2 | 0 | 0 | Matt opens by establishing Munjal's background in computer vision and recent fundraising success, asking for his origin story into healthcare. Munjal warmly details his academic background in CS, his ER visit at age 37, and his epiphany upon testing ChatGPT. | |
| Choosing Adherence Over Medical Diagnosis | 2 | 6 | 5 | 3 | Matt challenges Munjal on why generative AI shouldn't be used for medical diagnosis. Munjal aggressively dismisses AI diagnosis startups as dangerous and uninformed, breaking down the 4.2 trillion dollar healthcare market to show that non-diagnostic care adherence accounts for 3.6 trillion dollars. | |
| Three Healthcare LLM Use Cases and Super Staffing | 2 | 5 | 1 | 0 | Matt prompts Munjal to define Hippocratic AI's targeted use cases. Munjal educates the host on workflow vs staffing vs super staffing, explaining how 18-cents-per-hour voice LLMs enable continuous patient follow-ups that were previously economically impossible. | |
| Autonomous Agents and Voice Conversational Paradigms | 3 | 5 | 1 | 1 | Matt asks whether Hippocratic AI acts as a copilot or autonomous agent. Munjal explains that conversational LLMs require a new paradigm of instruction tuning for multi-turn interactions rather than single queries, quoting Yuval Noah Harari on the power of undivided attention. | |
| Model Certification Benchmarks and Clinical RLHF | 3 | 4 | 0 | 1 | Matt demonstrates preparation by citing specific benchmark stats where Hippocratic outperformed GPT-4 on medical licensing exams. Munjal uses an iceberg analogy to explain why healthcare domain data is hidden behind firewalls and requires proprietary pre-training and clinical RLHF. | |
| Custom Tokenization and Domain Data Acquisition | 3 | 4 | 1 | 1 | Matt probes into data acquisition strategies for domain-specific models. Munjal expands the definition of healthcare data to include restaurant menus for dietician agents and 200-page insurance plan PDFs, alongside custom tokenizers for synthetic drug names. | |
| Real-Time Infrastructure and Tone Classification | 3 | 4 | 1 | 1 | Matt asks about open source vs proprietary bases, hallucination risks, and vector database retrieval. Munjal argues that application choice is the primary defense against hallucinations and notes that document retrieval via RAG is far harder than most assume. | |
| Safety-First Philosophy and Bottom-Up Regulation | 3 | 4 | 1 | 0 | Matt guides the discussion to ethics, regulation, and release timelines. Munjal introduces his concept of bottoms-up regulation, where panels of active nurses and clinicians perform QA and RLHF to establish safety before launch. | |
| Team Structure, Clinician Roles, and Office Culture | 2 | 3 | 0 | 0 | Matt asks about regulator attitudes and startup culture. Munjal describes their team makeup, blending top AI researchers with full-time clinicians, and strongly advocates for a 100 percent in-person office culture. |