Aug 23, 2023 · 47m · mad

Hippocratic AI’s Munjal Shah: Building the First Safety-First LLM for Healthcare

Munjal Shah · 33m spoken Matt Turck · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Munjal Shah, Co-Founder and CEO of Hippocratic AI, about developing the first safety-focused large language model specifically designed for healthcare. Shah discusses the strategic decision to focus on non-diagnostic patient adherence and super staffing, detailing the technical, clinical, and regulatory frameworks required to deploy voice-enabled AI agents safely.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 17.7% of the talking time here. How this is scored →

Matt as informed peer 2.6 Guest teaching 4.1 Guest disagreement 1.1 Matt pushing back 0.8
05100:0015:0030:0045:001:03–4:01 · Matt as informed peer 2/10 Munjal Shah's Background and Origin Story Matt opens by establishing Munjal's background in computer vision and recent fundraising success, asking for his origin story into healthcare. Munjal warmly details his academic background in CS, his ER visit at age 37, and his epiphany upon testing ChatGPT.4:01–7:16 · Matt as informed peer 2/10 Choosing Adherence Over Medical Diagnosis Matt challenges Munjal on why generative AI shouldn't be used for medical diagnosis. Munjal aggressively dismisses AI diagnosis startups as dangerous and uninformed, breaking down the 4.2 trillion dollar healthcare market to show that non-diagnostic care adherence accounts for 3.6 trillion dollars.7:16–13:46 · Matt as informed peer 2/10 Three Healthcare LLM Use Cases and Super Staffing Matt prompts Munjal to define Hippocratic AI's targeted use cases. Munjal educates the host on workflow vs staffing vs super staffing, explaining how 18-cents-per-hour voice LLMs enable continuous patient follow-ups that were previously economically impossible.13:46–20:03 · Matt as informed peer 3/10 Autonomous Agents and Voice Conversational Paradigms Matt asks whether Hippocratic AI acts as a copilot or autonomous agent. Munjal explains that conversational LLMs require a new paradigm of instruction tuning for multi-turn interactions rather than single queries, quoting Yuval Noah Harari on the power of undivided attention.20:03–23:56 · Matt as informed peer 3/10 Model Certification Benchmarks and Clinical RLHF Matt demonstrates preparation by citing specific benchmark stats where Hippocratic outperformed GPT-4 on medical licensing exams. Munjal uses an iceberg analogy to explain why healthcare domain data is hidden behind firewalls and requires proprietary pre-training and clinical RLHF.23:56–29:32 · Matt as informed peer 3/10 Custom Tokenization and Domain Data Acquisition Matt probes into data acquisition strategies for domain-specific models. Munjal expands the definition of healthcare data to include restaurant menus for dietician agents and 200-page insurance plan PDFs, alongside custom tokenizers for synthetic drug names.29:32–34:43 · Matt as informed peer 3/10 Real-Time Infrastructure and Tone Classification Matt asks about open source vs proprietary bases, hallucination risks, and vector database retrieval. Munjal argues that application choice is the primary defense against hallucinations and notes that document retrieval via RAG is far harder than most assume.34:43–40:29 · Matt as informed peer 3/10 Safety-First Philosophy and Bottom-Up Regulation Matt guides the discussion to ethics, regulation, and release timelines. Munjal introduces his concept of bottoms-up regulation, where panels of active nurses and clinicians perform QA and RLHF to establish safety before launch.40:29–46:48 · Matt as informed peer 2/10 Team Structure, Clinician Roles, and Office Culture Matt asks about regulator attitudes and startup culture. Munjal describes their team makeup, blending top AI researchers with full-time clinicians, and strongly advocates for a 100 percent in-person office culture.1:03–4:01 · Guest teaching 2/10 Munjal Shah's Background and Origin Story Matt opens by establishing Munjal's background in computer vision and recent fundraising success, asking for his origin story into healthcare. Munjal warmly details his academic background in CS, his ER visit at age 37, and his epiphany upon testing ChatGPT.4:01–7:16 · Guest teaching 6/10 Choosing Adherence Over Medical Diagnosis Matt challenges Munjal on why generative AI shouldn't be used for medical diagnosis. Munjal aggressively dismisses AI diagnosis startups as dangerous and uninformed, breaking down the 4.2 trillion dollar healthcare market to show that non-diagnostic care adherence accounts for 3.6 trillion dollars.7:16–13:46 · Guest teaching 5/10 Three Healthcare LLM Use Cases and Super Staffing Matt prompts Munjal to define Hippocratic AI's targeted use cases. Munjal educates the host on workflow vs staffing vs super staffing, explaining how 18-cents-per-hour voice LLMs enable continuous patient follow-ups that were previously economically impossible.13:46–20:03 · Guest teaching 5/10 Autonomous Agents and Voice Conversational Paradigms Matt asks whether Hippocratic AI acts as a copilot or autonomous agent. Munjal explains that conversational LLMs require a new paradigm of instruction tuning for multi-turn interactions rather than single queries, quoting Yuval Noah Harari on the power of undivided attention.20:03–23:56 · Guest teaching 4/10 Model Certification Benchmarks and Clinical RLHF Matt demonstrates preparation by citing specific benchmark stats where Hippocratic outperformed GPT-4 on medical licensing exams. Munjal uses an iceberg analogy to explain why healthcare domain data is hidden behind firewalls and requires proprietary pre-training and clinical RLHF.23:56–29:32 · Guest teaching 4/10 Custom Tokenization and Domain Data Acquisition Matt probes into data acquisition strategies for domain-specific models. Munjal expands the definition of healthcare data to include restaurant menus for dietician agents and 200-page insurance plan PDFs, alongside custom tokenizers for synthetic drug names.29:32–34:43 · Guest teaching 4/10 Real-Time Infrastructure and Tone Classification Matt asks about open source vs proprietary bases, hallucination risks, and vector database retrieval. Munjal argues that application choice is the primary defense against hallucinations and notes that document retrieval via RAG is far harder than most assume.34:43–40:29 · Guest teaching 4/10 Safety-First Philosophy and Bottom-Up Regulation Matt guides the discussion to ethics, regulation, and release timelines. Munjal introduces his concept of bottoms-up regulation, where panels of active nurses and clinicians perform QA and RLHF to establish safety before launch.40:29–46:48 · Guest teaching 3/10 Team Structure, Clinician Roles, and Office Culture Matt asks about regulator attitudes and startup culture. Munjal describes their team makeup, blending top AI researchers with full-time clinicians, and strongly advocates for a 100 percent in-person office culture.1:03–4:01 · Guest disagreement 0/10 Munjal Shah's Background and Origin Story Matt opens by establishing Munjal's background in computer vision and recent fundraising success, asking for his origin story into healthcare. Munjal warmly details his academic background in CS, his ER visit at age 37, and his epiphany upon testing ChatGPT.4:01–7:16 · Guest disagreement 5/10 Choosing Adherence Over Medical Diagnosis Matt challenges Munjal on why generative AI shouldn't be used for medical diagnosis. Munjal aggressively dismisses AI diagnosis startups as dangerous and uninformed, breaking down the 4.2 trillion dollar healthcare market to show that non-diagnostic care adherence accounts for 3.6 trillion dollars.7:16–13:46 · Guest disagreement 1/10 Three Healthcare LLM Use Cases and Super Staffing Matt prompts Munjal to define Hippocratic AI's targeted use cases. Munjal educates the host on workflow vs staffing vs super staffing, explaining how 18-cents-per-hour voice LLMs enable continuous patient follow-ups that were previously economically impossible.13:46–20:03 · Guest disagreement 1/10 Autonomous Agents and Voice Conversational Paradigms Matt asks whether Hippocratic AI acts as a copilot or autonomous agent. Munjal explains that conversational LLMs require a new paradigm of instruction tuning for multi-turn interactions rather than single queries, quoting Yuval Noah Harari on the power of undivided attention.20:03–23:56 · Guest disagreement 0/10 Model Certification Benchmarks and Clinical RLHF Matt demonstrates preparation by citing specific benchmark stats where Hippocratic outperformed GPT-4 on medical licensing exams. Munjal uses an iceberg analogy to explain why healthcare domain data is hidden behind firewalls and requires proprietary pre-training and clinical RLHF.23:56–29:32 · Guest disagreement 1/10 Custom Tokenization and Domain Data Acquisition Matt probes into data acquisition strategies for domain-specific models. Munjal expands the definition of healthcare data to include restaurant menus for dietician agents and 200-page insurance plan PDFs, alongside custom tokenizers for synthetic drug names.29:32–34:43 · Guest disagreement 1/10 Real-Time Infrastructure and Tone Classification Matt asks about open source vs proprietary bases, hallucination risks, and vector database retrieval. Munjal argues that application choice is the primary defense against hallucinations and notes that document retrieval via RAG is far harder than most assume.34:43–40:29 · Guest disagreement 1/10 Safety-First Philosophy and Bottom-Up Regulation Matt guides the discussion to ethics, regulation, and release timelines. Munjal introduces his concept of bottoms-up regulation, where panels of active nurses and clinicians perform QA and RLHF to establish safety before launch.40:29–46:48 · Guest disagreement 0/10 Team Structure, Clinician Roles, and Office Culture Matt asks about regulator attitudes and startup culture. Munjal describes their team makeup, blending top AI researchers with full-time clinicians, and strongly advocates for a 100 percent in-person office culture.1:03–4:01 · Matt pushing back 0/10 Munjal Shah's Background and Origin Story Matt opens by establishing Munjal's background in computer vision and recent fundraising success, asking for his origin story into healthcare. Munjal warmly details his academic background in CS, his ER visit at age 37, and his epiphany upon testing ChatGPT.4:01–7:16 · Matt pushing back 3/10 Choosing Adherence Over Medical Diagnosis Matt challenges Munjal on why generative AI shouldn't be used for medical diagnosis. Munjal aggressively dismisses AI diagnosis startups as dangerous and uninformed, breaking down the 4.2 trillion dollar healthcare market to show that non-diagnostic care adherence accounts for 3.6 trillion dollars.7:16–13:46 · Matt pushing back 0/10 Three Healthcare LLM Use Cases and Super Staffing Matt prompts Munjal to define Hippocratic AI's targeted use cases. Munjal educates the host on workflow vs staffing vs super staffing, explaining how 18-cents-per-hour voice LLMs enable continuous patient follow-ups that were previously economically impossible.13:46–20:03 · Matt pushing back 1/10 Autonomous Agents and Voice Conversational Paradigms Matt asks whether Hippocratic AI acts as a copilot or autonomous agent. Munjal explains that conversational LLMs require a new paradigm of instruction tuning for multi-turn interactions rather than single queries, quoting Yuval Noah Harari on the power of undivided attention.20:03–23:56 · Matt pushing back 1/10 Model Certification Benchmarks and Clinical RLHF Matt demonstrates preparation by citing specific benchmark stats where Hippocratic outperformed GPT-4 on medical licensing exams. Munjal uses an iceberg analogy to explain why healthcare domain data is hidden behind firewalls and requires proprietary pre-training and clinical RLHF.23:56–29:32 · Matt pushing back 1/10 Custom Tokenization and Domain Data Acquisition Matt probes into data acquisition strategies for domain-specific models. Munjal expands the definition of healthcare data to include restaurant menus for dietician agents and 200-page insurance plan PDFs, alongside custom tokenizers for synthetic drug names.29:32–34:43 · Matt pushing back 1/10 Real-Time Infrastructure and Tone Classification Matt asks about open source vs proprietary bases, hallucination risks, and vector database retrieval. Munjal argues that application choice is the primary defense against hallucinations and notes that document retrieval via RAG is far harder than most assume.34:43–40:29 · Matt pushing back 0/10 Safety-First Philosophy and Bottom-Up Regulation Matt guides the discussion to ethics, regulation, and release timelines. Munjal introduces his concept of bottoms-up regulation, where panels of active nurses and clinicians perform QA and RLHF to establish safety before launch.40:29–46:48 · Matt pushing back 0/10 Team Structure, Clinician Roles, and Office Culture Matt asks about regulator attitudes and startup culture. Munjal describes their team makeup, blending top AI researchers with full-time clinicians, and strongly advocates for a 100 percent in-person office culture.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 52.7% · guest 47.3%0:00 · Matt 52.7% · guest 47.3%3:00 · Matt 26% · guest 74%3:00 · Matt 26% · guest 74%6:00 · Matt 22.7% · guest 77.3%6:00 · Matt 22.7% · guest 77.3%9:00 · Matt 0% · guest 100%9:00 · Matt 0% · guest 100%12:00 · Matt 12.7% · guest 87.3%12:00 · Matt 12.7% · guest 87.3%15:00 · Matt 0.1% · guest 99.9%15:00 · Matt 0.1% · guest 99.9%18:00 · Matt 21.3% · guest 78.7%18:00 · Matt 21.3% · guest 78.7%21:00 · Matt 10.4% · guest 89.6%21:00 · Matt 10.4% · guest 89.6%24:00 · Matt 16.6% · guest 83.4%24:00 · Matt 16.6% · guest 83.4%27:00 · Matt 9.2% · guest 90.8%27:00 · Matt 9.2% · guest 90.8%30:00 · Matt 9.6% · guest 90.4%30:00 · Matt 9.6% · guest 90.4%33:00 · Matt 33.4% · guest 66.6%33:00 · Matt 33.4% · guest 66.6%36:00 · Matt 0% · guest 100%36:00 · Matt 0% · guest 100%39:00 · Matt 28.7% · guest 71.3%39:00 · Matt 28.7% · guest 71.3%42:00 · Matt 11.1% · guest 88.9%42:00 · Matt 11.1% · guest 88.9%45:00 · Matt 31.4% · guest 68.6%45:00 · Matt 31.4% · guest 68.6%
Sharpest disagreement ▶ 4:49 Dismissal of Diagnostic AI Startups

Munjal forcefully rejects the premise of building LLMs for medical diagnosis, calling competing founders naive, comparing them to college dropouts building bar-line apps, and claiming they will kill somebody.

Hardest push from Matt ▶ 4:01 Questioning the Avoidance of Medical Diagnosis

Matt directly challenges Munjal's strategic positioning, asking why a technology as transformative as generative AI shouldn't be applied to medical diagnosis.

Biggest teaching moment ▶ 5:20 Healthcare Market Spend Breakdown

Munjal educates the host on healthcare macro-economics, breaking down the 4.2 trillion dollar market to demonstrate that diagnosis represents only 600 billion while non-diagnostic care represents 3.6 trillion dollars.

Matt holds his own ▶ 20:03 Host Citing Benchmark Statistics

Matt displays thorough research by citing precise performance metrics, noting that Hippocratic AI outperformed GPT-4 on 105 out of 114 medical board exams.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Munjal Shah's Background and Origin Story 2200 Matt opens by establishing Munjal's background in computer vision and recent fundraising success, asking for his origin story into healthcare. Munjal warmly details his academic background in CS, his ER visit at age 37, and his epiphany upon testing ChatGPT.
Choosing Adherence Over Medical Diagnosis 2653 Matt challenges Munjal on why generative AI shouldn't be used for medical diagnosis. Munjal aggressively dismisses AI diagnosis startups as dangerous and uninformed, breaking down the 4.2 trillion dollar healthcare market to show that non-diagnostic care adherence accounts for 3.6 trillion dollars.
Three Healthcare LLM Use Cases and Super Staffing 2510 Matt prompts Munjal to define Hippocratic AI's targeted use cases. Munjal educates the host on workflow vs staffing vs super staffing, explaining how 18-cents-per-hour voice LLMs enable continuous patient follow-ups that were previously economically impossible.
Autonomous Agents and Voice Conversational Paradigms 3511 Matt asks whether Hippocratic AI acts as a copilot or autonomous agent. Munjal explains that conversational LLMs require a new paradigm of instruction tuning for multi-turn interactions rather than single queries, quoting Yuval Noah Harari on the power of undivided attention.
Model Certification Benchmarks and Clinical RLHF 3401 Matt demonstrates preparation by citing specific benchmark stats where Hippocratic outperformed GPT-4 on medical licensing exams. Munjal uses an iceberg analogy to explain why healthcare domain data is hidden behind firewalls and requires proprietary pre-training and clinical RLHF.
Custom Tokenization and Domain Data Acquisition 3411 Matt probes into data acquisition strategies for domain-specific models. Munjal expands the definition of healthcare data to include restaurant menus for dietician agents and 200-page insurance plan PDFs, alongside custom tokenizers for synthetic drug names.
Real-Time Infrastructure and Tone Classification 3411 Matt asks about open source vs proprietary bases, hallucination risks, and vector database retrieval. Munjal argues that application choice is the primary defense against hallucinations and notes that document retrieval via RAG is far harder than most assume.
Safety-First Philosophy and Bottom-Up Regulation 3410 Matt guides the discussion to ethics, regulation, and release timelines. Munjal introduces his concept of bottoms-up regulation, where panels of active nurses and clinicians perform QA and RLHF to establish safety before launch.
Team Structure, Clinician Roles, and Office Culture 2300 Matt asks about regulator attitudes and startup culture. Munjal describes their team makeup, blending top AI researchers with full-time clinicians, and strongly advocates for a 100 percent in-person office culture.

Statements from this episode (19)

Insight
Shah: Pre-generative AI functioned as an idiot savant in narrow tasks
“It's been a bit of what I call an idiot savant till now. It could do one narrow thing well, but if you took it anything outside of that, it just like didn't work. You know, at least traditional classifier AI.”
Munjal Shah Aug 23, 2023 ▶ 3:38
Prediction Not checkable as stated
Shah: AI models attempting medical diagnoses will kill patients
“I mean, they just keep going there. I'm going to do diagnoses because it's grand or it's challenging, but I mean, they're going to kill somebody.”
Munjal Shah Aug 23, 2023 ▶ 4:53
Assertion Not checkable as stated
Shah: Patient non-adherence is healthcare's biggest problem, not diagnosis
“Our biggest issue in the U.S. And most of the developed world is the people don't follow the directions after they're diagnosed. Like, it's not an issue of diagnoses. It's an issue of adherence.”
Munjal Shah Aug 23, 2023 ▶ 5:59
Assertion Partly supported
Shah: The US healthcare system is 20% to 30% understaffed in nursing
“And so in the country today, we're 20%, 30% understaffed and like nurses alone. This isn't just true in the US. This is true in every single country out there.”
Munjal Shah Aug 23, 2023 ▶ 9:27
Assertion Supported
Shah: Operating a voice LLM costs roughly 18 cents per hour
“A large language model speaking at a hundred words per minute will cost somewhere around 18 cents an hour. Okay. So the LLM plus what's the ASR cost, the automatic speech recognition cost, the text-to-speech cost, the TTS cost. We kind of put it all in there. …”
Munjal Shah Aug 23, 2023 ▶ 10:49
Insight
Shah: US health plans avoid prevention because Americans change jobs triennially
“It turns out the average person in the US changes healthcare plans every three years because they change jobs every three years. And so the healthcare plan will bear the cost to do the prevention, but won't basically the benefit because actually most of the be…”
Munjal Shah Aug 23, 2023 ▶ 13:16
Assertion Contradicted
Shah: Not interrupting patients is the primary driver of perceived doctor empathy
“The number one driver of whether you thought you had an empathetic doctor was did they not cut you off when you started telling your story?”
Munjal Shah Aug 23, 2023 ▶ 16:18
Assertion Supported
Shah: Llama 2 instruction tuning datasets averaged one instruction per sample
“The average number of instructions on, there's, they have like seven data sets they showed for their instruction tuning. The average number of instructions? One.”
Munjal Shah Aug 23, 2023 ▶ 18:26
Assertion Not checkable as stated
Shah: Foundational LLMs are instruction-tuned for queries, not multi-turn conversations
“They've all instruction tuned for what I call a query, not a conversation.”
Munjal Shah Aug 23, 2023 ▶ 18:52
Disclosure
Shah: Hippocratic AI is building its healthcare LLM from scratch
“That's why we're building our own model and we're building it from scratch.”
Munjal Shah Aug 23, 2023 ▶ 19:59
Assertion Supported
Shah: Hippocratic AI outperformed GPT-4 on 105 of 114 healthcare exams
“And then we took it, and then we had GPT-IV take it, and we had all the other language models take it, and we beat them all. And we beat them on a 105 of a 114 for GPT-IV, for example.”
Munjal Shah Aug 23, 2023 ▶ 21:10
Disclosure
Shah: Hippocratic AI acquired 1.5 trillion healthcare tokens for training
“We've actually gone in and acquired 1.5 trillion healthcare tokens.”
Munjal Shah Aug 23, 2023 ▶ 22:44
Prediction Not checkable as stated
Shah: US health systems will buy finished AI applications, not APIs
“Most of the health systems in the US are going to need a ready finished product slash application, not just have a an API.”
Munjal Shah Aug 23, 2023 ▶ 25:40
Disclosure
Shah: Hippocratic AI scraped PDFs for every US healthcare plan
“We've actually taken the time to go and get every single, 200 page PDF of every health care plan in the country.”
Munjal Shah Aug 23, 2023 ▶ 28:13
Insight
Shah: Healthcare LLMs require tone classification to detect pain and anger
“You wouldn't need an, a tone classifier for most interactions with an LLM, but you do in healthcare. Because a lot of the information is in the anger, the pain, the frustration.”
Munjal Shah Aug 23, 2023 ▶ 30:47
Insight
Shah: Deploying in low-risk applications is the best hallucination mitigation
“The number one way to deal with hallucinations is pick the right applications.”
Munjal Shah Aug 23, 2023 ▶ 31:35
Assertion Not checkable as stated
Shah: Roughly 6,900 of 7,000 US hospitals use standardized medical protocols
“If you look at the 7000 hospitals in the country, like, you know, 6900 will probably use the same protocol.”
Munjal Shah Aug 23, 2023 ▶ 34:23
Assertion Supported
Shah: Almost all FDA AI regulation to date targets diagnostic products
“Almost all your AI regulation to date from the FDA has been on diagnostic products.”
Munjal Shah Aug 23, 2023 ▶ 40:47
Disclosure
Shah: Hippocratic AI operates fully in person five days a week
“We're a hundred percent in person. Our company is literally five days a week”
Munjal Shah Aug 23, 2023 ▶ 44:42
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.