Jan 2, 2019 · 29m · a16z

a16z Podcast | Putting AI in Medicine, in Practice

Mintu Turakhia · 12m spoken Brandon Ballinger · 9m spoken Vijay Pande · 2m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z podcast, host Hannah and industry experts explore the practical integration of artificial intelligence into clinical medicine, discussing technical capabilities, reimbursement incentives, and operational strategies for successful adoption.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 2.1 Guest teaching 4.0 Guest disagreement 1.9 The host pushing back 1.4
05100:0010:0020:000:59–4:24 · The host as informed peer 2/10 History of Medical AI and System Incentives The host facilitates a discussion on historical AI medical systems like MYCIN and financial incentives in healthcare. She asks helpful follow-ups and mildly questions how misdiagnosis incentives work in practice, while the guests explain system reimbursement structures.4:24–9:36 · The host as informed peer 2/10 Diagnostic Applications and Continuous Wearable Data Guests outline diagnostic applications, continuous wearable monitoring, and primary care access gaps. The host tracks the conversation, synthesizing key insights such as missing data acting as a health indicator.9:36–14:12 · The host as informed peer 3/10 AI Autonomy Human Error and Liability Mintu offers a counter-perspective that data collection rather than AI is the primary hurdle for early detection. The host presses him on how his view differs, leading into a discussion on AI recapitulating human error and levels of autonomy.14:12–19:18 · The host as informed peer 2/10 Multi-Modal Data Synthesis and Model Overfitting A notable technical debate occurs between guests when Mintu defines machine learning as statistical overfitting and Vijay directly corrects him. The host helps unpack the concept of generalization using the guest's classroom analogy.19:18–21:53 · The host as informed peer 2/10 Digital Health Research and System Adoption Brandon details the high dropout rates in mobile health studies and the necessity of interdisciplinary teams. The conversation is collaborative and informational as the host prompts questions on system incentives.21:53–23:58 · The host as informed peer 2/10 AI Model Versioning and Deployment Mintu raises a technical question regarding continuous learning versus batch versioning in regulated health AI. Vijay and Brandon explain holdout validation sets and speech recognition deployment practices.23:58–29:34 · The host as informed peer 2/10 Quality Improvement and Clinical Decision Support Mintu explains how AI can assist quality improvement in EKG reading and streamline primitive operating room scheduling workflows. The host reacts to hospital scheduling practices and concludes the interview.0:59–4:24 · Guest teaching 4/10 History of Medical AI and System Incentives The host facilitates a discussion on historical AI medical systems like MYCIN and financial incentives in healthcare. She asks helpful follow-ups and mildly questions how misdiagnosis incentives work in practice, while the guests explain system reimbursement structures.4:24–9:36 · Guest teaching 3/10 Diagnostic Applications and Continuous Wearable Data Guests outline diagnostic applications, continuous wearable monitoring, and primary care access gaps. The host tracks the conversation, synthesizing key insights such as missing data acting as a health indicator.9:36–14:12 · Guest teaching 5/10 AI Autonomy Human Error and Liability Mintu offers a counter-perspective that data collection rather than AI is the primary hurdle for early detection. The host presses him on how his view differs, leading into a discussion on AI recapitulating human error and levels of autonomy.14:12–19:18 · Guest teaching 5/10 Multi-Modal Data Synthesis and Model Overfitting A notable technical debate occurs between guests when Mintu defines machine learning as statistical overfitting and Vijay directly corrects him. The host helps unpack the concept of generalization using the guest's classroom analogy.19:18–21:53 · Guest teaching 3/10 Digital Health Research and System Adoption Brandon details the high dropout rates in mobile health studies and the necessity of interdisciplinary teams. The conversation is collaborative and informational as the host prompts questions on system incentives.21:53–23:58 · Guest teaching 4/10 AI Model Versioning and Deployment Mintu raises a technical question regarding continuous learning versus batch versioning in regulated health AI. Vijay and Brandon explain holdout validation sets and speech recognition deployment practices.23:58–29:34 · Guest teaching 4/10 Quality Improvement and Clinical Decision Support Mintu explains how AI can assist quality improvement in EKG reading and streamline primitive operating room scheduling workflows. The host reacts to hospital scheduling practices and concludes the interview.0:59–4:24 · Guest disagreement 1/10 History of Medical AI and System Incentives The host facilitates a discussion on historical AI medical systems like MYCIN and financial incentives in healthcare. She asks helpful follow-ups and mildly questions how misdiagnosis incentives work in practice, while the guests explain system reimbursement structures.4:24–9:36 · Guest disagreement 1/10 Diagnostic Applications and Continuous Wearable Data Guests outline diagnostic applications, continuous wearable monitoring, and primary care access gaps. The host tracks the conversation, synthesizing key insights such as missing data acting as a health indicator.9:36–14:12 · Guest disagreement 3/10 AI Autonomy Human Error and Liability Mintu offers a counter-perspective that data collection rather than AI is the primary hurdle for early detection. The host presses him on how his view differs, leading into a discussion on AI recapitulating human error and levels of autonomy.14:12–19:18 · Guest disagreement 5/10 Multi-Modal Data Synthesis and Model Overfitting A notable technical debate occurs between guests when Mintu defines machine learning as statistical overfitting and Vijay directly corrects him. The host helps unpack the concept of generalization using the guest's classroom analogy.19:18–21:53 · Guest disagreement 1/10 Digital Health Research and System Adoption Brandon details the high dropout rates in mobile health studies and the necessity of interdisciplinary teams. The conversation is collaborative and informational as the host prompts questions on system incentives.21:53–23:58 · Guest disagreement 1/10 AI Model Versioning and Deployment Mintu raises a technical question regarding continuous learning versus batch versioning in regulated health AI. Vijay and Brandon explain holdout validation sets and speech recognition deployment practices.23:58–29:34 · Guest disagreement 1/10 Quality Improvement and Clinical Decision Support Mintu explains how AI can assist quality improvement in EKG reading and streamline primitive operating room scheduling workflows. The host reacts to hospital scheduling practices and concludes the interview.0:59–4:24 · The host pushing back 2/10 History of Medical AI and System Incentives The host facilitates a discussion on historical AI medical systems like MYCIN and financial incentives in healthcare. She asks helpful follow-ups and mildly questions how misdiagnosis incentives work in practice, while the guests explain system reimbursement structures.4:24–9:36 · The host pushing back 1/10 Diagnostic Applications and Continuous Wearable Data Guests outline diagnostic applications, continuous wearable monitoring, and primary care access gaps. The host tracks the conversation, synthesizing key insights such as missing data acting as a health indicator.9:36–14:12 · The host pushing back 3/10 AI Autonomy Human Error and Liability Mintu offers a counter-perspective that data collection rather than AI is the primary hurdle for early detection. The host presses him on how his view differs, leading into a discussion on AI recapitulating human error and levels of autonomy.14:12–19:18 · The host pushing back 1/10 Multi-Modal Data Synthesis and Model Overfitting A notable technical debate occurs between guests when Mintu defines machine learning as statistical overfitting and Vijay directly corrects him. The host helps unpack the concept of generalization using the guest's classroom analogy.19:18–21:53 · The host pushing back 1/10 Digital Health Research and System Adoption Brandon details the high dropout rates in mobile health studies and the necessity of interdisciplinary teams. The conversation is collaborative and informational as the host prompts questions on system incentives.21:53–23:58 · The host pushing back 1/10 AI Model Versioning and Deployment Mintu raises a technical question regarding continuous learning versus batch versioning in regulated health AI. Vijay and Brandon explain holdout validation sets and speech recognition deployment practices.23:58–29:34 · The host pushing back 1/10 Quality Improvement and Clinical Decision Support Mintu explains how AI can assist quality improvement in EKG reading and streamline primitive operating room scheduling workflows. The host reacts to hospital scheduling practices and concludes the interview.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 15:00 Vijay rejects Mintu's definition of machine learning

Vijay directly interrupts and corrects Mintu's assertion that machine learning is by definition statistical overfitting, clarifying that proper ML actively avoids overfitting through regularization.

Hardest push from the host ▶ 9:36 Host challenges Mintu on AI versus data collection

Hannah pushes back when Mintu dissents on AI's necessity for early detection, demanding to know how his view differs from theirs and why earlier detection wouldn't be inherently beneficial.

Biggest teaching moment ▶ 2:07 Brandon explains why superior AI systems fail to deploy

Brandon educates the host on system-level deployment barriers, using the 1978 MYCIN case study to show how technical superiority over physicians fails to translate to hospital adoption without financial reimbursement alignment.

The host holds their own ▶ 3:18 Host presses financial logic behind misdiagnosis incentives

Hannah challenges Brandon's claim about hospital incentives by highlighting the moral and professional contradiction that no doctor or hospital actually wants incorrect diagnoses, forcing him to clarify fee-for-value nuance.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
History of Medical AI and System Incentives 2412 The host facilitates a discussion on historical AI medical systems like MYCIN and financial incentives in healthcare. She asks helpful follow-ups and mildly questions how misdiagnosis incentives work in practice, while the guests explain system reimbursement structures.
Diagnostic Applications and Continuous Wearable Data 2311 Guests outline diagnostic applications, continuous wearable monitoring, and primary care access gaps. The host tracks the conversation, synthesizing key insights such as missing data acting as a health indicator.
AI Autonomy Human Error and Liability 3533 Mintu offers a counter-perspective that data collection rather than AI is the primary hurdle for early detection. The host presses him on how his view differs, leading into a discussion on AI recapitulating human error and levels of autonomy.
Multi-Modal Data Synthesis and Model Overfitting 2551 A notable technical debate occurs between guests when Mintu defines machine learning as statistical overfitting and Vijay directly corrects him. The host helps unpack the concept of generalization using the guest's classroom analogy.
Digital Health Research and System Adoption 2311 Brandon details the high dropout rates in mobile health studies and the necessity of interdisciplinary teams. The conversation is collaborative and informational as the host prompts questions on system incentives.
AI Model Versioning and Deployment 2411 Mintu raises a technical question regarding continuous learning versus batch versioning in regulated health AI. Vijay and Brandon explain holdout validation sets and speech recognition deployment practices.
Quality Improvement and Clinical Decision Support 2411 Mintu explains how AI can assist quality improvement in EKG reading and streamline primitive operating room scheduling workflows. The host reacts to hospital scheduling practices and concludes the interview.

Statements from this episode (20)

Assertion Partly supported
Ballinger: Stanford's 1978 MYCIN AI system outperformed five pathologists
“An interesting case study is, is the Meissen system, which is from 1978, I believe, and so this was an expert system trained at Stanford. It would take inputs that were just typed in manually, and then it would essentially try to predict what a pathologist wou…”
Brandon Ballinger Jan 2, 2019 ▶ 2:07
Insight
Ballinger: Misdiagnosis generates more hospital revenue under fee-for-service models
“If you think about kind of a hospital from the CFO's perspective misdiagnosis actually earns them more money. Because when you misdiagnose, you do follow-up tests, right? And those, and our billing system is fee-for-service. So every little test that's done is…”
Brandon Ballinger Jan 2, 2019 ▶ 2:58
Prediction Not checkable as stated
Ballinger: Fee-for-value payment models will enable medical AI adoption
“Things like fee for value are interesting because now you're paying people for, say, an accurate diagnosis or for a reduction in hospitalizations, depending on the exact system, and so I think that's the case where actually accuracy is rewarded with greater pa…”
Brandon Ballinger Jan 2, 2019 ▶ 3:33
Assertion Supported
Ballinger: Doctors cannot diagnose sleep apnea from Fitbit data alone
“No doctor right now can read your Fitbit data and tell you whether you have a condition like sleep apnea.”
Brandon Ballinger Jan 2, 2019 ▶ 4:50
Assertion Not checkable as stated
Turakhia: AI cannot predict acute heart attacks days in advance
“You can predict a cumulative probability, like a probability of getting condition X or diagnosis X over a time horizon of five or 10 years. But we are nowhere near saying, you know, you're going to have a heart attack in the next three days.”
Mintu Turakhia Jan 2, 2019 ▶ 5:38
Insight
Turakhia: Missing wearable data is the strongest predictor of illness
“In fact, the biggest predictor, Of someone getting ill with a lot of wearable studies is missing data because they were too sick to wear the sensor.”
Mintu Turakhia Jan 2, 2019 ▶ 6:16
Assertion Contradicted
Ballinger: Most Americans do not get an EKG until age 65
“Most people actually in the U.S., you get your first D.K.G. When you turn 65 as part of your Medicare checkup.”
Brandon Ballinger Jan 2, 2019 ▶ 8:33
Assertion Contradicted
Ballinger: Only half of Americans have a primary care physician
“And if you look at that in aggregate, about half of people in the U.S. Have a primary care physician at all, which seems astonishingly low, but that's kind of the fact.”
Brandon Ballinger Jan 2, 2019 ▶ 8:58
Assertion Partly supported
Ballinger: Large percentages of common medical conditions remain undiagnosed
“About a third of people with diabetes don't realize they have it. About a fifth of people with hypertension. For AFib, it's 30 or 40%. For sleep apnea, it's like 80%.”
Brandon Ballinger Jan 2, 2019 ▶ 9:08
Assertion Supported
Turakhia: Neural networks replicate human error patterns in EKG and imaging studies
“Some of the most promising aspects of the imaging studies and the EKG studies are that the confusion matrices, the way humans misclassify things is recapitulated by the convolutional neural networks.”
Mintu Turakhia Jan 2, 2019 ▶ 10:49
Opinion
Turakhia: Fully autonomous medical AI faces societal, not technical, barriers
“That's a societal issue. That's not a technical hurdle at this point.”
Mintu Turakhia Jan 2, 2019 ▶ 12:43
Prediction Not publicly verifiable
Ballinger: Wearables will generate two trillion data points in 2017
“Wearables are an interesting case because they'll generate about Two trillion data points this year.”
Brandon Ballinger Jan 2, 2019 ▶ 13:46
Assertion Supported
Turakhia: Clinical documentation terminology varies between individual hospitals
“In natural language processing that's embedded in AI, the lexicon that people use, how doctors and clinicians write what it is that they're seeing with their patient is different from not even specialty to specialty, but hospital to hospital, sort of mini subc…”
Mintu Turakhia Jan 2, 2019 ▶ 16:50
Assertion Partly supported
Ballinger: Wearable device data like Fitbit data is globally standardized
“I think that's actually a nice thing about wearable data is that Fitbits are the same all over the world.”
Brandon Ballinger Jan 2, 2019 ▶ 17:10
Assertion Supported
Ballinger: Early Apple ResearchKit apps lost 90% of participants in 90 days
“The challenge that the first five ResearchKit apps had is that they got 40,000 people, and then they lost 90% of them in the first 90 days.”
Brandon Ballinger Jan 2, 2019 ▶ 19:39
Opinion
Turakhia: FDA's Digital Health Office effectively mitigates regulatory risk
“The regulatory risk thing is being largely addressed by this new Office of Digital Health and the FDA, and they're really doing, seem much more forward thinking about it.”
Mintu Turakhia Jan 2, 2019 ▶ 21:46
Insight
Turakhia: Continuous learning in medical AI risks patient harm from biased data
“Bad data could heavily bias the system and cause harm, right? So if you start learning from bad inputs that come into the system for whatever reason, you could intentionally or unintentionally, you know, cause harm.”
Mintu Turakhia Jan 2, 2019 ▶ 22:21
Assertion Supported
Turakhia: No standardized quality improvement metrics exist for EKG interpretation
“There's actually no standardized metrics for QI in any of this.”
Mintu Turakhia Jan 2, 2019 ▶ 24:35
Prediction Not checkable as stated
Ballinger: Full-stack healthcare provider startups will achieve fastest adoption
“I think that's a model that we're going to see probably get the quickest adoption.”
Brandon Ballinger Jan 2, 2019 ▶ 28:54
Prediction Not checkable as stated
Ballinger: Healthcare will reorganize vertically around specialized AI data networks
“So this is a case where probably we'll see the healthcare industry maybe reconstitute itself by vertical. With AI-based diagnostics or therapeutics, because if you think right now providers are geographically structured, but with AI, every data point makes the…”
Brandon Ballinger Jan 2, 2019 ▶ 29:07
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.