Brendan Fortuna (Ambience Healthcare) addresses whether AI engineering requires machine learning engineers or solely domain expert clinicians.
Insight
Fortuna: Healthcare AI evaluation must measure worst-case failures over best-of-N
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right…”
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Insight
Fortuna: Weighting Graders 75% Accuracy and 25% Style Mitigates Reward Hacking
“So what we did is in the grader, you know, in addition to just the content and like the semantic accuracy of what it's saying, we also started to add style. And we kind of weight them like 75, 25, and over time you can kind of harness and get the reward hackin…”
Assertion Partly supported
RFT boosted o3-mini to 57% F1 on medical coding versus clinicians' 40%
“And the clinicians using F-one score were scoring like, let's say around 40%, right, on the F-one score, which was surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model O-three mini up …”
Insight
Fortuna: Base LLMs unsafely infer unconfirmed diagnoses from patient symptoms
“Patient, you know, they start to make these medical inferences. They're so smart, but they start to, like, infer things that the doctor didn't actually explicitly say. You know, for instance, a patient will say, like, I'm feeling sad and stressed out. Difficul…”
Insight
Fortuna: Real-world clinical data is out of distribution for base AI models
“I actually think clinical real world clinical data is out of distribution. I think as much as the model is generalized, if you have no access to that data, it's really hard to learn. I think that the reasons are maybe twofold. The first is a lot of realistic c…”