Insight certainty 4/5 debate potential 3/5

Fortuna: Healthcare AI evaluation must measure worst-case failures over best-of-N

Brendan Fortuna · ⚡️Using RFT to Build Clinical Superintelligence · Jul 29, 2025 · at 19:09

Brendan Fortuna, Head of Engineering at Ambience AI, explains why standard AI benchmark metrics fail to capture the safety requirements needed in healthcare.

0:00 / 0:11exact quote · 11.8s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right?”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Brendan Fortuna

Insight
Fortuna: Domain experts cannot replace ML engineers in LLM fine-tuning
“I think the domain experts, like in our case, clinicians, they're really good at like debugging model outputs, meeting with users, distilling that feedback into something actionable, maybe annotating or doing evals, but they don't necessarily have like, you kn…”
Brendan Fortuna Jul 29, 2025 ▶ 17:16 ⚡️Using RFT to Build Clinical Superintelligence
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Brendan Fortuna Jul 29, 2025 ▶ 24:01 ⚡️Using RFT to Build Clinical Superintelligence
Insight
Fortuna: Weighting Graders 75% Accuracy and 25% Style Mitigates Reward Hacking
“So what we did is in the grader, you know, in addition to just the content and like the semantic accuracy of what it's saying, we also started to add style. And we kind of weight them like 75, 25, and over time you can kind of harness and get the reward hackin…”
Brendan Fortuna Jul 29, 2025 ▶ 10:28 ⚡️Using RFT to Build Clinical Superintelligence
Assertion Partly supported
RFT boosted o3-mini to 57% F1 on medical coding versus clinicians' 40%
“And the clinicians using F-one score were scoring like, let's say around 40%, right, on the F-one score, which was surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model O-three mini up …”
Brendan Fortuna Jul 29, 2025 ▶ 12:03 ⚡️Using RFT to Build Clinical Superintelligence
Insight
Fortuna: Base LLMs unsafely infer unconfirmed diagnoses from patient symptoms
“Patient, you know, they start to make these medical inferences. They're so smart, but they start to, like, infer things that the doctor didn't actually explicitly say. You know, for instance, a patient will say, like, I'm feeling sad and stressed out. Difficul…”
Brendan Fortuna Jul 29, 2025 ▶ 19:39 ⚡️Using RFT to Build Clinical Superintelligence
Insight
Fortuna: Real-world clinical data is out of distribution for base AI models
“I actually think clinical real world clinical data is out of distribution. I think as much as the model is generalized, if you have no access to that data, it's really hard to learn. I think that the reasons are maybe twofold. The first is a lot of realistic c…”
Brendan Fortuna Jul 29, 2025 ▶ 20:03 ⚡️Using RFT to Build Clinical Superintelligence
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.