Everything Brendan Fortuna said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Fortuna: Domain experts cannot replace ML engineers in LLM fine-tuning
“I think the domain experts, like in our case, clinicians, they're really good at like debugging model outputs, meeting with users, distilling that feedback into something actionable, maybe annotating or doing evals, but they don't necessarily have like, you kn…”
Fortuna: Healthcare AI evaluation must measure worst-case failures over best-of-N
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right…”
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Fortuna: Weighting Graders 75% Accuracy and 25% Style Mitigates Reward Hacking
“So what we did is in the grader, you know, in addition to just the content and like the semantic accuracy of what it's saying, we also started to add style. And we kind of weight them like 75, 25, and over time you can kind of harness and get the reward hackin…”
RFT boosted o3-mini to 57% F1 on medical coding versus clinicians' 40%
“And the clinicians using F-one score were scoring like, let's say around 40%, right, on the F-one score, which was surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model O-three mini up …”
Fortuna: Base LLMs unsafely infer unconfirmed diagnoses from patient symptoms
“Patient, you know, they start to make these medical inferences. They're so smart, but they start to, like, infer things that the doctor didn't actually explicitly say. You know, for instance, a patient will say, like, I'm feeling sad and stressed out. Difficul…”
Fortuna: Real-world clinical data is out of distribution for base AI models
“I actually think clinical real world clinical data is out of distribution. I think as much as the model is generalized, if you have no access to that data, it's really hard to learn. I think that the reasons are maybe twofold. The first is a lot of realistic c…”
Fortuna: Ambience saves clinicians up to two hours per day
“And the end result is like, we'll save doctors, you know, up to two hours a day.”
Fortuna: Ambient Scribing Is Only Five Percent of AI's Healthcare Value
“That's, we listen to audio and we take notes, but that's actually just like, I would say five percent of like the total value that we can kind of offer, you know, once you have the audio of a conversation, right. And once you have access to the EHR and you can…”
Fortuna: RFT is tremendously sample efficient compared to supervised fine-tuning
“And the second I think is like, it's tremendously sample efficient. Right. Each example sort of blooms into dozens of labels, right. And trajectories. So you can squeeze like X more signal right out of a dataset.”
Fortuna: LLM Graders for Prose Generation Are Highly Vulnerable to Reward Hacking
“And whenever using like an LLM grader, the task is like a little bit more pros or a little longer form generation. You could be very vulnerable to this. The models are super clever. They're incentivized to win, but they'll cheat and they'll do weird things.”
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Fortuna: Saturated academic medical benchmarks are no longer useful for LLMs
“These academic data sets that we've been kind of saturating for a long time are no longer useful. Models are gonna ace medical exams.”
Fortuna: Ambience sells to Cleveland Clinic, UCSF, Ardent, and John Muir
“Our customers are actually, like, big health systems. So, like, Cleveland Clinic and UCSF, Ardent, Sean Muir, these are the kind of customers we sell to”
Epic controls over 50% and Oracle Cerner holds 27% of EHR market
“Epic is obviously like the Goliath in the room, over 50% share. Cerner, that's another really big one owned by Oracle, maybe about 27%.”
Ambience AI Stack Uses Prompting, RAG, SFT, and RFT
“Ambience internally, we use prompting, we use chaining, we'll use RAG, we use fine tuning. It will use SFT and RFT.”
Fortuna: Ambience uses OpenAI's platform for reinforcement fine-tuning
“Our foray into RFT has primarily been through the OpenAI kind of platform. So we're using those self-service, you know, APIs.”
Fortuna: Oncologists and cardiologists spend up to 60 minutes pre-reviewing charts
“In certain specialties like oncology or cardiology, before they go into the visit with the patient, they can often spend like 10 minutes or up 30 minutes or 60 minutes looking at the chart. They're going to look at labs and images and, you know, other data abo…”
Ambience Healthcare uses Braintrust for on-premise LLM observability
“One of the cool tools that we use internally is Braintrust. I think they're doing some incredible work over there, building like a tool for domain experts. They do give some built in observability. They let you deploy on premise. It's a fantastic technology. H…”