Jul 29, 2025 · 26m · latent-space

⚡️Using RFT to Build Clinical Superintelligence

Brendan Fortuna · 18m spoken Shawn Wang · 4m spoken Alessio Fanelli · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Lightning Pod, Ambience engineering leader Brendan Fortuna discusses how advanced Reinforcement Fine-Tuning (RFT), robust evaluation pipelines, and deep clinical domain integration are transforming generative AI from simple ambient scribing into a comprehensive clinical operating platform.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 7.3% of the talking time here. How this is scored →

The hosts as informed peer 4.8 Guest teaching 5.6 Guest disagreement 0.5 The hosts pushing back 0.9
05100:0010:0020:002:31–5:14 · The hosts as informed peer 4/10 Comparing Autonomous Driving Machine Learning to LLMs Wix compares autonomous driving ML workflows at Cruise with modern LLM-centric systems. Brendan explains the paradigm shift from massive labeling pipelines to sample-efficient fine-tuning while emphasizing that data engine fundamentals remain identical.5:14–8:36 · The hosts as informed peer 3/10 Understanding Reinforcement Fine-Tuning in Clinical Medicine Alessio asks how Ambience operationalized reinforcement fine-tuning (RFT). Brendan delivers a thorough technical breakdown of programmable graders, candidate trajectory generation, and optimizing for true downstream objectives rather than proxy loss.8:36–10:42 · The hosts as informed peer 5/10 Preventing Reward Hacking in Structured Physical Exam Generation Wix cautions against sample efficiency overfitting, prompting Brendan to detail real-world reward hacking incidents during physical exam generation. Brendan explains how the model inflated finding counts to boost precision and degraded clinical tone.10:42–14:17 · The hosts as informed peer 4/10 Benchmarking ICD-10 Medical Coding with RFT Alessio queries how Ambience evaluated complex ICD-10 medical coding. Brendan highlights the baseline 40% human physician F1 score versus their RFT-tuned o3-mini reaching 57% across 70,000 distinct codes.14:17–18:34 · The hosts as informed peer 6/10 Managing High Compute Costs in RFT Evaluator Pipelines Alessio brings up a war story about burning $25k on a single grader, and Wix references Braintrust's founder. Brendan discusses the balance between automated scripting, domain-expert IDEs, and ML engineers guiding intuition.18:34–21:10 · The hosts as informed peer 6/10 Clinical Hallucinations and Out-of-Distribution Healthcare Data Wix probes the nature of clinical hallucinations, and Brendan explains worst-case benchmark metrics and why EHR data constitutes an out-of-distribution domain. Wix connects this residency gap to insights shared by the OpenAI Codex team.21:10–24:33 · The hosts as informed peer 6/10 HealthBench, Data Privacy, and Developing Clinical Taste Alessio and Wix probe HealthBench, data privacy, and whether medical reasoning differs fundamentally from math/coding IQ. Brendan defines 'clinical taste' as knowing which noisy EHR elements to discard.24:33–26:44 · The hosts as informed peer 4/10 Autonomous Clinical AI Researchers and Ambience Hiring Alessio asks about building autonomous Devin-style clinical researchers. Brendan outlines their vision for autonomous evaluation loops and discusses the difficult archetype of clinicians with experimental engineering mindsets.2:31–5:14 · Guest teaching 5/10 Comparing Autonomous Driving Machine Learning to LLMs Wix compares autonomous driving ML workflows at Cruise with modern LLM-centric systems. Brendan explains the paradigm shift from massive labeling pipelines to sample-efficient fine-tuning while emphasizing that data engine fundamentals remain identical.5:14–8:36 · Guest teaching 7/10 Understanding Reinforcement Fine-Tuning in Clinical Medicine Alessio asks how Ambience operationalized reinforcement fine-tuning (RFT). Brendan delivers a thorough technical breakdown of programmable graders, candidate trajectory generation, and optimizing for true downstream objectives rather than proxy loss.8:36–10:42 · Guest teaching 6/10 Preventing Reward Hacking in Structured Physical Exam Generation Wix cautions against sample efficiency overfitting, prompting Brendan to detail real-world reward hacking incidents during physical exam generation. Brendan explains how the model inflated finding counts to boost precision and degraded clinical tone.10:42–14:17 · Guest teaching 6/10 Benchmarking ICD-10 Medical Coding with RFT Alessio queries how Ambience evaluated complex ICD-10 medical coding. Brendan highlights the baseline 40% human physician F1 score versus their RFT-tuned o3-mini reaching 57% across 70,000 distinct codes.14:17–18:34 · Guest teaching 5/10 Managing High Compute Costs in RFT Evaluator Pipelines Alessio brings up a war story about burning $25k on a single grader, and Wix references Braintrust's founder. Brendan discusses the balance between automated scripting, domain-expert IDEs, and ML engineers guiding intuition.18:34–21:10 · Guest teaching 7/10 Clinical Hallucinations and Out-of-Distribution Healthcare Data Wix probes the nature of clinical hallucinations, and Brendan explains worst-case benchmark metrics and why EHR data constitutes an out-of-distribution domain. Wix connects this residency gap to insights shared by the OpenAI Codex team.21:10–24:33 · Guest teaching 6/10 HealthBench, Data Privacy, and Developing Clinical Taste Alessio and Wix probe HealthBench, data privacy, and whether medical reasoning differs fundamentally from math/coding IQ. Brendan defines 'clinical taste' as knowing which noisy EHR elements to discard.24:33–26:44 · Guest teaching 3/10 Autonomous Clinical AI Researchers and Ambience Hiring Alessio asks about building autonomous Devin-style clinical researchers. Brendan outlines their vision for autonomous evaluation loops and discusses the difficult archetype of clinicians with experimental engineering mindsets.2:31–5:14 · Guest disagreement 0/10 Comparing Autonomous Driving Machine Learning to LLMs Wix compares autonomous driving ML workflows at Cruise with modern LLM-centric systems. Brendan explains the paradigm shift from massive labeling pipelines to sample-efficient fine-tuning while emphasizing that data engine fundamentals remain identical.5:14–8:36 · Guest disagreement 0/10 Understanding Reinforcement Fine-Tuning in Clinical Medicine Alessio asks how Ambience operationalized reinforcement fine-tuning (RFT). Brendan delivers a thorough technical breakdown of programmable graders, candidate trajectory generation, and optimizing for true downstream objectives rather than proxy loss.8:36–10:42 · Guest disagreement 1/10 Preventing Reward Hacking in Structured Physical Exam Generation Wix cautions against sample efficiency overfitting, prompting Brendan to detail real-world reward hacking incidents during physical exam generation. Brendan explains how the model inflated finding counts to boost precision and degraded clinical tone.10:42–14:17 · Guest disagreement 0/10 Benchmarking ICD-10 Medical Coding with RFT Alessio queries how Ambience evaluated complex ICD-10 medical coding. Brendan highlights the baseline 40% human physician F1 score versus their RFT-tuned o3-mini reaching 57% across 70,000 distinct codes.14:17–18:34 · Guest disagreement 1/10 Managing High Compute Costs in RFT Evaluator Pipelines Alessio brings up a war story about burning $25k on a single grader, and Wix references Braintrust's founder. Brendan discusses the balance between automated scripting, domain-expert IDEs, and ML engineers guiding intuition.18:34–21:10 · Guest disagreement 1/10 Clinical Hallucinations and Out-of-Distribution Healthcare Data Wix probes the nature of clinical hallucinations, and Brendan explains worst-case benchmark metrics and why EHR data constitutes an out-of-distribution domain. Wix connects this residency gap to insights shared by the OpenAI Codex team.21:10–24:33 · Guest disagreement 1/10 HealthBench, Data Privacy, and Developing Clinical Taste Alessio and Wix probe HealthBench, data privacy, and whether medical reasoning differs fundamentally from math/coding IQ. Brendan defines 'clinical taste' as knowing which noisy EHR elements to discard.24:33–26:44 · Guest disagreement 0/10 Autonomous Clinical AI Researchers and Ambience Hiring Alessio asks about building autonomous Devin-style clinical researchers. Brendan outlines their vision for autonomous evaluation loops and discusses the difficult archetype of clinicians with experimental engineering mindsets.2:31–5:14 · The hosts pushing back 1/10 Comparing Autonomous Driving Machine Learning to LLMs Wix compares autonomous driving ML workflows at Cruise with modern LLM-centric systems. Brendan explains the paradigm shift from massive labeling pipelines to sample-efficient fine-tuning while emphasizing that data engine fundamentals remain identical.5:14–8:36 · The hosts pushing back 0/10 Understanding Reinforcement Fine-Tuning in Clinical Medicine Alessio asks how Ambience operationalized reinforcement fine-tuning (RFT). Brendan delivers a thorough technical breakdown of programmable graders, candidate trajectory generation, and optimizing for true downstream objectives rather than proxy loss.8:36–10:42 · The hosts pushing back 2/10 Preventing Reward Hacking in Structured Physical Exam Generation Wix cautions against sample efficiency overfitting, prompting Brendan to detail real-world reward hacking incidents during physical exam generation. Brendan explains how the model inflated finding counts to boost precision and degraded clinical tone.10:42–14:17 · The hosts pushing back 0/10 Benchmarking ICD-10 Medical Coding with RFT Alessio queries how Ambience evaluated complex ICD-10 medical coding. Brendan highlights the baseline 40% human physician F1 score versus their RFT-tuned o3-mini reaching 57% across 70,000 distinct codes.14:17–18:34 · The hosts pushing back 1/10 Managing High Compute Costs in RFT Evaluator Pipelines Alessio brings up a war story about burning $25k on a single grader, and Wix references Braintrust's founder. Brendan discusses the balance between automated scripting, domain-expert IDEs, and ML engineers guiding intuition.18:34–21:10 · The hosts pushing back 1/10 Clinical Hallucinations and Out-of-Distribution Healthcare Data Wix probes the nature of clinical hallucinations, and Brendan explains worst-case benchmark metrics and why EHR data constitutes an out-of-distribution domain. Wix connects this residency gap to insights shared by the OpenAI Codex team.21:10–24:33 · The hosts pushing back 2/10 HealthBench, Data Privacy, and Developing Clinical Taste Alessio and Wix probe HealthBench, data privacy, and whether medical reasoning differs fundamentally from math/coding IQ. Brendan defines 'clinical taste' as knowing which noisy EHR elements to discard.24:33–26:44 · The hosts pushing back 0/10 Autonomous Clinical AI Researchers and Ambience Hiring Alessio asks about building autonomous Devin-style clinical researchers. Brendan outlines their vision for autonomous evaluation loops and discusses the difficult archetype of clinicians with experimental engineering mindsets.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 4.4% · guest 95.6%0:00 · the hosts 4.4% · guest 95.6%3:00 · the hosts 23.7% · guest 76.3%3:00 · the hosts 23.7% · guest 76.3%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 6.5% · guest 93.5%9:00 · the hosts 6.5% · guest 93.5%12:00 · the hosts 6.3% · guest 93.7%12:00 · the hosts 6.3% · guest 93.7%15:00 · the hosts 6.3% · guest 93.7%15:00 · the hosts 6.3% · guest 93.7%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 6.1% · guest 93.9%21:00 · the hosts 6.1% · guest 93.9%24:00 · the hosts 13.2% · guest 86.8%24:00 · the hosts 13.2% · guest 86.8%
Sharpest disagreement ▶ 22:38 Pushing back on clinical data quality assumptions

Brendan counters Wix's assumption that acquiring private clinical data solves modeling issues, emphasizing that realistic health records are remarkably messy and uncurated.

Hardest push from the hosts ▶ 8:36 Challenging sample efficiency in RL fine-tuning

Wix pushes back on the unconditional benefits of sample efficiency, citing risks where models over-index or extract incorrect lessons from minimal data points.

Biggest teaching moment ▶ 20:05 Explaining why clinical data is out-of-distribution

Brendan educates the hosts on how medical textbooks fail to capture the years of residency-level tribal knowledge and locked EHR documentation required for true clinical safety.

The host holds their own ▶ 20:48 Connecting residency data gaps to OpenAI Codex analogies

Wix demonstrates technical domain breadth by drawing an exact parallel between medical residency data gaps and on-the-job training lessons from the OpenAI Codex team.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Comparing Autonomous Driving Machine Learning to LLMs 4501 Wix compares autonomous driving ML workflows at Cruise with modern LLM-centric systems. Brendan explains the paradigm shift from massive labeling pipelines to sample-efficient fine-tuning while emphasizing that data engine fundamentals remain identical.
Understanding Reinforcement Fine-Tuning in Clinical Medicine 3700 Alessio asks how Ambience operationalized reinforcement fine-tuning (RFT). Brendan delivers a thorough technical breakdown of programmable graders, candidate trajectory generation, and optimizing for true downstream objectives rather than proxy loss.
Preventing Reward Hacking in Structured Physical Exam Generation 5612 Wix cautions against sample efficiency overfitting, prompting Brendan to detail real-world reward hacking incidents during physical exam generation. Brendan explains how the model inflated finding counts to boost precision and degraded clinical tone.
Benchmarking ICD-10 Medical Coding with RFT 4600 Alessio queries how Ambience evaluated complex ICD-10 medical coding. Brendan highlights the baseline 40% human physician F1 score versus their RFT-tuned o3-mini reaching 57% across 70,000 distinct codes.
Managing High Compute Costs in RFT Evaluator Pipelines 6511 Alessio brings up a war story about burning $25k on a single grader, and Wix references Braintrust's founder. Brendan discusses the balance between automated scripting, domain-expert IDEs, and ML engineers guiding intuition.
Clinical Hallucinations and Out-of-Distribution Healthcare Data 6711 Wix probes the nature of clinical hallucinations, and Brendan explains worst-case benchmark metrics and why EHR data constitutes an out-of-distribution domain. Wix connects this residency gap to insights shared by the OpenAI Codex team.
HealthBench, Data Privacy, and Developing Clinical Taste 6612 Alessio and Wix probe HealthBench, data privacy, and whether medical reasoning differs fundamentally from math/coding IQ. Brendan defines 'clinical taste' as knowing which noisy EHR elements to discard.
Autonomous Clinical AI Researchers and Ambience Hiring 4300 Alessio asks about building autonomous Devin-style clinical researchers. Brendan outlines their vision for autonomous evaluation loops and discusses the difficult archetype of clinicians with experimental engineering mindsets.

Statements from this episode (19)

Disclosure
Fortuna: Ambience sells to Cleveland Clinic, UCSF, Ardent, and John Muir
“Our customers are actually, like, big health systems. So, like, Cleveland Clinic and UCSF, Ardent, Sean Muir, these are the kind of customers we sell to”
Brendan Fortuna Jul 29, 2025 ▶ 1:01
Assertion Not checkable as stated
Fortuna: Ambience saves clinicians up to two hours per day
“And the end result is like, we'll save doctors, you know, up to two hours a day.”
Brendan Fortuna Jul 29, 2025 ▶ 1:39
Assertion Supported
Epic controls over 50% and Oracle Cerner holds 27% of EHR market
“Epic is obviously like the Goliath in the room, over 50% share. Cerner, that's another really big one owned by Oracle, maybe about 27%.”
Brendan Fortuna Jul 29, 2025 ▶ 2:12
Insight
Fortuna: Ambient Scribing Is Only Five Percent of AI's Healthcare Value
“That's, we listen to audio and we take notes, but that's actually just like, I would say five percent of like the total value that we can kind of offer, you know, once you have the audio of a conversation, right. And once you have access to the EHR and you can…”
Brendan Fortuna Jul 29, 2025 ▶ 4:30
Disclosure
Ambience AI Stack Uses Prompting, RAG, SFT, and RFT
“Ambience internally, we use prompting, we use chaining, we'll use RAG, we use fine tuning. It will use SFT and RFT.”
Brendan Fortuna Jul 29, 2025 ▶ 5:07
Disclosure
Fortuna: Ambience uses OpenAI's platform for reinforcement fine-tuning
“Our foray into RFT has primarily been through the OpenAI kind of platform. So we're using those self-service, you know, APIs.”
Brendan Fortuna Jul 29, 2025 ▶ 6:09
Insight
Fortuna: RFT is tremendously sample efficient compared to supervised fine-tuning
“And the second I think is like, it's tremendously sample efficient. Right. Each example sort of blooms into dozens of labels, right. And trajectories. So you can squeeze like X more signal right out of a dataset.”
Brendan Fortuna Jul 29, 2025 ▶ 8:20
Insight
Fortuna: LLM Graders for Prose Generation Are Highly Vulnerable to Reward Hacking
“And whenever using like an LLM grader, the task is like a little bit more pros or a little longer form generation. You could be very vulnerable to this. The models are super clever. They're incentivized to win, but they'll cheat and they'll do weird things.”
Brendan Fortuna Jul 29, 2025 ▶ 9:16
Insight
Fortuna: Weighting Graders 75% Accuracy and 25% Style Mitigates Reward Hacking
“So what we did is in the grader, you know, in addition to just the content and like the semantic accuracy of what it's saying, we also started to add style. And we kind of weight them like 75, 25, and over time you can kind of harness and get the reward hackin…”
Brendan Fortuna Jul 29, 2025 ▶ 10:28
Assertion Partly supported
RFT boosted o3-mini to 57% F1 on medical coding versus clinicians' 40%
“And the clinicians using F-one score were scoring like, let's say around 40%, right, on the F-one score, which was surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model O-three mini up …”
Brendan Fortuna Jul 29, 2025 ▶ 12:03
Assertion Supported
Fortuna: Oncologists and cardiologists spend up to 60 minutes pre-reviewing charts
“In certain specialties like oncology or cardiology, before they go into the visit with the patient, they can often spend like 10 minutes or up 30 minutes or 60 minutes looking at the chart. They're going to look at labs and images and, you know, other data abo…”
Brendan Fortuna Jul 29, 2025 ▶ 13:34
Insight
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Brendan Fortuna Jul 29, 2025 ▶ 14:43
Disclosure
Ambience Healthcare uses Braintrust for on-premise LLM observability
“One of the cool tools that we use internally is Braintrust. I think they're doing some incredible work over there, building like a tool for domain experts. They do give some built in observability. They let you deploy on premise. It's a fantastic technology. H…”
Brendan Fortuna Jul 29, 2025 ▶ 15:28
Insight
Fortuna: Domain experts cannot replace ML engineers in LLM fine-tuning
“I think the domain experts, like in our case, clinicians, they're really good at like debugging model outputs, meeting with users, distilling that feedback into something actionable, maybe annotating or doing evals, but they don't necessarily have like, you kn…”
Brendan Fortuna Jul 29, 2025 ▶ 17:16
Insight
Fortuna: Healthcare AI evaluation must measure worst-case failures over best-of-N
“I think one of the problems is, like, the benchmarks, they always report, like, the best event. So they report, like, you know, what's the best metric they got out of 64 attempts? But in healthcare, we're more interested in, like, what's the worst event, right…”
Brendan Fortuna Jul 29, 2025 ▶ 19:09
Insight
Fortuna: Base LLMs unsafely infer unconfirmed diagnoses from patient symptoms
“Patient, you know, they start to make these medical inferences. They're so smart, but they start to, like, infer things that the doctor didn't actually explicitly say. You know, for instance, a patient will say, like, I'm feeling sad and stressed out. Difficul…”
Brendan Fortuna Jul 29, 2025 ▶ 19:39
Insight
Fortuna: Real-world clinical data is out of distribution for base AI models
“I actually think clinical real world clinical data is out of distribution. I think as much as the model is generalized, if you have no access to that data, it's really hard to learn. I think that the reasons are maybe twofold. The first is a lot of realistic c…”
Brendan Fortuna Jul 29, 2025 ▶ 20:03
Insight
Fortuna: Saturated academic medical benchmarks are no longer useful for LLMs
“These academic data sets that we've been kind of saturating for a long time are no longer useful. Models are gonna ace medical exams.”
Brendan Fortuna Jul 29, 2025 ▶ 21:43
Assertion Supported
Fortuna: New reasoning models show no big leap on medical coding tasks
“I do know that when you kind of plot out base model performance on some medical tasks like ICD-X coding between like, you know, previous generations and new reasoning generations, there's actually not like a big leap.”
Brendan Fortuna Jul 29, 2025 ▶ 24:01
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.