Jul 29, 2025 · 26m · latent-space
⚡️Using RFT to Build Clinical Superintelligence
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Latent Space Lightning Pod, Ambience engineering leader Brendan Fortuna discusses how advanced Reinforcement Fine-Tuning (RFT), robust evaluation pipelines, and deep clinical domain integration are transforming generative AI from simple ambient scribing into a comprehensive clinical operating platform.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 7.3% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Brendan counters Wix's assumption that acquiring private clinical data solves modeling issues, emphasizing that realistic health records are remarkably messy and uncurated.
Hardest push from the hosts ▶ 8:36 Challenging sample efficiency in RL fine-tuningWix pushes back on the unconditional benefits of sample efficiency, citing risks where models over-index or extract incorrect lessons from minimal data points.
Biggest teaching moment ▶ 20:05 Explaining why clinical data is out-of-distributionBrendan educates the hosts on how medical textbooks fail to capture the years of residency-level tribal knowledge and locked EHR documentation required for true clinical safety.
The host holds their own ▶ 20:48 Connecting residency data gaps to OpenAI Codex analogiesWix demonstrates technical domain breadth by drawing an exact parallel between medical residency data gaps and on-the-job training lessons from the OpenAI Codex team.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Comparing Autonomous Driving Machine Learning to LLMs | 4 | 5 | 0 | 1 | Wix compares autonomous driving ML workflows at Cruise with modern LLM-centric systems. Brendan explains the paradigm shift from massive labeling pipelines to sample-efficient fine-tuning while emphasizing that data engine fundamentals remain identical. | |
| Understanding Reinforcement Fine-Tuning in Clinical Medicine | 3 | 7 | 0 | 0 | Alessio asks how Ambience operationalized reinforcement fine-tuning (RFT). Brendan delivers a thorough technical breakdown of programmable graders, candidate trajectory generation, and optimizing for true downstream objectives rather than proxy loss. | |
| Preventing Reward Hacking in Structured Physical Exam Generation | 5 | 6 | 1 | 2 | Wix cautions against sample efficiency overfitting, prompting Brendan to detail real-world reward hacking incidents during physical exam generation. Brendan explains how the model inflated finding counts to boost precision and degraded clinical tone. | |
| Benchmarking ICD-10 Medical Coding with RFT | 4 | 6 | 0 | 0 | Alessio queries how Ambience evaluated complex ICD-10 medical coding. Brendan highlights the baseline 40% human physician F1 score versus their RFT-tuned o3-mini reaching 57% across 70,000 distinct codes. | |
| Managing High Compute Costs in RFT Evaluator Pipelines | 6 | 5 | 1 | 1 | Alessio brings up a war story about burning $25k on a single grader, and Wix references Braintrust's founder. Brendan discusses the balance between automated scripting, domain-expert IDEs, and ML engineers guiding intuition. | |
| Clinical Hallucinations and Out-of-Distribution Healthcare Data | 6 | 7 | 1 | 1 | Wix probes the nature of clinical hallucinations, and Brendan explains worst-case benchmark metrics and why EHR data constitutes an out-of-distribution domain. Wix connects this residency gap to insights shared by the OpenAI Codex team. | |
| HealthBench, Data Privacy, and Developing Clinical Taste | 6 | 6 | 1 | 2 | Alessio and Wix probe HealthBench, data privacy, and whether medical reasoning differs fundamentally from math/coding IQ. Brendan defines 'clinical taste' as knowing which noisy EHR elements to discard. | |
| Autonomous Clinical AI Researchers and Ambience Hiring | 4 | 3 | 0 | 0 | Alessio asks about building autonomous Devin-style clinical researchers. Brendan outlines their vision for autonomous evaluation loops and discusses the difficult archetype of clinicians with experimental engineering mindsets. |