reinforcement fine-tuning

also referred to as: reinforcement fine tuning · rft

11 statements across 6 episodes · 7 bullish · 0 bearish · 6 people on the record · first statement Jan 2, 2025 by Nathan Lambert · across every show →

Everything said about reinforcement fine-tuning, oldest first

Jan 2, 2025 positive
Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Nathan Lambert Jan 2, 2025 ▶ 9:57 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Jan 2, 2025 bullish
Opinion
Lambert: Reinforcement fine-tuning will succeed where answer correctness matters over style
“It is just a new paradigm for fine tuning, and I have seen some of this work, and I'm pretty Optimistic that it'll work for kind of kind of really specific capabilities where answers matter rather than features in your style of text mattering.”
Nathan Lambert Jan 2, 2025 ▶ 9:18 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
Mar 11, 2025 positive
Disclosure
OpenAI plans to connect agent traces to evals and reinforcement fine-tuning
“Like you got to tie the traces to the evals product so that you can generate good evals. Once you have good evals and graders and tasks, You can use that to do reinforcement fine tuning and you know, lots of details to be figured out over here, but that's the …”
Nikunj Handa Mar 11, 2025 ▶ 24:44 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Apr 15, 2025 neutral
Assertion Supported
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Michelle Pokrass Apr 15, 2025 ▶ 39:32 GPT 4.1: The New OpenAI Workhorse
Jun 19, 2025 bullish
Insight
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Noam Brown Jun 19, 2025 ▶ 21:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
Jul 11, 2025 bullish
Prediction Not checkable as stated
Hsu: Top Curriculum-Writing AI Agents Will Likely Rely on Reinforcement Fine-Tuning
“And I also think like in the future, a really good curriculum or lesson writer agent will probably be like reinforcement fine-tuned on a lot of our internal data as well.”
Andrew Hsu Jul 11, 2025 ▶ 34:32 Personalized AI Language Education — with Andrew Hsu, Speak
Jul 29, 2025 positive
Insight
Fortuna: RFT is tremendously sample efficient compared to supervised fine-tuning
“And the second I think is like, it's tremendously sample efficient. Right. Each example sort of blooms into dozens of labels, right. And trajectories. So you can squeeze like X more signal right out of a dataset.”
Brendan Fortuna Jul 29, 2025 ▶ 8:20 ⚡️Using RFT to Build Clinical Superintelligence
Jul 29, 2025
Insight
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Brendan Fortuna Jul 29, 2025 ▶ 14:43 ⚡️Using RFT to Build Clinical Superintelligence
Jul 29, 2025
Disclosure
Ambience AI Stack Uses Prompting, RAG, SFT, and RFT
“Ambience internally, we use prompting, we use chaining, we'll use RAG, we use fine tuning. It will use SFT and RFT.”
Brendan Fortuna Jul 29, 2025 ▶ 5:07 ⚡️Using RFT to Build Clinical Superintelligence
Jul 29, 2025 bullish
Assertion Partly supported
RFT boosted o3-mini to 57% F1 on medical coding versus clinicians' 40%
“And the clinicians using F-one score were scoring like, let's say around 40%, right, on the F-one score, which was surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model O-three mini up …”
Brendan Fortuna Jul 29, 2025 ▶ 12:03 ⚡️Using RFT to Build Clinical Superintelligence
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.