Reinforcement Fine Tuning
topic on 4 shows · 16 statements across 10 episodes
BG2 Pod
Latent Space
No Priors
the Official SaaStr Podcast
15 statements about Reinforcement Fine Tuning, every show
Wu: Reinforcement fine-tuning is an order of magnitude more powerful than SFT
“Reinforcement fine tuning introduces like RL or reinforcement learning to this loop. Way more complex, way more finicky, but an order of magnitude more powerful.”
Godement: Reinforcement fine-tuning will become the norm for pushing AI capability frontiers
“Pushing the frontier on, like, actual capabilities, my hunch is that RFT will pretty much become the norm. Like, you know, if you are actually pushing in your field, like, you know, intelligence, like, you know, to a pretty high point, like, at some point, lik…”
OpenAI Reinforcement Fine-Tuning Makes Domain-Specific Models Viable for Vertical Apps
“In the last few months, OpenAI shipped reinforcement fine tuning, which is kind of a new way of fine tuning that they offer. And it's actually quite good. And so it actually is starting to seem again being able at least to not pre-train, but fine tune data on …”
Ambience AI Stack Uses Prompting, RAG, SFT, and RFT
“Ambience internally, we use prompting, we use chaining, we'll use RAG, we use fine tuning. It will use SFT and RFT.”
Fortuna: RFT is tremendously sample efficient compared to supervised fine-tuning
“And the second I think is like, it's tremendously sample efficient. Right. Each example sort of blooms into dozens of labels, right. And trajectories. So you can squeeze like X more signal right out of a dataset.”
RFT boosted o3-mini to 57% F1 on medical coding versus clinicians' 40%
“And the clinicians using F-one score were scoring like, let's say around 40%, right, on the F-one score, which was surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model O-three mini up …”
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Hsu: Top Curriculum-Writing AI Agents Will Likely Rely on Reinforcement Fine-Tuning
“And I also think like in the future, a really good curriculum or lesson writer agent will probably be like reinforcement fine-tuned on a lot of our internal data as well.”
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Fulford: RFT is only worthwhile for out-of-distribution or make-or-break tasks
“I think if you have a very specific task that you think is so different to anything that the model was likely trained on and you try it a bunch of times yourself and you've tried a lot of different prompts and it's just really not good at it. So maybe it's gen…”
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Foody: Reinforcement fine-tuning makes application-layer AI customization viable
“The reason I'm so optimistic about it taking off is that it's, like, profoundly data efficient, right? And it finally makes sense to customize models at the application layer.”
OpenAI plans to connect agent traces to evals and reinforcement fine-tuning
“Like you got to tie the traces to the evals product so that you can generate good evals. Once you have good evals and graders and tasks, You can use that to do reinforcement fine tuning and you know, lots of details to be figured out over here, but that's the …”
Lambert: Reinforcement fine-tuning will succeed where answer correctness matters over style
“It is just a new paradigm for fine tuning, and I have seen some of this work, and I'm pretty Optimistic that it'll work for kind of kind of really specific capabilities where answers matter rather than features in your style of text mattering.”