Reinforcement Fine Tuning

topic on 4 shows · 16 statements across 10 episodes

BG2 Pod Latent Space No Priors the Official SaaStr Podcast

15 statements about Reinforcement Fine Tuning, every show

BG2 Insight
Wu: Reinforcement fine-tuning is an order of magnitude more powerful than SFT
“Reinforcement fine tuning introduces like RL or reinforcement learning to this loop. Way more complex, way more finicky, but an order of magnitude more powerful.”
Sherwin Wu Sep 11, 2025 ▶ 40:09 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
BG2 Prediction Not checkable as stated
Godement: Reinforcement fine-tuning will become the norm for pushing AI capability frontiers
“Pushing the frontier on, like, actual capabilities, my hunch is that RFT will pretty much become the norm. Like, you know, if you are actually pushing in your field, like, you know, intelligence, like, you know, to a pretty high point, like, at some point, lik…”
Olivier Godman Sep 11, 2025 ▶ 42:10 Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview · Bg2 Pod
SAASTR Opinion
OpenAI Reinforcement Fine-Tuning Makes Domain-Specific Models Viable for Vertical Apps
“In the last few months, OpenAI shipped reinforcement fine tuning, which is kind of a new way of fine tuning that they offer. And it's actually quite good. And so it actually is starting to seem again being able at least to not pre-train, but fine tune data on …”
Jacob Effron Aug 13, 2025 ▶ 10:32 Redpoint Ventures Playbook: How Top VCs Are Really Investing in AI with Jacob Effron
LATENT SPACE Disclosure
Ambience AI Stack Uses Prompting, RAG, SFT, and RFT
“Ambience internally, we use prompting, we use chaining, we'll use RAG, we use fine tuning. It will use SFT and RFT.”
Brendan Fortuna Jul 29, 2025 ▶ 5:07 ⚡️Using RFT to Build Clinical Superintelligence
Fortuna: RFT is tremendously sample efficient compared to supervised fine-tuning
“And the second I think is like, it's tremendously sample efficient. Right. Each example sort of blooms into dozens of labels, right. And trajectories. So you can squeeze like X more signal right out of a dataset.”
Brendan Fortuna Jul 29, 2025 ▶ 8:20 ⚡️Using RFT to Build Clinical Superintelligence
LATENT SPACE Assertion Partly supported
RFT boosted o3-mini to 57% F1 on medical coding versus clinicians' 40%
“And the clinicians using F-one score were scoring like, let's say around 40%, right, on the F-one score, which was surprisingly low, lower than we thought. We were able to use RFT to kind of hill climb and get that, you know, get a small model O-three mini up …”
Brendan Fortuna Jul 29, 2025 ▶ 12:03 ⚡️Using RFT to Build Clinical Superintelligence
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Brendan Fortuna Jul 29, 2025 ▶ 14:43 ⚡️Using RFT to Build Clinical Superintelligence
LATENT SPACE Prediction Not checkable as stated
Hsu: Top Curriculum-Writing AI Agents Will Likely Rely on Reinforcement Fine-Tuning
“And I also think like in the future, a really good curriculum or lesson writer agent will probably be like reinforcement fine-tuned on a lot of our internal data as well.”
Andrew Hsu Jul 11, 2025 ▶ 34:32 Personalized AI Language Education — with Andrew Hsu, Speak
Brown: Data for reinforcement fine-tuning survives future model scaling
“I think the difference is that like for reinforcement fine tuning, you're collecting data that's going to be useful As the models improve as well. So if we come out with, like, future models that are even more capable, you could still fine tune them on your da…”
Noam Brown Jun 19, 2025 ▶ 21:26 Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI
NO PRIORS Insight
Fulford: RFT is only worthwhile for out-of-distribution or make-or-break tasks
“I think if you have a very specific task that you think is so different to anything that the model was likely trained on and you try it a bunch of times yourself and you've tried a lot of different prompts and it's just really not good at it. So maybe it's gen…”
Isa Fulford Apr 24, 2025 ▶ 7:50 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
LATENT SPACE Assertion Supported
OpenAI currently restricts reinforcement fine-tuning exclusively to its reasoning models
“No, that's reinforcement fine tuning is only for reasoning models.”
Michelle Pokrass Apr 15, 2025 ▶ 39:32 GPT 4.1: The New OpenAI Workhorse
NO PRIORS Opinion
Foody: Reinforcement fine-tuning makes application-layer AI customization viable
“The reason I'm so optimistic about it taking off is that it's, like, profoundly data efficient, right? And it finally makes sense to customize models at the application layer.”
Brendan Foody Apr 10, 2025 ▶ 32:40 No Priors Ep. 110 | With Mercor CEO and Co-Founder Brendan Foody
LATENT SPACE Disclosure
OpenAI plans to connect agent traces to evals and reinforcement fine-tuning
“Like you got to tie the traces to the evals product so that you can generate good evals. Once you have good evals and graders and tasks, You can use that to do reinforcement fine tuning and you know, lots of details to be figured out over here, but that's the …”
Nikunj Handa Mar 11, 2025 ▶ 24:44 The new OpenAI Agents Platform: CUA, Web Search, Responses API, Agents SDK!!
Lambert: Reinforcement fine-tuning will succeed where answer correctness matters over style
“It is just a new paradigm for fine tuning, and I have seen some of this work, and I'm pretty Optimistic that it'll work for kind of kind of really specific capabilities where answers matter rather than features in your style of text mattering.”
Nathan Lambert Jan 2, 2025 ▶ 9:18 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]
LATENT SPACE Assertion Supported
Lambert: Reinforcement fine-tuning requires only dozens of labeled samples
“This reinforcement fine tuning does many passes over the data, which is why they can say you only need dozens of labeled samples to actually learn from it, which is very different than. Previous training regimes”
Nathan Lambert Jan 2, 2025 ▶ 9:57 The State of Reasoning — from Nathan Lambert, Interconnects/AI2 [LS Live @ NeurIPS 2024]

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.