SFT
topic on 2 shows · 6 statements across 5 episodes
6 statements about SFT, every show
Dubois: Starting post-training with RL without SFT is extremely inefficient
“Because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically.”
Dubois: Effective reinforcement learning pipelines prevent AI hallucinations caused by SFT
“So, so hallucination at least the intuition that people have is that it can come, for example, from SFT, and it can come from this, like, pursuing pipeline, but if you have good reinforcement in pipeline, that shouldn't happen too often.”
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Pokrass: Developers are sleeping on preference fine-tuning for model style steering
“One thing I will say is that I think people have slept on the preference fine tuning offering or the, I think that's what we call the product. So SFT is, people know it pretty well. It's the original fine tuning we had, whereas this preference fine tuning is s…”