SFT

topic on 2 shows · 6 statements across 5 episodes

Latent Space the MAD Podcast

6 statements about SFT, every show

MAD Insight
Dubois: Starting post-training with RL without SFT is extremely inefficient
“Because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically.”
Yann Dubois May 21, 2026 ▶ 36:48 OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
MAD Insight
Dubois: Effective reinforcement learning pipelines prevent AI hallucinations caused by SFT
“So, so hallucination at least the intuition that people have is that it can come, for example, from SFT, and it can come from this, like, pursuing pipeline, but if you have good reinforcement in pipeline, that shouldn't happen too often.”
Yann Dubois May 21, 2026 ▶ 55:42 OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real
Zhang: Superhuman computer vision requires RLHF rather than human SFT data
“But if you only do SFT and the SFT data is annotated by human, then your performance is funded by human. You cannot get, kind of, superhuman performance just by, kind of, this kind of data engine approach to use human annotated data and then learn from that. Y…”
Pengchuan Zhang Dec 18, 2025 ▶ 44:33 SAM 3: The Eyes for AI — Nikhila & Pengchuan (Meta Superintelligence), ft. Joseph Nelson (Roboflow)
Lambert: SFT cannot teach emergent tool use; models must learn via RL environments
“It's very easy to get the model to do tools if you prompt it to, but it's very hard to get the like RL model to learn that the tool is useful. And that's why it's to go through these things where it's like 80 failed tool uses and it still gets it or like it st…”
Nathan Lambert Jul 31, 2025 ▶ 24:35 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)
Fortuna: RFT on 100 examples costs thousands vs. $100 for SFT
“With like SFT, let's say you're using the OpenAI, you know, to do some supervised fine turning, you'll probably have like, you know, maybe a few thousand examples. The job takes a few hours. It costs you like a hundred bucks, right? With RFT, maybe you have li…”
Brendan Fortuna Jul 29, 2025 ▶ 14:43 ⚡️Using RFT to Build Clinical Superintelligence
Pokrass: Developers are sleeping on preference fine-tuning for model style steering
“One thing I will say is that I think people have slept on the preference fine tuning offering or the, I think that's what we call the product. So SFT is, people know it pretty well. It's the original fine tuning we had, whereas this preference fine tuning is s…”
Michelle Pokrass Apr 15, 2025 ▶ 39:04 GPT 4.1: The New OpenAI Workhorse

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.