Insight certainty 3/5 debate potential 2/5

Fulford: RFT is only worthwhile for out-of-distribution or make-or-break tasks

Isa Fulford · No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford · Apr 24, 2025 · at 7:50

OpenAI's Isa Fulford advises Sarah Guo on when startups should pursue reinforcement fine-tuning versus waiting for foundation model improvements.

0:00 / 0:50exact quote · 50.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I think if you have a very specific task that you think is so different to anything that the model was likely trained on and you try it a bunch of times yourself and you've tried a lot of different prompts and it's just really not good at it. So maybe it's genetic sequencing task or something that's just so out of distribution for the model that it doesn't know how to figure it out. I think that is a good time to try reinforcement fine tuning or If you have a task that is so critical to your, like, business workflow that getting the extra 10, 15% performance is really make or break, then probably try it. But if it's something that you think, oh, the model's pretty good at, but it gets things wrong, you know, some percentage of the time, and then you see with every next model that's released, it gets a little bit better, it might not be worth the effort if the model naturally is just going to get better at those things.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Isa Fulford

Insight
Fulford: Information Synthesis Is a Prerequisite for Scientific AI Discovery
“Secondly, I think the overall goal for OpenAI is to create an AGI that can make new scientific discoveries, and we kind of felt that a prerequisite to that is to be able to synthesize information. You know, if you can't write a literature review, you're not go…”
Isa Fulford Apr 24, 2025 ▶ 2:58 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Not checkable as stated
OpenAI: Deep Research learned upfront planning without explicit instruction
“We didn't teach it To plan up front, but sometimes we'll see it does end up making a plan up front before starting its research.”
Isa Fulford Apr 24, 2025 ▶ 10:30 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Not checkable as stated
Fulford: OpenAI Deep Research attempts reward hacking around tool restrictions
“Sometimes the model will do smart things and try to get around restrictions you put on it. So you have to make sure that it's not hacking, you know, and trying to use a different search engine other than the search engine that you gave it or something like tha…”
Isa Fulford Apr 24, 2025 ▶ 10:42 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Not checkable as stated
Fulford: Deep Research Hallucinates Less Than Any Prior OpenAI Model
“While this model is hallucinates less than any model that we've ever released, it is still possible for it to hallucinate most times because it will infer something incorrectly from one of its sources.”
Isa Fulford Apr 24, 2025 ▶ 11:38 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Supported
Fulford: Deep Research Completes Multi-Hour Human Work in 5 to 30 Minutes
“Right now, in five or 30 minutes, it can do what human experts rate take many hours.”
Isa Fulford Apr 24, 2025 ▶ 27:01 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Disclosure
Fulford: OpenAI Prioritized Read-Only Synthesis Over Action Agents
“Yeah, so I think before we focused on taking right actions, which those are examples of taking right actions, we wanted to get really good at synthesizing information from a large number of sources and mostly read-only tasks.”
Isa Fulford Apr 24, 2025 ▶ 2:38 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.