Insight certainty 4/5 debate potential 2/5

Fulford: Training reasoning models on math and coding generalizes to writing

Isa Fulford · No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford · Apr 24, 2025 · at 7:23

OpenAI researcher Isa Fulford discusses AI model generalization and reinforcement fine-tuning with Sarah Guo on No Priors.

0:00 / 0:26exact quote · 26.4s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“So I think in general you will always get a model better, better at a specific task if you train on that task, but we also see a lot of generalization from training on one kind of task to, you know, other domains. So you can train a reasoning model on mostly math, coding, other reasoning kind of problems, and it will be good at writing, but if you know, trained on that specific task, it would be better at it.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Isa Fulford

Insight
Fulford: Information Synthesis Is a Prerequisite for Scientific AI Discovery
“Secondly, I think the overall goal for OpenAI is to create an AGI that can make new scientific discoveries, and we kind of felt that a prerequisite to that is to be able to synthesize information. You know, if you can't write a literature review, you're not go…”
Isa Fulford Apr 24, 2025 ▶ 2:58 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Insight
Fulford: RFT is only worthwhile for out-of-distribution or make-or-break tasks
“I think if you have a very specific task that you think is so different to anything that the model was likely trained on and you try it a bunch of times yourself and you've tried a lot of different prompts and it's just really not good at it. So maybe it's gen…”
Isa Fulford Apr 24, 2025 ▶ 7:50 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Not checkable as stated
OpenAI: Deep Research learned upfront planning without explicit instruction
“We didn't teach it To plan up front, but sometimes we'll see it does end up making a plan up front before starting its research.”
Isa Fulford Apr 24, 2025 ▶ 10:30 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Not checkable as stated
Fulford: OpenAI Deep Research attempts reward hacking around tool restrictions
“Sometimes the model will do smart things and try to get around restrictions you put on it. So you have to make sure that it's not hacking, you know, and trying to use a different search engine other than the search engine that you gave it or something like tha…”
Isa Fulford Apr 24, 2025 ▶ 10:42 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Not checkable as stated
Fulford: Deep Research Hallucinates Less Than Any Prior OpenAI Model
“While this model is hallucinates less than any model that we've ever released, it is still possible for it to hallucinate most times because it will infer something incorrectly from one of its sources.”
Isa Fulford Apr 24, 2025 ▶ 11:38 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Assertion Supported
Fulford: Deep Research Completes Multi-Hour Human Work in 5 to 30 Minutes
“Right now, in five or 30 minutes, it can do what human experts rate take many hours.”
Isa Fulford Apr 24, 2025 ▶ 27:01 No Priors Ep. 112 | With OpenAI Deep Research, Isa Fulford
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.