Opinion certainty 3/5 debate potential 3/5

Nair: 2017–2022 academic RL breakthroughs failed because researchers overfit to benchmarks

Ashvin Nair · [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor · Dec 30, 2025 · at 9:26

Ashvin Nair, former OpenAI reasoning researcher and engineer at Cursor, discusses his PhD research at UC Berkeley and the disconnect between academic RL hype and real-world utility.

0:00 / 0:35exact quote · 35.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“A lot of the methods that people were really excited about is, like you know, off policy learning, like, value functions, like, these kind of things, and somehow that, that stuff hasn't really panned out, I would say, and it's not exactly clear why, but in the academic literature, we thought we were making a ton of progress. And I think in retrospect, I had to say that we probably kind of overfit to the benchmarks pretty heavily, and, you know, how I see this in retrospect is that we gave ourselves a lot of, like, new knobs to tune, and then implicitly kind of tuned those to fit the benchmarks.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Ashvin Nair

Prediction Didn’t hold up
Nair: LLM agents will hit $1T before robotics hits $10B
“It feels like LLM agents are going to be like a trillion dollar market before robotics is maybe even like a ten billion dollar market.”
Ashvin Nair Dec 30, 2025 ▶ 3:59 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Nair: Academia rewards complex math over simple, generalizable solutions
“One of the pitfalls of academia is that it doesn't really reward, like, simple ideas that work, and instead kind of tends to reward, like, kind of mathier ideas. Those mathier ideas also give you these, like, kind of implicit knobs to tune that allow you to, l…”
Ashvin Nair Dec 30, 2025 ▶ 10:52 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Nair: RL on LLMs is peaky and fails to generalize beyond training
“RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. It generalizes to some extent, and generalizes in interesting ways, but It's like very peaky, right? Like …”
Ashvin Nair Dec 30, 2025 ▶ 12:26 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Nair: Context integration, not model intelligence, bottlenecks useful automation
“A big thing that needs to happen is, like, it's not, it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, …”
Ashvin Nair Dec 30, 2025 ▶ 13:09 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Opinion
Nair: OpenAI model splits happen because it ships its org chart
“OpenAI has a tendency to ship the org chart, basically.”
Ashvin Nair Dec 30, 2025 ▶ 17:45 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Insight
Nair: RLHF Is a Side Branch Because Compute Cannot Be Scaled
“I think human feedback is kind of like a bit of like a side branch, because you can't really pour that much compute Into it, right? It's like, you take the model, and you, like, elicit it to be a little bit better in terms of personality”
Ashvin Nair Dec 30, 2025 ▶ 23:16 [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.