Dec 30, 2025 · 45m · latent-space

[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor

Ashvin Nair · 30m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Former OpenAI reasoning researcher and Cursor engineer Ashvin Nair analyzes the evolution of reinforcement learning, the disconnect between competitive benchmarks and real-world automation, frontier AI lab dynamics, and Cursor's mission to automate the end-to-end software development lifecycle.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.4 Guest teaching 2.7 Guest disagreement 1.4 The hosts pushing back 2.2
05100:0015:0030:0045:000:00–5:55 · The hosts as informed peer 5/10 Transitioning from Robotics Research to Language Models The host brings up industry context including dinner conversations with Lex Fridman and OpenAI restarting robotics, while questioning market valuations between software AI and robotics. Ashvin clarifies recent robotics funding figures and compares the current state of robotics to GPT-1 and GPT-2.5:56–9:15 · The hosts as informed peer 4/10 Early Work at OpenAI and Benchmark Goodharting The host and guest discuss early CodeGen work at OpenAI and how competitive programming benchmarks like IOI Gold are achieved. Ashvin explains community-level Goodharting of benchmarks, while the host pushes back that optimizing test-time compute is not necessarily cheating.9:15–11:42 · The hosts as informed peer 4/10 Academic RL Pitfalls and the Reality of Scaling Ashvin reflects on his PhD research under Sergey Levin, explaining why academic RL overfit benchmarks with hyperparameter tuning instead of simple scalable methods. The host readily affirms the analysis, pointing out the historical RL winter and startup pivots.11:42–16:47 · The hosts as informed peer 6/10 RL Bottlenecks, Context Integration, and Useful Automation Ashvin argues that RL does not generalize beyond training distributions and requires bringing real-world context into products. The host demonstrates domain knowledge by detailing the GDPval benchmark, source document handling, and distinguishing hyperparameter scaling from neural architecture search.16:47–21:23 · The hosts as informed peer 5/10 Monolithic Models Versus Specialized Organizational Architectures The host challenges the single-model AGI narrative by citing blog posts from OpenAI leadership about abandoning one-model-fits-all architectures. Ashvin reframes this by explaining that OpenAI often ships its org chart rather than fundamental scientific truths.21:23–27:35 · The hosts as informed peer 4/10 The Evolution of OpenAI's Reasoning Paradigm and o-Series Ashvin explains the internal conviction led by Ilya Sutskever and Jakub Pachocki that drove OpenAI toward RL reasoning models. The host inquires into internal prototypes and the diminishing gap between internal research leads and external releases.27:35–30:35 · The hosts as informed peer 4/10 AI Capability Forecasting and Calibration Misconceptions Ashvin describes forecasting discrepancies at AI conferences where short-term capabilities are underestimated while long-term timelines are overly sci-fi. The host adds historical context about human bias toward predicting transformative events within one's own lifetime.30:35–36:19 · The hosts as informed peer 5/10 The DeepSeek Moment and Frontier RL Convergence The host aggressively challenges Ashvin's decision to leave OpenAI's massive resources for Cursor and dismisses Tab's two-hour policy updates as mere autocomplete. Ashvin firmly pushes back, arguing that small co-located product and ML teams enable workflows that bureaucratic frontier labs cannot match.36:20–40:23 · The hosts as informed peer 4/10 Engineering Cursor Composer and End-to-End Dev Automation Ashvin details the engineering focus behind Cursor Composer, emphasizing low-latency synchronous iteration over slow frontier model calls. The host references specific internal cluster visualizations and developer workflows.40:23–43:51 · The hosts as informed peer 5/10 Theoretical Limits, Continual Learning, and Neural Memory The host questions the feasibility of continual learning into model weights, arguing finite parameter capacity and information theoretic limits lead to catastrophic forgetting. Ashvin refutes the capacity bottleneck by calculating the small proportion of deployment tokens relative to trillions of pretraining tokens.43:52–44:59 · The hosts as informed peer 2/10 Interviewing for RL Roles and Cursor Hiring Pitch The host asks for an effective RL interview question, prompting Ashvin to discuss Cursor's work trials and probe candidate knowledge on the instability of off-policy reinforcement learning.0:00–5:55 · Guest teaching 2/10 Transitioning from Robotics Research to Language Models The host brings up industry context including dinner conversations with Lex Fridman and OpenAI restarting robotics, while questioning market valuations between software AI and robotics. Ashvin clarifies recent robotics funding figures and compares the current state of robotics to GPT-1 and GPT-2.5:56–9:15 · Guest teaching 3/10 Early Work at OpenAI and Benchmark Goodharting The host and guest discuss early CodeGen work at OpenAI and how competitive programming benchmarks like IOI Gold are achieved. Ashvin explains community-level Goodharting of benchmarks, while the host pushes back that optimizing test-time compute is not necessarily cheating.9:15–11:42 · Guest teaching 2/10 Academic RL Pitfalls and the Reality of Scaling Ashvin reflects on his PhD research under Sergey Levin, explaining why academic RL overfit benchmarks with hyperparameter tuning instead of simple scalable methods. The host readily affirms the analysis, pointing out the historical RL winter and startup pivots.11:42–16:47 · Guest teaching 3/10 RL Bottlenecks, Context Integration, and Useful Automation Ashvin argues that RL does not generalize beyond training distributions and requires bringing real-world context into products. The host demonstrates domain knowledge by detailing the GDPval benchmark, source document handling, and distinguishing hyperparameter scaling from neural architecture search.16:47–21:23 · Guest teaching 3/10 Monolithic Models Versus Specialized Organizational Architectures The host challenges the single-model AGI narrative by citing blog posts from OpenAI leadership about abandoning one-model-fits-all architectures. Ashvin reframes this by explaining that OpenAI often ships its org chart rather than fundamental scientific truths.21:23–27:35 · Guest teaching 3/10 The Evolution of OpenAI's Reasoning Paradigm and o-Series Ashvin explains the internal conviction led by Ilya Sutskever and Jakub Pachocki that drove OpenAI toward RL reasoning models. The host inquires into internal prototypes and the diminishing gap between internal research leads and external releases.27:35–30:35 · Guest teaching 2/10 AI Capability Forecasting and Calibration Misconceptions Ashvin describes forecasting discrepancies at AI conferences where short-term capabilities are underestimated while long-term timelines are overly sci-fi. The host adds historical context about human bias toward predicting transformative events within one's own lifetime.30:35–36:19 · Guest teaching 4/10 The DeepSeek Moment and Frontier RL Convergence The host aggressively challenges Ashvin's decision to leave OpenAI's massive resources for Cursor and dismisses Tab's two-hour policy updates as mere autocomplete. Ashvin firmly pushes back, arguing that small co-located product and ML teams enable workflows that bureaucratic frontier labs cannot match.36:20–40:23 · Guest teaching 2/10 Engineering Cursor Composer and End-to-End Dev Automation Ashvin details the engineering focus behind Cursor Composer, emphasizing low-latency synchronous iteration over slow frontier model calls. The host references specific internal cluster visualizations and developer workflows.40:23–43:51 · Guest teaching 4/10 Theoretical Limits, Continual Learning, and Neural Memory The host questions the feasibility of continual learning into model weights, arguing finite parameter capacity and information theoretic limits lead to catastrophic forgetting. Ashvin refutes the capacity bottleneck by calculating the small proportion of deployment tokens relative to trillions of pretraining tokens.43:52–44:59 · Guest teaching 2/10 Interviewing for RL Roles and Cursor Hiring Pitch The host asks for an effective RL interview question, prompting Ashvin to discuss Cursor's work trials and probe candidate knowledge on the instability of off-policy reinforcement learning.0:00–5:55 · Guest disagreement 2/10 Transitioning from Robotics Research to Language Models The host brings up industry context including dinner conversations with Lex Fridman and OpenAI restarting robotics, while questioning market valuations between software AI and robotics. Ashvin clarifies recent robotics funding figures and compares the current state of robotics to GPT-1 and GPT-2.5:56–9:15 · Guest disagreement 2/10 Early Work at OpenAI and Benchmark Goodharting The host and guest discuss early CodeGen work at OpenAI and how competitive programming benchmarks like IOI Gold are achieved. Ashvin explains community-level Goodharting of benchmarks, while the host pushes back that optimizing test-time compute is not necessarily cheating.9:15–11:42 · Guest disagreement 1/10 Academic RL Pitfalls and the Reality of Scaling Ashvin reflects on his PhD research under Sergey Levin, explaining why academic RL overfit benchmarks with hyperparameter tuning instead of simple scalable methods. The host readily affirms the analysis, pointing out the historical RL winter and startup pivots.11:42–16:47 · Guest disagreement 1/10 RL Bottlenecks, Context Integration, and Useful Automation Ashvin argues that RL does not generalize beyond training distributions and requires bringing real-world context into products. The host demonstrates domain knowledge by detailing the GDPval benchmark, source document handling, and distinguishing hyperparameter scaling from neural architecture search.16:47–21:23 · Guest disagreement 2/10 Monolithic Models Versus Specialized Organizational Architectures The host challenges the single-model AGI narrative by citing blog posts from OpenAI leadership about abandoning one-model-fits-all architectures. Ashvin reframes this by explaining that OpenAI often ships its org chart rather than fundamental scientific truths.21:23–27:35 · Guest disagreement 1/10 The Evolution of OpenAI's Reasoning Paradigm and o-Series Ashvin explains the internal conviction led by Ilya Sutskever and Jakub Pachocki that drove OpenAI toward RL reasoning models. The host inquires into internal prototypes and the diminishing gap between internal research leads and external releases.27:35–30:35 · Guest disagreement 1/10 AI Capability Forecasting and Calibration Misconceptions Ashvin describes forecasting discrepancies at AI conferences where short-term capabilities are underestimated while long-term timelines are overly sci-fi. The host adds historical context about human bias toward predicting transformative events within one's own lifetime.30:35–36:19 · Guest disagreement 3/10 The DeepSeek Moment and Frontier RL Convergence The host aggressively challenges Ashvin's decision to leave OpenAI's massive resources for Cursor and dismisses Tab's two-hour policy updates as mere autocomplete. Ashvin firmly pushes back, arguing that small co-located product and ML teams enable workflows that bureaucratic frontier labs cannot match.36:20–40:23 · Guest disagreement 1/10 Engineering Cursor Composer and End-to-End Dev Automation Ashvin details the engineering focus behind Cursor Composer, emphasizing low-latency synchronous iteration over slow frontier model calls. The host references specific internal cluster visualizations and developer workflows.40:23–43:51 · Guest disagreement 2/10 Theoretical Limits, Continual Learning, and Neural Memory The host questions the feasibility of continual learning into model weights, arguing finite parameter capacity and information theoretic limits lead to catastrophic forgetting. Ashvin refutes the capacity bottleneck by calculating the small proportion of deployment tokens relative to trillions of pretraining tokens.43:52–44:59 · Guest disagreement 0/10 Interviewing for RL Roles and Cursor Hiring Pitch The host asks for an effective RL interview question, prompting Ashvin to discuss Cursor's work trials and probe candidate knowledge on the instability of off-policy reinforcement learning.0:00–5:55 · The hosts pushing back 3/10 Transitioning from Robotics Research to Language Models The host brings up industry context including dinner conversations with Lex Fridman and OpenAI restarting robotics, while questioning market valuations between software AI and robotics. Ashvin clarifies recent robotics funding figures and compares the current state of robotics to GPT-1 and GPT-2.5:56–9:15 · The hosts pushing back 2/10 Early Work at OpenAI and Benchmark Goodharting The host and guest discuss early CodeGen work at OpenAI and how competitive programming benchmarks like IOI Gold are achieved. Ashvin explains community-level Goodharting of benchmarks, while the host pushes back that optimizing test-time compute is not necessarily cheating.9:15–11:42 · The hosts pushing back 1/10 Academic RL Pitfalls and the Reality of Scaling Ashvin reflects on his PhD research under Sergey Levin, explaining why academic RL overfit benchmarks with hyperparameter tuning instead of simple scalable methods. The host readily affirms the analysis, pointing out the historical RL winter and startup pivots.11:42–16:47 · The hosts pushing back 2/10 RL Bottlenecks, Context Integration, and Useful Automation Ashvin argues that RL does not generalize beyond training distributions and requires bringing real-world context into products. The host demonstrates domain knowledge by detailing the GDPval benchmark, source document handling, and distinguishing hyperparameter scaling from neural architecture search.16:47–21:23 · The hosts pushing back 3/10 Monolithic Models Versus Specialized Organizational Architectures The host challenges the single-model AGI narrative by citing blog posts from OpenAI leadership about abandoning one-model-fits-all architectures. Ashvin reframes this by explaining that OpenAI often ships its org chart rather than fundamental scientific truths.21:23–27:35 · The hosts pushing back 1/10 The Evolution of OpenAI's Reasoning Paradigm and o-Series Ashvin explains the internal conviction led by Ilya Sutskever and Jakub Pachocki that drove OpenAI toward RL reasoning models. The host inquires into internal prototypes and the diminishing gap between internal research leads and external releases.27:35–30:35 · The hosts pushing back 1/10 AI Capability Forecasting and Calibration Misconceptions Ashvin describes forecasting discrepancies at AI conferences where short-term capabilities are underestimated while long-term timelines are overly sci-fi. The host adds historical context about human bias toward predicting transformative events within one's own lifetime.30:35–36:19 · The hosts pushing back 6/10 The DeepSeek Moment and Frontier RL Convergence The host aggressively challenges Ashvin's decision to leave OpenAI's massive resources for Cursor and dismisses Tab's two-hour policy updates as mere autocomplete. Ashvin firmly pushes back, arguing that small co-located product and ML teams enable workflows that bureaucratic frontier labs cannot match.36:20–40:23 · The hosts pushing back 1/10 Engineering Cursor Composer and End-to-End Dev Automation Ashvin details the engineering focus behind Cursor Composer, emphasizing low-latency synchronous iteration over slow frontier model calls. The host references specific internal cluster visualizations and developer workflows.40:23–43:51 · The hosts pushing back 4/10 Theoretical Limits, Continual Learning, and Neural Memory The host questions the feasibility of continual learning into model weights, arguing finite parameter capacity and information theoretic limits lead to catastrophic forgetting. Ashvin refutes the capacity bottleneck by calculating the small proportion of deployment tokens relative to trillions of pretraining tokens.43:52–44:59 · The hosts pushing back 0/10 Interviewing for RL Roles and Cursor Hiring Pitch The host asks for an effective RL interview question, prompting Ashvin to discuss Cursor's work trials and probe candidate knowledge on the instability of off-policy reinforcement learning.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 0%45:00 · the hosts 0% · guest 0%
Sharpest disagreement ▶ 34:34 Ashvin rejects host's dismissal of Tab automation

Ashvin explicitly counters the host's attempt to minimize Cursor's two-hour policy training as simple autocomplete, arguing instead that organizational agility drives real ML iteration.

Hardest push from the hosts ▶ 33:06 Host presses why Ashvin left OpenAI's compute

The host bluntly refuses the premise of leaving OpenAI, pointing out their infinite data, Codex assets, and massive compute compared to an early-stage startup.

Biggest teaching moment ▶ 41:41 Ashvin calculates token ratios to disprove weight capacity limits

Ashvin educates the host on why catastrophic forgetting is overstated for continual deployment, demonstrating that millions of task tokens are negligible against trillions of pre-trained tokens.

The host holds their own ▶ 13:33 Host outlines GDPval methodology and agent evaluation

The host demonstrates deep technical knowledge by breaking down GDPval's 128-task white-collar suite and explaining the need for raw uncleaned data inputs.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Transitioning from Robotics Research to Language Models 5223 The host brings up industry context including dinner conversations with Lex Fridman and OpenAI restarting robotics, while questioning market valuations between software AI and robotics. Ashvin clarifies recent robotics funding figures and compares the current state of robotics to GPT-1 and GPT-2.
Early Work at OpenAI and Benchmark Goodharting 4322 The host and guest discuss early CodeGen work at OpenAI and how competitive programming benchmarks like IOI Gold are achieved. Ashvin explains community-level Goodharting of benchmarks, while the host pushes back that optimizing test-time compute is not necessarily cheating.
Academic RL Pitfalls and the Reality of Scaling 4211 Ashvin reflects on his PhD research under Sergey Levin, explaining why academic RL overfit benchmarks with hyperparameter tuning instead of simple scalable methods. The host readily affirms the analysis, pointing out the historical RL winter and startup pivots.
RL Bottlenecks, Context Integration, and Useful Automation 6312 Ashvin argues that RL does not generalize beyond training distributions and requires bringing real-world context into products. The host demonstrates domain knowledge by detailing the GDPval benchmark, source document handling, and distinguishing hyperparameter scaling from neural architecture search.
Monolithic Models Versus Specialized Organizational Architectures 5323 The host challenges the single-model AGI narrative by citing blog posts from OpenAI leadership about abandoning one-model-fits-all architectures. Ashvin reframes this by explaining that OpenAI often ships its org chart rather than fundamental scientific truths.
The Evolution of OpenAI's Reasoning Paradigm and o-Series 4311 Ashvin explains the internal conviction led by Ilya Sutskever and Jakub Pachocki that drove OpenAI toward RL reasoning models. The host inquires into internal prototypes and the diminishing gap between internal research leads and external releases.
AI Capability Forecasting and Calibration Misconceptions 4211 Ashvin describes forecasting discrepancies at AI conferences where short-term capabilities are underestimated while long-term timelines are overly sci-fi. The host adds historical context about human bias toward predicting transformative events within one's own lifetime.
The DeepSeek Moment and Frontier RL Convergence 5436 The host aggressively challenges Ashvin's decision to leave OpenAI's massive resources for Cursor and dismisses Tab's two-hour policy updates as mere autocomplete. Ashvin firmly pushes back, arguing that small co-located product and ML teams enable workflows that bureaucratic frontier labs cannot match.
Engineering Cursor Composer and End-to-End Dev Automation 4211 Ashvin details the engineering focus behind Cursor Composer, emphasizing low-latency synchronous iteration over slow frontier model calls. The host references specific internal cluster visualizations and developer workflows.
Theoretical Limits, Continual Learning, and Neural Memory 5424 The host questions the feasibility of continual learning into model weights, arguing finite parameter capacity and information theoretic limits lead to catastrophic forgetting. Ashvin refutes the capacity bottleneck by calculating the small proportion of deployment tokens relative to trillions of pretraining tokens.
Interviewing for RL Roles and Cursor Hiring Pitch 2200 The host asks for an effective RL interview question, prompting Ashvin to discuss Cursor's work trials and probe candidate knowledge on the instability of off-policy reinforcement learning.

Statements from this episode (25)

Prediction Didn’t hold up
Nair: LLM agents will hit $1T before robotics hits $10B
“It feels like LLM agents are going to be like a trillion dollar market before robotics is maybe even like a ten billion dollar market.”
Ashvin Nair Dec 30, 2025 ▶ 3:59
Opinion
Nair: AI robotics is currently in its 'GPT-1 to GPT-2' era
“Yeah, like I would say that robotics is in kind of like the GPT-one to GPT-two area right now.”
Ashvin Nair Dec 30, 2025 ▶ 4:54
Opinion
Nair: Robotics investments today back teams, not proven technology
“Yeah, I think at this point, especially, it still feels like in robotics, you're not exactly investing in a technology, probably, you're just investing in a team.”
Ashvin Nair Dec 30, 2025 ▶ 5:34
What-if
Nair: Previously Believed IOI Gold Would Mean AI Was Solved
“If you told me that we could have gotten IOI Gold then, I would have just assumed that we could all just go on vacation, like, you know, it's all over, like, AI is solved, like, no point in working anymore.”
Ashvin Nair Dec 30, 2025 ▶ 7:09
Insight
Nair: IOI Gold Mastery Has Not Solved Real-World Coding Automation
“Most programmers in the world cannot do IOI at any decent level. But, like, we're still struggling to, like, automate most programming jobs or, like, you know, there's a lot left to do.”
Ashvin Nair Dec 30, 2025 ▶ 8:48
Opinion
Nair: 2017–2022 academic RL breakthroughs failed because researchers overfit to benchmarks
“A lot of the methods that people were really excited about is, like you know, off policy learning, like, value functions, like, these kind of things, and somehow that, that stuff hasn't really panned out, I would say, and it's not exactly clear why, but in the…”
Ashvin Nair Dec 30, 2025 ▶ 9:26
Insight
Nair: Academia rewards complex math over simple, generalizable solutions
“One of the pitfalls of academia is that it doesn't really reward, like, simple ideas that work, and instead kind of tends to reward, like, kind of mathier ideas. Those mathier ideas also give you these, like, kind of implicit knobs to tune that allow you to, l…”
Ashvin Nair Dec 30, 2025 ▶ 10:52
Insight
Nair: RL on LLMs is peaky and fails to generalize beyond training
“RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. It generalizes to some extent, and generalizes in interesting ways, but It's like very peaky, right? Like …”
Ashvin Nair Dec 30, 2025 ▶ 12:26
Insight
Nair: Context integration, not model intelligence, bottlenecks useful automation
“A big thing that needs to happen is, like, it's not, it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, …”
Ashvin Nair Dec 30, 2025 ▶ 13:09
Opinion
Nair: OpenAI model splits happen because it ships its org chart
“OpenAI has a tendency to ship the org chart, basically.”
Ashvin Nair Dec 30, 2025 ▶ 17:45
Opinion
Nair: Public corporate boards may govern AI more democratically than non-profit boards
“When the blip happened, one of my reactions was like, well, you know, this nonprofit board stuff, like, actually, if it takes such somewhat, like, surprising, Maybe erratic actions, like maybe you'd rather just have, like, you know, a thing like the Microsoft …”
Ashvin Nair Dec 30, 2025 ▶ 20:27
Opinion
Nair: Sutskever and Pachocki Drove OpenAI's First-Principles Research Conviction
“I think in general, OpenAI is really good about, like, having conviction in something, and just, like, really, like, from first principles, like, going after it, and I think, like, the people who are kind of most responsible for that is probably, like, Ilya Se…”
Ashvin Nair Dec 30, 2025 ▶ 22:28
Insight
Nair: RLHF Is a Side Branch Because Compute Cannot Be Scaled
“I think human feedback is kind of like a bit of like a side branch, because you can't really pour that much compute Into it, right? It's like, you take the model, and you, like, elicit it to be a little bit better in terms of personality”
Ashvin Nair Dec 30, 2025 ▶ 23:16
Insight
Nair: OpenAI Progress Feels Smooth Internally, Not Like Sudden Leaps
“It seems like externally people are kind of very, like, oh like, research seems to come in these, like, big leaps. But I think internally at OpenAI, it feels very smooth.”
Ashvin Nair Dec 30, 2025 ▶ 25:48
Assertion Not checkable as stated
Nair: OpenAI internal models surpassed forecasters' 2027 benchmark targets before o1 launch
“And their estimates were, like, oh, we'll be at, like, 10, 20% in, like, 20, 27, and I think at the time, there was, like, you know, models internally that were, like, already better than their estimates, so that, like, it's, like, off by, like, you know, two …”
Ashvin Nair Dec 30, 2025 ▶ 28:28
Prediction Not checkable as stated
Nair: AI will probably reach human-level intelligence around 2030
“And actually, you know, it is, it's somewhere in the, like, twenty-thirty-ish thing that, like, it will probably reach, like, human level intelligence.”
Ashvin Nair Dec 30, 2025 ▶ 29:52
Assertion Not checkable as stated
Nair: OpenAI already possessed a superior model during the DeepSeek release
“The feeling in OpenAI is that like, well, I think we had a better model already at the time, right?”
Ashvin Nair Dec 30, 2025 ▶ 31:21
Opinion
Nair: Frontier AI labs have converged on similar reinforcement learning methods
“Well, it does seem like basically a lot of the labs have kind of like converged onto some similar-ish way of doing RL, and they're all kind of back at the same level of like Frontier again”
Ashvin Nair Dec 30, 2025 ▶ 31:54
Assertion Supported
Nair: Cursor updates its tab autocomplete model policy every two hours
“Recently Jacob Jackson had this blog post about like, online tab, where like, you know, we're doing policy- Because there is every two hours. Exactly, like, a policy update every two hours or something”
Ashvin Nair Dec 30, 2025 ▶ 34:06
Insight
Nair: AI models lag orders of magnitude behind human one-shot error learning
“It seems like we're kind of, like, a few orders of magnitude of, like, kind of data efficiency, basically, away from, like, that kind of, like, you know, you do something once, or, like, you make a mistake like, you, yeah, you introduce, like, a bug in your co…”
Ashvin Nair Dec 30, 2025 ▶ 35:45
Prediction Not checkable as stated
Nair: Continual learning breakthroughs will be paradigm-shifting within a year
“I suspect that it will be kind of, like, paradigm shifting in the next, like, year or something, but I have no idea, like, you know, what it might be.”
Ashvin Nair Dec 30, 2025 ▶ 36:11
Disclosure
Nair: Cursor aims to automate end-to-end software engineering process
“What we're really aiming for is, like, more, like, you know, automate software engineering as a process where you, like, write code, you go look at Datadog look at what's, like, happening, then come back and, like, you know, maybe have some hypotheses about wh…”
Ashvin Nair Dec 30, 2025 ▶ 38:34
Disclosure
Nair: Internal Slack posts replaced reading external papers at OpenAI
“Unfortunately I've like, kind of gotten the habit, especially at OpenAI, of like, not reading that much external work, and just like reading people's like, Slack posts internally. That's like the main, like, way to like, you know like, learn new stuff.”
Ashvin Nair Dec 30, 2025 ▶ 40:39
Opinion
Nair: Continual learning during deployment does not risk model capacity overload
“If you could learn enough about those million tokens that you're actually in deployment on I don't think you should need, like, I don't think there's a risk of overloading the capacity of your model, right? Because you can train on a trillion tokens, and it's …”
Ashvin Nair Dec 30, 2025 ▶ 42:00
Opinion
Nair: Fundamental science is less fruitful than empirical work for AI progress
“There's like actually so many of these kind of more scientific questions that I would like love to explore sometime, but then it really kind of conflicts with like empirical stuff, you know, like unfortunately at any given moment in time, it doesn't seem like …”
Ashvin Nair Dec 30, 2025 ▶ 43:16
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.