Dec 30, 2025 · 45m · latent-space
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
Former OpenAI reasoning researcher and Cursor engineer Ashvin Nair analyzes the evolution of reinforcement learning, the disconnect between competitive benchmarks and real-world automation, frontier AI lab dynamics, and Cursor's mission to automate the end-to-end software development lifecycle.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Ashvin explicitly counters the host's attempt to minimize Cursor's two-hour policy training as simple autocomplete, arguing instead that organizational agility drives real ML iteration.
Hardest push from the hosts ▶ 33:06 Host presses why Ashvin left OpenAI's computeThe host bluntly refuses the premise of leaving OpenAI, pointing out their infinite data, Codex assets, and massive compute compared to an early-stage startup.
Biggest teaching moment ▶ 41:41 Ashvin calculates token ratios to disprove weight capacity limitsAshvin educates the host on why catastrophic forgetting is overstated for continual deployment, demonstrating that millions of task tokens are negligible against trillions of pre-trained tokens.
The host holds their own ▶ 13:33 Host outlines GDPval methodology and agent evaluationThe host demonstrates deep technical knowledge by breaking down GDPval's 128-task white-collar suite and explaining the need for raw uncleaned data inputs.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Transitioning from Robotics Research to Language Models | 5 | 2 | 2 | 3 | The host brings up industry context including dinner conversations with Lex Fridman and OpenAI restarting robotics, while questioning market valuations between software AI and robotics. Ashvin clarifies recent robotics funding figures and compares the current state of robotics to GPT-1 and GPT-2. | |
| Early Work at OpenAI and Benchmark Goodharting | 4 | 3 | 2 | 2 | The host and guest discuss early CodeGen work at OpenAI and how competitive programming benchmarks like IOI Gold are achieved. Ashvin explains community-level Goodharting of benchmarks, while the host pushes back that optimizing test-time compute is not necessarily cheating. | |
| Academic RL Pitfalls and the Reality of Scaling | 4 | 2 | 1 | 1 | Ashvin reflects on his PhD research under Sergey Levin, explaining why academic RL overfit benchmarks with hyperparameter tuning instead of simple scalable methods. The host readily affirms the analysis, pointing out the historical RL winter and startup pivots. | |
| RL Bottlenecks, Context Integration, and Useful Automation | 6 | 3 | 1 | 2 | Ashvin argues that RL does not generalize beyond training distributions and requires bringing real-world context into products. The host demonstrates domain knowledge by detailing the GDPval benchmark, source document handling, and distinguishing hyperparameter scaling from neural architecture search. | |
| Monolithic Models Versus Specialized Organizational Architectures | 5 | 3 | 2 | 3 | The host challenges the single-model AGI narrative by citing blog posts from OpenAI leadership about abandoning one-model-fits-all architectures. Ashvin reframes this by explaining that OpenAI often ships its org chart rather than fundamental scientific truths. | |
| The Evolution of OpenAI's Reasoning Paradigm and o-Series | 4 | 3 | 1 | 1 | Ashvin explains the internal conviction led by Ilya Sutskever and Jakub Pachocki that drove OpenAI toward RL reasoning models. The host inquires into internal prototypes and the diminishing gap between internal research leads and external releases. | |
| AI Capability Forecasting and Calibration Misconceptions | 4 | 2 | 1 | 1 | Ashvin describes forecasting discrepancies at AI conferences where short-term capabilities are underestimated while long-term timelines are overly sci-fi. The host adds historical context about human bias toward predicting transformative events within one's own lifetime. | |
| The DeepSeek Moment and Frontier RL Convergence | 5 | 4 | 3 | 6 | The host aggressively challenges Ashvin's decision to leave OpenAI's massive resources for Cursor and dismisses Tab's two-hour policy updates as mere autocomplete. Ashvin firmly pushes back, arguing that small co-located product and ML teams enable workflows that bureaucratic frontier labs cannot match. | |
| Engineering Cursor Composer and End-to-End Dev Automation | 4 | 2 | 1 | 1 | Ashvin details the engineering focus behind Cursor Composer, emphasizing low-latency synchronous iteration over slow frontier model calls. The host references specific internal cluster visualizations and developer workflows. | |
| Theoretical Limits, Continual Learning, and Neural Memory | 5 | 4 | 2 | 4 | The host questions the feasibility of continual learning into model weights, arguing finite parameter capacity and information theoretic limits lead to catastrophic forgetting. Ashvin refutes the capacity bottleneck by calculating the small proportion of deployment tokens relative to trillions of pretraining tokens. | |
| Interviewing for RL Roles and Cursor Hiring Pitch | 2 | 2 | 0 | 0 | The host asks for an effective RL interview question, prompting Ashvin to discuss Cursor's work trials and probe candidate knowledge on the instability of off-policy reinforcement learning. |