OpenAI studied reinforcement learning reasoning models internally long before releasing o1
“So before we released a one and thinking reasoning models, we were studying this internally”
Tworek: OpenAI's o1 was mostly a tech demo for solving puzzles
“O-one like, to be perfectly honest, it was really mostly good at solving puzzles and like maybe a few kind of thinking problems here and there, but it wasn't like, it wasn't a very useful model. It was almost more like a technology demonstration.”
Tworek: OpenAI's o1 release caught US AI labs unprepared for RL
“As far as I know, like our O-one release mostly caught a lot of us labs by surprise. They didn't have like similarly advanced RL research program to my knowledge, basically no one.”
Hays: An o3-level reasoning model was open-sourced within a year of o1
“So less than one year between O one announced, which was September of 20, 24. And we have an O three level model open sourced that's runnable on consumer hardware, wild progress.”
Marcus: Models like o1 are not systematically better than GPT-4
“Models like O-I are not systematically better than GPT-IV. There's, they're better in certain use cases. Especially ones where you can create data in advance.”
Srinivas: Most ChatGPT users do not know o1 or GPT-4 differences
“In fact, most people using tracks between the world don't even know there's a model called O one or O three and don't even know what the difference is for GPT four.”
Patel: GPT-5 will simultaneously scale pre-training and post-training reasoning
“And so now GPT-Five, as Sam calls it, is, is gonna be a model that has huge pre-training scale, right? Like GPT-Five, but also huge post-training scale, Like O-one and O-three and continuing to scale that up, right? This would be the first time we see a model …”
Taylor: OpenAI o1 used reinforcement learning on chains of thought
“What at OpenAI, what we did with the O-one model, which is to do some reinforcement learning those chains of thought to really reach new levels of intelligence.”
Hu: OpenAI reasoning models are catching up to Claude 3.5 Sonnet
“The thing about CodeGen, the big game in town that we saw, ah, six months ago was Clawed Sonnet. It's still actually a big contender. Most are still using it. But O-one, O-one Pro, and O-three meaning all these resilient models are starting to see it's almost …”
OpenAI used o1 synthetic data to train Canvas commenting behaviors
“The way we used it is, like, we would use a one model to produce, to, like, simulate, like, use a conversation. Let's say, like, write me a document about XYZ, but then we used a one to, like, produce the document, and then we kind of injected, like, user prom…”
Gerstner: DeepSeek's training run was up to 50% more efficient than o1
“Now, of course, they talked about the six million dollar final training run, and this, it's important to understand, this is also correct. Right? And as I said on CNBC, this compares to about 10 or fifteen million for O-one out of OpenAI. So apples to apples, …”
Karina Nguyen: OpenAI o1 excels when given explicit hard constraints
“If you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and match like the candidate that is most like, fulfill the criteria that you…”
Coogan: OpenAI's $200 o1 Pro uses the same model as $20 o1
“In fact, the O-one model is 20 dollars a month, is basically the same model used in the O-one Pro model for 10 X the price at 200 dollars a month, which raised plenty of eyebrows. The main difference is that O-one Pro thinks for a lot longer before responding”
Coogan: DeepSeek-R1 API is 27 times cheaper than OpenAI's o1
“The deep seek R one API is currently 27%, 27 times cheaper than open AI's O one for a similar level of quality.”
DeepSeek used model distillation to match OpenAI's o1 at lower cost
“And, you know, they've just used the process of distillation to, you know, effectively bring those bigger versions of the sort of state-of-the-art models and distill them down into you know, smaller models, which eventually led to this R-one, you know, the equ…”
Beauchamp: AI intelligence is generative LLMs combined with tree search
“I think if you want to talk about what would intelligence look like, it looks much more like tree search. Combining the generative nature of these LLMs with a really good tree search. And that's what opening I've done with O-one and O-three.”
OpenAI likely reaches frontier capabilities via search, then distills into mini models
“The only way you reach the frontier with the full size models of O-one and O-three is with that stuff. And then you can distill to the minis, the O-one mini, O-three mini. So in my writeup, I said like, maybe this is the formula for O-one mini, O-three mini. T…”
McAteer: OpenAI o1 is the first model that grows more impressive over time
“O-one is actually the first model where I'm getting more impressed by it the more I use it. So, like, when ChatGPT first came out, right, I think it was GPT-III. And at first it seemed like, oh wow, this is amazing, it can actually create text that sounds like…”
McAteer: o1 is the first model to achieve one-shot codebase implementation
“Using O-one was, it was the first time where I would connect it to my IDE. I would provide the full context of my code base. I'll just create a file that concatenates all my files into one just simple text file, give it to O-one, and then say, hey, based on th…”
Hylak: OpenAI o1 struggles to match personal tone and writing styles
“I think that I've had a very hard time getting it to actually write stuff. I know that I've heard of people using it for writing where it's like processing diffs, more like providing critiques or feedback, but at least for myself, I haven't found a good way to…”
McAteer: Asking o1 for full code files instead of diffs is ineffective
“At first I was asking it to generate full files of code, like as I'm making changes. That can get a little bit confusing, and then if you're trying to use like a coding agent to bring it over into your projects. So I feel like that's not the best format for co…”
Hylak: Hidden reasoning tokens create an information asymmetry between OpenAI and developers
“I think that what makes a one even trickier than other models is that there is actually an asymmetric miss to how well open AI understands the model and how well we, for example, the fact that like reasoning tokens are hidden, right? So there's all this stuff …”
Hylak: OpenAI o1 is OpenAI's most capable yet hardest model to use
“We're finding that like, oh, one is the most capable model. I think that opening eye has made. And it's also, I think the hardest to use.”
Hylak: Users willing to wait five minutes for AI will wait an hour
“I think that it's like the number of tasks that you're willing to wait, you know, like 3:05 minutes for it. It's like probably pretty similar to the number of tasks you're willing to wait like an hour for, which is interesting.”
Lambert: OpenAI o1 uses large-scale RL on verifiable outcomes, not MCTS
“You should take open AI at their face value, which they are doing very large scale RL on the verifiable outcomes is what I've added, especially in context of the RL API that they've released, which I'll talk about more, but most of the reasons to believe in mo…”
Lambert: Early DeepSeek and Qwen reasoning models are substantially narrower than o1
“And I think that these models are really substantially narrower than these full O-one models from OpenAI. So OpenAI is, if you use O-one, you can do it for a lot more tasks. If you use, like I was using the DeepSeq model, and it's supposed to be for math or co…”
Wang: China's DeepSeek produced the first replication of OpenAI's o1 model
“OpenAI released O-one and released the O-one preview a number of months ago... Yeah, this is OpenAI's advanced reasoning model, which is great at sort of scientific reasoning and mathematical reasoning and reasoning and code, et cetera. And the very first repl…”
Coogan: OpenAI o1 is built on the GPT-4 foundation model
“The foundation model, yeah, it's still running GPT-IV, but obviously there's a lot of there's a lot of training to have that happens on top and a lot of UI stuff.”
Fireworks AI will release a reasoning model inspired by OpenAI's o1
“So another announcement is we will also announce a, our next
Declarative system is going to be appear as a model that has extremely high quality, and this model is inspired by O-one announcement from OpenAI.
You should see that by the time we announce this o…”
Hu: OpenAI o1 Combines Next-Token LLMs with RL Reward Functions
“GPT is all generative based on predicting the next token and patterns and then getting those results to check that they're correct. So I think a lot of it is you had to have a lot of data that was factually correct and Fed into probably the model and the train…”
Friedman: Full OpenAI o1 Model Is a Huge Step Function Above o1-Preview
“Like the full O-one model, which is coming out any day now is a huge step function above even O-one preview, which is what enabled all these incredible results at the hackathon.”
Weil: AI reasoning models are currently only in their GPT-1 phase
“So, it's basically a new way to scale intelligence, and we feel like we're just at the very beginning, you know, we're at the, like, GPT-I phase of this new form of reasoning.”
Altman: Competing AI labs will successfully replicate OpenAI's o1 model
“After, after a research lab does something, even if you don't know exactly how they did it, it's, I won't say easy, but it's doable to go off and copy it, and you can see this in the replications of GPT-IV, and I'm sure you'll see this in replications of O-one…”
Tan: OpenAI enabled internal model distillation as a developer lock-in strategy
“OpenAI itself has now enabled distillation internal to its own API. So you can use O-one, you can use even GPT-IV or IV-O to distill it down into a much cheaper model that's internal to them, like GPT-IV, IV-O-Mini. And that's sort of their, you know, lock-in …”
Hu: OpenAI's o1 massively increases GPU compute requirements for inference
“I think the other thing that's interesting about one is that it makes a lot of the GPU needs even bigger because it's moving a lot of the computation needs a lot higher for inference because it's taking a lot more time to do a lot of the inference. So I think …”
Sarah Guo says test-time compute scaling unlocks new AI competition
“Another school of thought is which I do subscribe to, by the way, is you know, new scaling law, right? So will allow us to do an important range of new tasks, and how good it is exactly at this moment is not the important thing. It's a new dimension of competi…”
Competitors will catch up to OpenAI's o1 reasoning model within six months
“O1 just came out. We haven't fully benchmarked on our side yet. It's definitely an improvement. And so open AI is ahead of everybody else. It's an improvement in reasoning which is good. We can talk about that later on, but I expect people will catch up within…”
Calacanis: OpenAI's o1 model is a year ahead of competitors
“And the previous version I felt was three to six months ahead of competitors. This is a year ahead of competitors.”
OpenAI o1 Benchmarks Show Major Gains Across Math, Coding, and Science
“So if you think about, ah, competition math, ah, GPT-IV-O was getting a 13.4 accuracy score on competition math. But this thing and the way that it can think through the different problems is getting 83.3, ah, score on accuracy. Ah, with competition math. So i…”
OpenAI Notably Avoided Using the Word 'Agent' in o1 Announcement
“Open AI hasn't actually used the word agent once in its announcement, and it doesn't sound like any of its scientists have talked about that either.”
Kantrowitz: OpenAI's Blog Post States o1 Does Not Solve Hallucinations
“Yeah, and OpenAI even in its blog post says that it does not solve, this does not solve hallucinations, and then you can”