Chaubard: HRM scored 70% on ARC Prize 1 without pre-training
“There is no pre-training at all. This starts from, like, literally Tagula-Rasa weights, and it can outperform at that time, if we go back, you know, we had O-three, if you remember back, way back when. And it, O-three gets zero. Literally zero, and this got, l…”
Sherman Wu: OpenAI's o3 model stands out for diligent tool execution
“One of my favorite models is actually O three. Cause it was like one of the most diligent models. It would just like do all these tool calls and it's like really the intelligence itself trying to like do the, you know, tool calls or reg or anything like that o…”
Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”
OpenAI o3 model completes 40% of internal research engineer pull requests
“That's another data point, by the way, from this was from the O three system card. They showed a jump from like low to mid single digits to roughly 40% of PRs actually checked in by Research engineers at OpenAI that the model could do. So prior to O three, not…”
Chen: Earlier Codex models spent too little time on hard problems
“What we found is the latest, the previous generation of the codex models, they were spending too little time solving the hardest problems and too much time solving the easy, easy problems. And I think that, that is actually just probably out of the box what yo…”
Patel: OpenAI o1 and o3 share basic architecture with GPT-4o
“For a long time, OpenAI was charging more per token for the reasoning model, right, O-one and O-three than they were for GPT-Four-O, even though the architecture is, like, basically the same. It's just the weights are different.”
OpenAI's o3 model failed to calculate a basic spatial measurement
“I tried O-three out the other day. Yeah. I took a photo of a thing I had hung up, And I said, how much space from the bottom of that photo, of that picture, the poster, to the floor? It took four minutes. It wrote multiple Python scripts to give me the wrong a…”
Tan: OpenAI o3 operates around 130 IQ, smarter than many past hires
“I think of O-three as basically about a 130 IQ, maybe O-three pro can be even smarter than that. When I really think about that, it's like, oh yeah, like a lot of the people who I've ever hired in my lifetime are like, yeah, O-three is smarter than that person…”
Schulhoff: Explicit chain-of-thought prompting is still needed for GPT-4 and GPT-4o
“Actually for those models, I'd say no need, but if you're using GPT-IV, GPT-IV-O, then it's still worth it.”
Greg Isenberg: OpenAI's o3 model beats GPT-4o for finding startup ideas
“Now you're gonna wanna make sure that O three is clicked here. Sometimes by default, I think four O is clicked. But I find for questions like this, you're gonna wanna use O three. It's just gonna be a bit more effective.”
Claude loses AI Diplomacy games because it refuses to deceive opponents
“I haven't seen Claude with any game yet because they won't do it. Like there's like, O three has managed to get them on board for like draws, even though they all know the only win condition in the game is, is 18 supply centers.”
Meng To: OpenAI's o3 is best for image-to-HTML conversion
“In fact, I do use O three, which is one of the best models for taking the image, a screenshot of your favorite website or your own design Figma, and then turn that to HTML and then bring that to lovable to V zero or to aura, whichever you want.”
OpenAI's o3 acts as a '10-minute AGI' for human tasks
“Whether or not you want to call this, like, end minute AGI, I kind of like the phrase, 10 minute AGI, for just, like, how to think about O-three is that anything that you can do as a human in 10 minutes, O-three is usually going to be able to do reasonably wel…”
Mitchell: o3 output distribution makes single-prompt evaluations misleading
“O-three can do really cool things, like when it chains together a lot of tool calls, and then, like, sometimes for the same prompt, it won't have that, you know, moment of magic, or it will, you know, just take a little, it'll do a little less work for you, an…”
Jin: RL enables models to surpass expert labelers and develop self-direction
“The model outperforming expert labelers is, is possible. The model learning, like, self-direction is, like, expected. And yeah, we've seen, like, kind of cool emergent behaviors with, like, you know, like, O-one, O-three, R-one, kind of, like, these, like, thi…”
Coogan: Reasoning benchmarks are unfair without inference cost and compute time
“There's this big like value trade off between, yeah, if you let the L, if you let the LLM reason for hours and you spend 2000 dollars per query, you can get remarkable results. But then is that really a fair benchmark against something that costs a dollar to i…”
David Sacks argues DeepSeek-R1 only matches a four-month-old OpenAI model
“The R-one model is, is basically comparable to O-one, which OpenAI released four months ago and was training on internally, call it nine or 10 months ago. So OpenAI is on O-three now. Its frontier is ahead of where R-one is.”
Fernando: OpenAI o3 launch will likely trigger reasoning model price drops
“Open AI is currently promised that the O three model will come out and the mini model will come out, which would be on par with this model. So that prices will also probably significantly drop as well because they just get more efficient with time.”