O-three
also referred to as: o three
4 statements across 3 episodes · 3 bullish · 1 bearish · 3 people on the record · first statement Sep 25, 2025 by Mark Chen · across every show →
Everything said about O-three, oldest first
Sep 25, 2025 negative
Chen: Earlier Codex models spent too little time on hard problems
“What we found is the latest, the previous generation of the codex models, they were spending too little time solving the hardest problems and too much time solving the easy, easy problems. And I think that, that is actually just probably out of the box what yo…”
Oct 14, 2025 positive
OpenAI o3 model completes 40% of internal research engineer pull requests
“That's another data point, by the way, from this was from the O three system card. They showed a jump from like low to mid single digits to roughly 40% of PRs actually checked in by Research engineers at OpenAI that the model could do. So prior to O three, not…”
Oct 14, 2025 positive
Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”
Nov 28, 2025 positive
Sherman Wu: OpenAI's o3 model stands out for diligent tool execution
“One of my favorite models is actually O three. Cause it was like one of the most diligent models. It would just like do all these tool calls and it's like really the intelligence itself trying to like do the, you know, tool calls or reg or anything like that o…”