Mitchell: o3 output distribution makes single-prompt evaluations misleading
Eric Mitchell · No Priors Ep. 113 | With OpenAI's Eric Mitchell and Brandon McKinzie · May 1, 2025 · at 35:21
OpenAI research scientist Eric Mitchell discusses how frontier reasoning models should be evaluated and tested by users.
“O-three can do really cool things, like when it chains together a lot of tool calls, and then, like, sometimes for the same prompt, it won't have that, you know, moment of magic, or it will, you know, just take a little, it'll do a little less work for you, and so, yeah, though, like, the peak performance is really impressive, but there is a distribution of behavior, and I think people often don't appreciate that there is this distribution of outcomes when you put the same prompt in, and getting intuition about that is useful.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →