Eric Mitchell was not on this episode. A recording of them was played into it, so these
are their words but not an appearance on No Priors. It still counts as said, and it is kept out of every score on their page.
OpenAI researcher Eric Mitchell explains why reasoning models write and execute code rather than computing calculations within their context window.
Opinion
Mitchell: AI reasoning improvements will not be limited to math and code
“So like there, I think there's some reason for spikiness, but I think some people will probably go too far with this and saying like, oh yes, these models will only be really good at math and code. And like, not, you know, like everything else is like, you can…”
Assertion Supported
Mitchell: o3 autonomously executes multi-step tasks using integrated tools
“Not only is the model it's on its own smarter than our previous O series models, which is great, but it's also able to use all these tools that like further enhance its abilities and whether that's doing like research on something where you want up-to-date inf…”
Disclosure
Mitchell: OpenAI plans to unify models and remove the ChatGPT switcher
“You know, I think for us, like unification of our models is something that, you know, Sam has talked about publicly that, you know, we have this big crazy model switcher in ChatGPT and there are a lot of choices and you know, we have a model that might be good…”
Insight
Mitchell: OpenAI limits model agency due to asymmetric error costs
“There's a reason we don't go hog wild and say, like, oh yes, here's, like, the keys to the kingdom, like, have at it. There are still, you know, asymmetric costs to, like, the time you can save and the types of errors you can make, and so we're trying to, like…”
Insight
Mitchell: Physical time bottlenecks make AI tasks harder than simulatable domains
“Stuff that is really bottlenecked by like time, like the physical world is also, you know, just harder than stuff that we can simulate really well.”
Insight
Mitchell: High-quality evaluation benchmarks are underappreciated compared to training data
“I mean, yeah, like you want, you know, good data to train on and that's of course valuable for making the model better, but I think it is often neglected how also important it is to have high quality data, which is like a different definition of high quality w…”