Everything Eric Mitchell said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Mitchell: AI reasoning improvements will not be limited to math and code
“So like there, I think there's some reason for spikiness, but I think some people will probably go too far with this and saying like, oh yes, these models will only be really good at math and code. And like, not, you know, like everything else is like, you can…”
Mitchell: LLMs allocate compute efficiently by deferring tasks to specialized tools
“I think, like, part of this is you can just allocate compute a lot more efficiently because you can defer stuff that the model doesn't have comparative advantage to doing to a tool that is, like, really well suited to doing that thing.”
Mitchell: o3 autonomously executes multi-step tasks using integrated tools
“Not only is the model it's on its own smarter than our previous O series models, which is great, but it's also able to use all these tools that like further enhance its abilities and whether that's doing like research on something where you want up-to-date inf…”
Mitchell: OpenAI plans to unify models and remove the ChatGPT switcher
“You know, I think for us, like unification of our models is something that, you know, Sam has talked about publicly that, you know, we have this big crazy model switcher in ChatGPT and there are a lot of choices and you know, we have a model that might be good…”
Mitchell: OpenAI limits model agency due to asymmetric error costs
“There's a reason we don't go hog wild and say, like, oh yes, here's, like, the keys to the kingdom, like, have at it. There are still, you know, asymmetric costs to, like, the time you can save and the types of errors you can make, and so we're trying to, like…”
Mitchell: Physical time bottlenecks make AI tasks harder than simulatable domains
“Stuff that is really bottlenecked by like time, like the physical world is also, you know, just harder than stuff that we can simulate really well.”
Mitchell: High-quality evaluation benchmarks are underappreciated compared to training data
“I mean, yeah, like you want, you know, good data to train on and that's of course valuable for making the model better, but I think it is often neglected how also important it is to have high quality data, which is like a different definition of high quality w…”
Mitchell: o3 output distribution makes single-prompt evaluations misleading
“O-three can do really cool things, like when it chains together a lot of tool calls, and then, like, sometimes for the same prompt, it won't have that, you know, moment of magic, or it will, you know, just take a little, it'll do a little less work for you, an…”
Mitchell: AI offloads tasks lacking comparative advantage to external tools
“I think like part of this is you can just allocate compute a lot more efficiently because you can defer stuff that the model doesn't have comparative advantage to doing to a tool that is like really well suited to doing that thing.”
Mitchell: Real-world robotics imposes strict latency constraints absent in disembodied AI
“The real world is like an interesting litmus test because at the end of the day, like there is a, you know, frame rate in the real world you need to live on. And it doesn't matter if you get the right answer after you think for two minutes, like, You know, the…”