why aren't all 10 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Mitchell: AI reasoning improvements will not be limited to math and code
“So like there, I think there's some reason for spikiness, but I think some people will probably go too far with this and saying like, oh yes, these models will only be really good at math and code. And like, not, you know, like everything else is like, you can…”
Insight
Mitchell: LLMs allocate compute efficiently by deferring tasks to specialized tools
“I think, like, part of this is you can just allocate compute a lot more efficiently because you can defer stuff that the model doesn't have comparative advantage to doing to a tool that is, like, really well suited to doing that thing.”
Assertion Supported
Mitchell: o3 autonomously executes multi-step tasks using integrated tools
“Not only is the model it's on its own smarter than our previous O series models, which is great, but it's also able to use all these tools that like further enhance its abilities and whether that's doing like research on something where you want up-to-date inf…”
Disclosure
Mitchell: OpenAI plans to unify models and remove the ChatGPT switcher
“You know, I think for us, like unification of our models is something that, you know, Sam has talked about publicly that, you know, we have this big crazy model switcher in ChatGPT and there are a lot of choices and you know, we have a model that might be good…”
Insight
Mitchell: OpenAI limits model agency due to asymmetric error costs
“There's a reason we don't go hog wild and say, like, oh yes, here's, like, the keys to the kingdom, like, have at it. There are still, you know, asymmetric costs to, like, the time you can save and the types of errors you can make, and so we're trying to, like…”
Insight
Mitchell: Physical time bottlenecks make AI tasks harder than simulatable domains
“Stuff that is really bottlenecked by like time, like the physical world is also, you know, just harder than stuff that we can simulate really well.”
Insight
Mitchell: High-quality evaluation benchmarks are underappreciated compared to training data
“I mean, yeah, like you want, you know, good data to train on and that's of course valuable for making the model better, but I think it is often neglected how also important it is to have high quality data, which is like a different definition of high quality w…”
Insight
Mitchell: o3 output distribution makes single-prompt evaluations misleading
“O-three can do really cool things, like when it chains together a lot of tool calls, and then, like, sometimes for the same prompt, it won't have that, you know, moment of magic, or it will, you know, just take a little, it'll do a little less work for you, an…”
Insight
Mitchell: AI offloads tasks lacking comparative advantage to external tools
“I think like part of this is you can just allocate compute a lot more efficiently because you can defer stuff that the model doesn't have comparative advantage to doing to a tool that is like really well suited to doing that thing.”
Insight
Mitchell: Real-world robotics imposes strict latency constraints absent in disembodied AI
“The real world is like an interesting litmus test because at the end of the day, like there is a, you know, frame rate in the real world you need to live on. And it doesn't matter if you get the right answer after you think for two minutes, like, You know, the…”