why aren't all 7 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Opinion
Zelikman: Most AI labs treat humans as mere intermediates to full automation
“And maybe this is a strong statement, but I say for most labs, like the human is kind of, you know, the intermediate until you have like this fully automated, like, you know, system. And so spending a lot of time optimizing things for being really good at unde…”
Insight
Zelikman: AI field underinvests in memory due to task-centric training regimes
“I would say that memory is definitely like a feature that has been under, under-invested in by the field. But I would say that it is kind of difficult to invest in memory in this very, like, task-centric regime. Because if you have, like, A bunch of these, lik…”
Insight
Zelikman: Training reasoning models on just positive examples causes a plateau
“So if you only train on like the positive examples, then you end up in this kind of like potential minimum where there's just no more data that it can actually solve.”
Assertion Supported
Zelikman: Frontier AI models solve questions that stump actual PhD researchers
“Some of the HLE questions that these models are able to solve are genuinely things that are, like, non-trivial for, like, actual, like, PhD researchers.”
Opinion
Zelikman: Current frontier AI models fundamentally lack emotional intelligence
“One of the core things is that they're not smart Like, emotionally, or, like, they're not smart on the level of, like, actually understanding kind of what people care about, or kind of, like, how to actually, like, help people accomplish the things that they c…”
Insight
Zelikman: AI performance gap persists between verifiable and non-verifiable tasks
“There's still a gap between how well these models perform on verifiable tasks versus not verifiable tasks.”
Assertion Supported
Zelikman: Language models can be trained to simulate students for test design
“Like, even back in my PhD, I think one of my, I guess, less well-known works was actually about, we showed that you can train language models to simulate different kinds of students. For tests. Yeah, yeah. And by simulating students, you can actually design be…”