why aren't all 16 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Disclosure
Chen: OpenAI's three-year goal is models conducting end-to-end research
“When we look at our kind of three-year roadmap, right the end goal that we want to reach is one where You know, the models are just doing end-to-end research, and I think a part of that problem is just being able to have the model come up with good taste.”
Opinion
Chen: Meta poaching has calmed down and OpenAI came out on top
“I think that met us calmed down a little bit. I think we came out on top”
Opinion
Mark Chen: AI models are producing 'Move 37' breakthroughs in math and coding
“There's move-thirty-sevens in, in math. There's in computer science and coding. I think even, yeah, just, it feels like a lot of people woke up at the start of this year and were like, man, agents are working in my profession. And you know, they're essentially…”
Assertion Not checkable as stated
OpenAI's Chen: AI models already discover novel theorems and advance sciences
“The initial direction we took was you should move it to real world research, right? And we've seen that the models, they've gotten a lot better at just kind of discovering novel theorems and pushing the frontiers of hard sciences. Even today, right, that's no …”
Prediction Not checkable as stated
Mark Chen: AI scaling laws will continue to hold
“And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering, more careful scaling, and it always unlocks that next ability to scale further. So I mean, it's held for You know, almost 10 order…”
Disclosure
Chen: OpenAI favors unifying modalities in as few architectures as possible
“For a research lab, I think there are a lot of advantages for it to being under one. So you just have to maintain one infrastructure stack, for instance. I think the cost to, like, maintaining and scaling many infrastructure stacks at once I think that's somet…”
Opinion
Chen: Pre-training is not dead and remains underrated in AI research
“Well, I think if you still have a pre-training is dead view of the world I think pre-training is definitely yeah, yeah, not, not dead. It's underrated.”
Insight
Mark Chen: A PhD is not necessary to excel in AI research
“There are a lot of researchers who just started out without formal training in machine learning or AI research. We've very much believed in training people up to do this. I think the real hard thing is the ability to creatively solve problems and think outside…”
Assertion Not checkable as stated
Jakob Pachocki and Ilya Sutskever overcame internal inertia to build OpenAI o1
“Even at a company like OpenAI, you would have people ask naturally, why do something when you have a machine that works? And fundamentally, you know, it's to the credit of, you know, Jakob, Ilya, many of the people who really had conviction and vision in this …”
Assertion Not checkable as stated
OpenAI's Mark Chen: AI is in an evals crisis with saturated benchmarks
“So I think beyond that, the other scary thing in the field is the number of canonical gold standard benchmarks is low. And we really are kind of in an evals crisis, right? Where all the really great evals that we all know, like growing up, like taking the SAT …”
Insight
Chen: Eval teams and model optimization teams must remain separate
“Yeah, I think there's a kind of interesting philosophy of separate the teams that are creating the evals from the teams that are optimizing the models themselves, because that way you don't, like, co-incentivize them, right? Like, the way the evals theme can w…”
Insight
Mark Chen: Paper replication is the best way to develop AI research taste
“The best mechanism I've found for developing that is really just replication. So I think you should take papers that you really look up to, and just try to fully replicate it.”
Insight
OpenAI's Chen: Reinforcement learning struggles in subjective, hard-to-grade fields
“RLs traditionally had headwinds when it's come to fields that, you know, it's more kind of, Subjective than objective. So if you kind of think of, you know, one kind of, you know example of this is creative writing, where, you know, you could take two pieces o…”
Insight
Chen: Compaction shortcuts expensive native long context in coding products
“Many, many coding products today have features like compaction, right? Where you can compress kind of either insights or working state and stuff like that, you know, it just shortcuts a lot of the very brutally difficult and expensive permit is that you have t…”
Disclosure
Mark Chen brought soup to researchers to counter Zuckerberg's talent poaching
“Oh, you know, it's absolutely a true story. And I have brought soup to our own researchers.”
Disclosure
OpenAI's three research pillars are pre-training, RL, and alignment
“At the very highest level, right, we have an org that focuses on pre-training, right, which is, you know, giving models a lot of world knowledge. We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights toge…”