why aren't all 11 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Prediction Not checkable as stated
McGrath: Specialized and Frontier Reasoning AI Models Will Eventually Converge
“You know, I think if you look at like deep research, the original one and GPT-Five thinking on like high reasoning today, I think you'll see that like eventually the models all sort of converge in their capabilities.”
Insight
McGrath: RLHF and RLVR differ by data quality, not optimization math
“Really, at the end of the day, like, RLHF, RLVR,
They're both policy gradient methods, but the, what's different is just like the input data.”
Insight
McGrath: DeepSeek Math's real breakthrough is verifiable reward trust, not GRPO
“As you said, it came out in the deep seek math paper, and like, it's an interesting optimization method, but it's like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust. Like when, you know…”
Insight
McGrath: RL runs have far more infrastructure failure points than pre-training
“The issue with RL is, like, you're doing tasks, and each task could have, like, a different grading setup, and each one of those different grading setups, that's, like, more infrastructure, and so, You know, when I'm staying up late trying to figure out what's…”
Insight
McGrath: Design specs let Codex complete hours of coding in 15 minutes
“If I spend, like, you know, 30, 40 minutes writing something that looks like a design doc or something, Codex can do more work than I can do in a few hours in, like, 15 minutes.”
Assertion Supported
McGrath: GPT-5 Thinking Matches or Beats Deep Research on Published Evals
“I mean, I think if you look at our published evals, they're, they look, like, basically on par if it's not better, so, like, I mean, that's personally what I do.”
Assertion Partly supported
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Prediction Not checkable as stated
McGrath: AGI will be a single tool that decides its own thinking time
“Yeah, I think, like, eventually, you know, we'll have AGI, and like, you're not gonna have to worry too much about how hard to think directly. It'll just, you know, we'll have a one tool that you always go to, and it knows how long to think for, and things lik…”
Insight
McGrath: AI Frontier Is Bottlenecked by Shortage of Hybrid Systems-ML Talent
“I think we're still having trouble not at OpenAI, but I think as a whole, producing lots of people that do lot, want to do lots of both systems work and ML work. And I think if you're trying to push the frontier, you don't know which Place is currently bottlen…”
Assertion Partly supported
McGrath: OpenAI 10xed Effective Context Window for GPT-4.1
“I worked on long context, that was why I was on last, was for 4.1, where we, you know, I think, tenxed the effective context window for 4.1”
Disclosure
McGrath: OpenAI Continues to Release Non-Thinking Models for Specific APIs
“No, we're still, we still are releasing non-thinking models but that one was the one that we did that was like API-specific non-thinking so, you know, focus has shifted a little.”