why aren't all 4 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Contradicted
Frontier models are cheaper for agentic tasks because they require fewer turns
“Interestingly, in Tau Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the some of the GBD five, for instance got to the answer faster. And …”
Assertion Partly supported
Early GPT-5 testers report noticeable gains across coding, science, and writing
“And the story that we wrote, we kind of talked about how at least the people that we've talked to who tested it so far have been pretty impressed. They seem to think that it's been, you know, there's been improvements in a number of domains and both like scien…”
Assertion Partly supported
McGrath: GPT-5.1 dramatically reduced token usage over GPT-5 while boosting evals
“Yeah, and so you can see, like, from five to 5.1, our overall evals, you know, we bumped some. But if you look at a two D plot of how many tokens it takes for us to get that, it went way down.”
Assertion Supported
Fioca: GPT-5 matches Codex coding capability but adds step-by-step preambles
“With the five series, because it's more general, and it's just about as good as coding as codex for a lot of things. We've taught it to be more communicative. And so it has preambles before tool calls. It'll say things like, I'm about to go look for this.”