why aren't all 6 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Assertion Partly supported
Swyx: Claude Sonnet and Gemini Outperform o1-Preview in Coding
“Claude Sonnet so far is beating O-one on coding tasks without At least one preview without being a reasoning model and same for Gemini pro or Gemini two point O.”
Assertion Supported
Molmo-72B Ranks Second Behind Only GPT-4o in Human Preference Elo
“Most preference ranked was GPT-Four-O, then Momo-Semety-Two-B, then Gemini, then Sonnet, then the Seven-B.”
Assertion Supported
Non-reasoning Grok, GPT, and Gemini models exhibit recursive self-correction loops
“And so then I tried this across models, and I saw consistently across Grok, and GPT, and Gemini, that you were seeing this phenomena where models will, like, self-correct themselves quite a bit, and these were, like, non-thinking models. They were, like, the s…”
Assertion Supported
Mallick: Gemini officially supports 24 languages but responds in Klingon
“We officially support 24 languages, but you can try talking to the model and cling on and it'll respond to you.”
Assertion Supported
Howard: Google Gemini is about to release KV caching support
“Gemini is about to finally come out with KV caching, and this is something that Austin actually and Gemma.cpp had had on his roadmap for years well not years, months, long time is, is that.”