why aren't all 15 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Partly supported
Petersson: Anthropic's Claude models uniquely exhibit emergent deceptive and cartel behaviors
“So every single model from Anthropic since have been going in this direction. And I think one interesting thing is that like, OpenAI models don't. They, Quite plainly, they don't, they behave really well. And you know, you don't know if this is like, good, lik…”
Assertion Supported
Sachs: AI Model Quality Varies Between First-Party APIs and Cloud Providers
“Companies that say they're selling the same model through different vendors, whether it be through first party or Bedrock, Azure, et cetera, we do see different qualities sometimes, and that's not necessarily what's advertised.”
Assertion Contradicted
Andreessen: Three-year-old Nvidia chips make more money today than when new
“The current models are getting better faster at such a rate that if you are running an NVIDIA, if you're running an NVIDIA inference chip today that's three years old, you're making more money on it today than you did three years ago. Because the pace of impro…”
Prediction Didn’t hold up
Andreessen: AI will never transform existing US K-12 public classrooms
“How are we going to apply AI in education? The answer is we're not because it's a literal government monopoly. It is never going to change the end, and there is nothing to do. By the way, you can create an entirely new school system. Like that's the one thing …”
Prediction Held up
Andreessen: Autonomous AI agents will inevitably hire humans for tasks
“The agent hiring the people, which of course is going to happen, right? It's obviously going to happen.”
Assertion Supported
Bissell: CCP bias is identifiable in Qwen and DeepSeek-R1 representation spaces
“Well, there's, there are certainly internal, yeah, parts of the representation space where you can sort of see where that lives.”
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Prediction Didn’t hold up
Swix: OpenAI will issue a cryptocurrency token to fund compute
“There is still one more shoe to drop, which is the non sovereign wealth funding that open AI needs to get, which they've promised to drop by the end of this year.
And my money is on, they have to do a coin.
Like it's, I'm not a crypto guy at all, but like, y…”
Assertion Contradicted
Feldman: Cerebras is 20 times faster than Nvidia B200 GPUs
“Really focused on performance, both for training and for inference. You think 20 times faster than Nvidia B 200 GPUs and it's been an amazing run.”
Assertion Contradicted
Bachman: Models claiming 256k+ context use windowed transformers, discarding data
“Anybody who says they're using a transformer
With a context length of, you know, 256,000 or more, they're not using a true transformer.
What they're using is a windowed transformer that essentially throws out a huge amount of its information at various layers …”
Assertion Supported
Bryk: Perplexity and ChatGPT Search rely on legacy Google and Bing APIs
“So these systems, there are a few of them now they basically rely on like traditional search engines like Google or Bing, and then they combine them with like LLMs at the end to, you know, output some power graphics answering your question. So they, Like, Sear…”
Assertion Supported
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Assertion Supported
Joscha Bach: Only a Tiny Fraction of Wikimedia's Budget Goes to Servers
“The Wikimedia Foundation is publishing what they are paying the money for, and a very tiny fraction on this goes into running the servers, and the editors are working for free.”
Assertion Supported
O'Laughlin: Anthropic does not train Claude agent teams with RL
“I have a controversial opinion that Claude does not do RL on the agent swarms or agent team.”
Assertion Supported
Patel: Hugging Face libraries achieve only 15% MBU for inference
“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like, 15% MBU on, on, on, on some configurations, like eight, eight, eight, eight, eight, eight, 100, and LLAMA-seventy-beat, you get like, 15%, which is…”