why aren't all 9 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Assertion Supported
Ben Allal: Hugging Face SmolLM2-1.7B outperforms Llama 3.2 models
“So it's a series of three models, which are the best in class in each model size. For example, our 1.7 B model outperforms Lama one B and also .2.”
Assertion Supported
Cerebras WSE-3 runs Llama inference 70x faster than NVIDIA GPUs
“Cerebris came out that the wafer scale engine three can serve llama 70 B at 2.1 thousand sorry, 202,100 tokens per second and serves llama four or five B at nearly 1000 tokens per second. So this, you know, to give you an understanding, like this is about 70 t…”
Assertion Supported
Patel: Hugging Face libraries achieve only 15% MBU for inference
“Hugging Face's libraries are actually very inefficient, like incredibly inefficient for inference. You get like, 15% MBU on, on, on, on some configurations, like eight, eight, eight, eight, eight, eight, 100, and LLAMA-seventy-beat, you get like, 15%, which is…”
Assertion Supported
Soldani: Llama and Qwen models fail OSI open source AI definition
“Under this definition, for example, Lama or some of the Quen models are not open source because the license says you can, you can't use this model for this, or it says if you use this model, you have to name the output this way or derivative needs to be named …”
Assertion Supported
Angelopoulos: The Chatbot Arena leaderboard is currently not an apples-to-apples comparison
“None of the leaderboard currently is apples to apples, because you have, like, Gemini Flash, you have, you know, all sorts of tiny models, like Llama Like, eight B and four or five B are not apples to apples.”
Assertion Supported
DeepSeek 128k context fits in 8GB KV cache versus Llama's 80GB
“For context like the total I think the total context length of DeepSeq is a 128,000 tokens, or it might be 256,000 with rope extension. That entire context, I think it's a 128,000, fits into eight gigabytes. And previously context, like I think the Lama four o…”
Assertion Supported
Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA
“LAMA was trained on one trillion tokens, but LAMA-III was trained on 15 trillion tokens.”
Assertion Supported
Patel: Running LLaMA-70B at reading speed requires 2.1 TB/s memory bandwidth
“Hey, to run Llama's seventy billion requires two terabytes a second of memory bandwidth, 2.1, at reading, human reading speed.”