Moonshot's Kimi K3 Ranks Top Three Across Major AI Benchmarks
“This is from Nathan Lambert's Substack Interconnects. It comes number two on the VALS AI Index. Number three overall on artificial analysis is Intelligence Index number one in the front end cone arena, and it has many more impressive results.”
Artificial Analysis uses mystery shopper accounts to prevent endpoint manipulation
“We have what we call a mystery shopper policy, and so, and we're totally transparent with all the labs we work with about this, that we will register accounts not on our own domain and run both intelligence evals and performance benchmarks without them being u…”
Artificial Analysis Intelligence Index synthesizes 10 evaluation datasets
“The artificial analysis intelligence index, which is our synthesis metric that we pulled together currently from 10 different eval data sets to give what? We're pretty confident is the best single number to look at for how smart the models are.”
Hill-Smith: Early Benchmarks Like HumanEval Are Saturated and Trivial
“Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Hill-Smith: OpenAI Was Untouchable for Well Over a Year
“If we go back even a little bit before then, we're in the era where, when you look at this chart, like, OpenAI was untouchable for well over a year.”
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Cameron: General model intelligence does not correlate with hallucination rates
“One interesting aspect is that we've found that there's not really a, not a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they do…”
Hill-Smith: Omniscience Factual Accuracy Tracks Model Parameter Count Most Closely
“If you go back up to accuracy on Omniscience, an interesting thing about this accuracy metric is that it tracks more closely than anything else that we measure, the total parameter count of models.”
Hill-Smith estimates Gemini 3 Pro parameter count hits 5 to 10 trillion
“You might reasonably form a view that there's a pretty good chance that Gemini three pro is bigger than that, that it could be in the five to 10 trillion parameter range. To be clear, I have absolutely no idea, but just based on this chart, like that's where y…”
Gemini 3 Pro performs poorly on GDPval-AA benchmark evaluator tasks
“One data point there is that even as the, as an evaluator, Gemini three pro interestingly doesn't do actually that well in GDP val AA.”
Models perform better in custom agent harnesses than native web chatbots
“And what's really interesting is that if you compare, for instance, Claude, 4.5 Opus using the Claude web chatbot, it performs worse than the model in our Agentic harness. And so in every case, the model performs better in our agentic harness than its web chat…”
Hill-Smith: Multi-tool MCP agent workflows barely work right now
“I would say that this stuff like barely works in fairness right now.”
Artificial Analysis open-sources minimalist agent harness Stirrup on GitHub
“We released that on, on GitHub yesterday. It's called Stirrup, so if people want to check it out, and it's a great you know, base for, you know, generalist building a generalist agent.”
AI2's OLMo 3 32B leads Artificial Analysis's 18-point Openness Index
“It's out of 18 currently. And so we've got an openness index page, but essentially these are points. You get points for being more open across these different categories and the maximum you can achieve is 18. So AI two with their extremely open OMO three, 32 B…”
GPT-4-level intelligence is now over 100 times cheaper than at launch
“The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now.”
Cameron: Model performance correlates with total parameters, not active parameters
“We, in our benchmark, see a lot of performance correlated more with total parameters than active, and not that correlated with how sparse like the models are. Our accuracy benchmark is part of a omniscience. It's very correlated with total. It's not correlated…”
Reasoning models consume 10x more tokens on average than non-reasoning models
“So, earlier this year, and probably when you and George last spoke for the AI engineers world's fair, we had this great slide that was super easy, where we would show that the average reasoning model is using 10 times the number of tokens per query in our inte…”
Arena's organic user prompts provide realism that Artificial Analysis lacks
“They have arenas, but the arenas are not based on organic usage. Like the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a level of rea…”
swyx: Artificial Analysis video arena uses pre-generated videos instead of user inputs
“So like, for example, for AA, their video arena is pre-generated videos. You can't enter in your own video.”
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Inception generalist model matches Claude Haiku quality at 5-10x speed
“We had our generalist model evaluated by artificial analysis and the intelligence score from AA artificial analysis around 40. So it's comparable to GPT, 4.1 nano, cloud haiku, kind of like Close source speed optimized models. It's roughly comparable in terms …”
Feldman: Cerebras inference has been the fastest platform since August 2024 launch
“And every day since August 26th when we launched Inference, our way has been the fastest way across a whole set of models tested by artificial analysis and others.”