Artificial Analysis

includes Artificial Analysis Intelligence Index, Artificial Analysis Quality Index

20 statements across 4 episodes · 8 bullish · 2 bearish · 6 people on the record · first statement Aug 4, 2025 by Stefano Ermon · said 77 times in 17 episodes since 2024 · across every show →

Mentions by year, the whole family

brought up most by Shawn Wang (31), George Cameron (16), Micah Hill-Smith (14), Andrew Feldman (4), Stefano Ermon (2), Lin Qiao (2), Pranav Reddy (1), Philip Kiely (1)

tap a year for its mentions
00254508202420252026episodesmentions
048202420252026episodes it came up in
007.54158202420252026episodesmentions per episode
2026 47 mentions in 4 episodes 12 per episode
2025 24 mentions in 8 episodes 3 per episode
2024 6 mentions in 5 episodes 1 per episode

every mention, scene by scene, with the transcript →

Everything said about Artificial Analysis, oldest first

Aug 4, 2025 bullish
Assertion Partly supported
Inception generalist model matches Claude Haiku quality at 5-10x speed
“We had our generalist model evaluated by artificial analysis and the intelligence score from AA artificial analysis around 40. So it's comparable to GPT, 4.1 nano, cloud haiku, kind of like Close source speed optimized models. It's roughly comparable in terms …”
Stefano Ermon Aug 4, 2025 ▶ 16:55 ⚡️Mercury: Ultra-Fast Diffusion LLMs — Estefano Ermon, CEO Inception Labs
Oct 1, 2025 bullish
Assertion Supported
Cerebras leads all Artificial Analysis inference benchmarks by a large margin
“I think also just go up and look at artificial analysis. Wherever we are, we're the fastest not by a little bit, but by a lot.”
Andrew Feldman Oct 1, 2025 ▶ 11:03 ⚡️Raising $1.1b to build the fastest LLM Chips on Earth — Andrew Feldman, Cerebras
Dec 31, 2025 neutral
Assertion Supported
swyx: Artificial Analysis video arena uses pre-generated videos instead of user inputs
“So like, for example, for AA, their video arena is pre-generated videos. You can't enter in your own video.”
Shawn Wang Dec 31, 2025 ▶ 7:50 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Dec 31, 2025 positive
Opinion
Arena's organic user prompts provide realism that Artificial Analysis lacks
“They have arenas, but the arenas are not based on organic usage. Like the thing that distinguishes our platform versus theirs is that the users are actually inputting their own use case. They're actually asking their own question. And that gives a level of rea…”
Anastasios Angelopoulos Dec 31, 2025 ▶ 7:24 [State of Evals] LMArena's $1.7B Vision — Anastasios Angelopoulos, LMArena
Jan 9, 2026 neutral
Assertion Supported
Cameron: Model performance correlates with total parameters, not active parameters
“We, in our benchmark, see a lot of performance correlated more with total parameters than active, and not that correlated with how sparse like the models are. Our accuracy benchmark is part of a omniscience. It's very correlated with total. It's not correlated…”
George Cameron Jan 9, 2026 ▶ 1:05:08 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026
Assertion Not publicly verifiable
Hill-Smith: Omniscience Factual Accuracy Tracks Model Parameter Count Most Closely
“If you go back up to accuracy on Omniscience, an interesting thing about this accuracy metric is that it tracks more closely than anything else that we measure, the total parameter count of models.”
Micah Hill-Smith Jan 9, 2026 ▶ 36:24 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026
Disclosure
Artificial Analysis Intelligence Index synthesizes 10 evaluation datasets
“The artificial analysis intelligence index, which is our synthesis metric that we pulled together currently from 10 different eval data sets to give what? We're pretty confident is the best single number to look at for how smart the models are.”
Micah Hill-Smith Jan 9, 2026 ▶ 19:13 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 neutral
Disclosure
Artificial Analysis uses mystery shopper accounts to prevent endpoint manipulation
“We have what we call a mystery shopper policy, and so, and we're totally transparent with all the labs we work with about this, that we will register accounts not on our own domain and run both intelligence evals and performance benchmarks without them being u…”
Micah Hill-Smith Jan 9, 2026 ▶ 13:43 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 negative
Assertion Not checkable as stated
Gemini 3 Pro performs poorly on GDPval-AA benchmark evaluator tasks
“One data point there is that even as the, as an evaluator, Gemini three pro interestingly doesn't do actually that well in GDP val AA.”
George Cameron Jan 9, 2026 ▶ 41:41 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 positive
Assertion Supported
AI2's OLMo 3 32B leads Artificial Analysis's 18-point Openness Index
“It's out of 18 currently. And so we've got an openness index page, but essentially these are points. You get points for being more open across these different categories and the maximum you can achieve is 18. So AI two with their extremely open OMO three, 32 B…”
George Cameron Jan 9, 2026 ▶ 53:35 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026
Assertion Not checkable as stated
Hill-Smith: Early Benchmarks Like HumanEval Are Saturated and Trivial
“Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Micah Hill-Smith Jan 9, 2026 ▶ 20:51 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 bullish
Assertion Not publicly verifiable
Models perform better in custom agent harnesses than native web chatbots
“And what's really interesting is that if you compare, for instance, Claude, 4.5 Opus using the Claude web chatbot, it performs worse than the model in our Agentic harness. And so in every case, the model performs better in our agentic harness than its web chat…”
George Cameron Jan 9, 2026 ▶ 45:40 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 positive
Disclosure
Artificial Analysis open-sources minimalist agent harness Stirrup on GitHub
“We released that on, on GitHub yesterday. It's called Stirrup, so if people want to check it out, and it's a great you know, base for, you know, generalist building a generalist agent.”
George Cameron Jan 9, 2026 ▶ 50:12 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 neutral
Assertion Supported
Reasoning models consume 10x more tokens on average than non-reasoning models
“So, earlier this year, and probably when you and George last spoke for the AI engineers world's fair, we had this great slide that was super easy, where we would show that the average reasoning model is using 10 times the number of tokens per query in our inte…”
Micah Hill-Smith Jan 9, 2026 ▶ 1:06:27 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 positive
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Micah Hill-Smith Jan 9, 2026 ▶ 30:09 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 neutral
Assertion Supported
Cameron: General model intelligence does not correlate with hallucination rates
“One interesting aspect is that we've found that there's not really a, not a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they do…”
George Cameron Jan 9, 2026 ▶ 31:28 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 positive
Assertion Supported
GPT-4-level intelligence is now over 100 times cheaper than at launch
“The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 58:55 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 neutral
Opinion
Hill-Smith estimates Gemini 3 Pro parameter count hits 5 to 10 trillion
“You might reasonably form a view that there's a pretty good chance that Gemini three pro is bigger than that, that it could be in the five to 10 trillion parameter range. To be clear, I have absolutely no idea, but just based on this chart, like that's where y…”
Micah Hill-Smith Jan 9, 2026 ▶ 37:42 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 neutral
Assertion Supported
Hill-Smith: OpenAI Was Untouchable for Well Over a Year
“If we go back even a little bit before then, we're in the era where, when you look at this chart, like, OpenAI was untouchable for well over a year.”
Micah Hill-Smith Jan 9, 2026 ▶ 23:52 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Jan 9, 2026 bearish
Opinion
Hill-Smith: Multi-tool MCP agent workflows barely work right now
“I would say that this stuff like barely works in fairness right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 48:18 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.