Jan 9, 2026 · 1h 18m · latent-space
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of Latent Space, Artificial Analysis co-founders George Cameron and Micah Hill-Smith join Swyx to discuss the creation and mechanics of their independent AI benchmarking platform, detailing their intelligence and openness indices, the economics of model inference, and the evaluation of agentic reasoning workflows.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 25.1% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
When Swyx claims models cannot practically drop below 5% active parameter sparsity, Micah flatly challenges the assertion and cites active models running at 3%.
Hardest push from the hosts ▶ 1:04:02 Swyx pushing back on sparsity expansion limitsSwyx directly intervenes on the smiling curve premise to argue that fine-grained expert sparsity has reached its mathematical and architectural limits.
Biggest teaching moment ▶ 13:00 Micah explaining statistical variance in reasoning evalsMicah educates Swyx on the necessity of high repeat counts and confidence intervals in 4-option evals to prevent noisy leaderboards.
The host holds their own ▶ 6:52 Swyx detailing eval harness tooling historySwyx lays out the exact historical lineage of open-source eval tooling, citing Stanford HELM and EleutherAI's harness to frame the benchmarking problem space.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Origins of Artificial Analysis and the Mixtral Catalyst | 6 | 2 | 1 | 2 | Swyx recounts early interactions and contextualizes the Mixtral release, while Micah and George explain the genesis of the company as an internal tooling side-project. | |
| Limitations of Traditional Lab Benchmarks and Eval Discrepancies | 7 | 3 | 1 | 1 | Swyx demonstrates deep domain knowledge of historical eval setups including EleutherAI harness and Stanford HELM. Micah elaborates on how labs manipulated prompts and chain-of-thought to skew early benchmark scores. | |
| Benchmarking Mechanics, Eval Costs, and Mystery Shoppers | 7 | 4 | 1 | 2 | Swyx brings up real practitioner issues around response parsing, regex fallback, and multiple choice ordering bias. Micah explains statistical variance, confidence intervals, and the mystery shopper mechanism to stop provider manipulation. | |
| Scaling Through AI Grant and Power User Feedback | 5 | 4 | 2 | 3 | Swyx questions whether frontier AI Grant batchmates were really the right target audience. Micah pushes back constructively, explaining that batch founders functioned as ideal power users testing multi-model switching. | |
| Evolution and Composition of the Intelligence Index | 6 | 3 | 1 | 1 | Swyx and Micah discuss benchmark saturation from V1 to V3. Micah details how datasets like HumanEval were solved by small models, forcing index redesign toward agentic workflows. | |
| Charting LLM History from OpenAI Dominance to DeepSeek | 6 | 2 | 1 | 1 | Swyx and the guests review the interactive index charts, charting the transition from OpenAI's monopoly to the multi-model frontier and DeepSeek's V3/R1 breakthrough over Boxing Day. | |
| Omniscience Index, Hallucination Tracking, and Hard Science Evals | 7 | 5 | 2 | 4 | Swyx challenges the guests on why they made their own hallucination metric instead of adopting lab cards, and notes model calibration nuances. George explains how omniscience penalizes false confidence and why Claude models perform distinctly well. | |
| Estimating Frontier Model Sizes and Scaling Law Horizons | 5 | 4 | 2 | 3 | Micah reveals that omniscience accuracy correlates tightly with total parameter count, using it to infer frontier model sizes. Swyx pushes back that size guessing is mostly creator hype and questions active scaling. | |
| GDPval-AA and Autonomous Agent Evaluation | 6 | 4 | 1 | 3 | George breaks down GDPval-AA and their agentic harness. Swyx probes the LLM-as-judge self-preference issues and suggests benchmarking against calibrated human baseline scores. | |
| Enterprise Workflows, MCP Tooling, and Open-Sourcing Stirrup | 6 | 2 | 1 | 2 | Micah describes MCP tooling integrations for personal data sources, while George announces open-sourcing their Stirrup agent harness. Swyx cross-references Terminal Bench and Harbor. | |
| Quantifying Open Source with the Openness Index | 7 | 3 | 2 | 4 | Swyx critiques the openness weighting scheme and calls out omissions like Hugging Face. Micah defends the objective scoring methodology across training code, data transparency, and license tiers. | |
| The Smiling Curve: Cost Collapse and Inference Expansion | 7 | 5 | 4 | 5 | Micah and George introduce the smiling curve dynamics. Swyx pushes back asserting a practical floor on sparsity around 5%, but Micah and George point to real-world models operating below 3% active parameters. | |
| Reasoning Architectures, Token Efficiency, and Turn Count Dynamics | 7 | 4 | 2 | 3 | Swyx probes the fuzzy distinction between reasoning and non-reasoning models and trade-offs between token efficiency and turn efficiency. George uses Tau-Bench data to explain why higher per-token prices can be cheaper per task. |