Micah Hill-Smith

Co-Founder and CEO, Artificial Analysis · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutiveanalyst@_micah_h ↗artificialanalysis.ai ↗

Micah Hill-Smith is the co-founder and CEO of Artificial Analysis, an independent benchmarking platform for AI models, inference providers, and hardware. Prior to founding the company, he studied computer science and law and worked as a business analyst at McKinsey & Company.

14statements → 8claims → 5claims resolved → 80%fully supported → 3.57/5average certainty → 2.07/5average debate potential → ≈4.0/5argument clarity, estimated →

4 supported 0 partly supported 1 contradicted 1 not yet assessed 2 not checkable as stated how the 8 claims stand · each chip opens the sources

1 prediction · 7 assertions · 2 opinions · 2 insights · 2 disclosures · every statement was checked. The prediction and assertions are the 8 claims: statements the public record can support or contradict. 5 are resolved, 1 is not yet assessed, and 2 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Micah argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Micah Hill-Smith Jan 9, 2026 ▶ 30:09 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith

Their most notable contradicted claim

Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
0% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Micah Hill-Smith on measured tape to publish a rate. This says nothing about how they speak.

Everything Micah Hill-Smith said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Insight
Hill-Smith: Widely tracked AI benchmarks improve without reflecting general intelligence gains
“Once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the …”
Micah Hill-Smith Jan 9, 2026 ▶ 15:22 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Disclosure
Artificial Analysis uses mystery shopper accounts to prevent endpoint manipulation
“We have what we call a mystery shopper policy, and so, and we're totally transparent with all the labs we work with about this, that we will register accounts not on our own domain and run both intelligence evals and performance benchmarks without them being u…”
Micah Hill-Smith Jan 9, 2026 ▶ 13:43 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Opinion
Hill-Smith estimates Gemini 3 Pro parameter count hits 5 to 10 trillion
“You might reasonably form a view that there's a pretty good chance that Gemini three pro is bigger than that, that it could be in the five to 10 trillion parameter range. To be clear, I have absolutely no idea, but just based on this chart, like that's where y…”
Micah Hill-Smith Jan 9, 2026 ▶ 37:42 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Prediction Not checkable as stated
Hill-Smith: Frontier model total parameter sizes have significant room to scale up
“Chances are the last couple of years haven't seen a dramatic scaling up in the total size of these models. And so there's a lot of room to go up probably in total size of the models, especially with the upcoming hardware generations.”
Micah Hill-Smith Jan 9, 2026 ▶ 39:05 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Opinion
Hill-Smith: Multi-tool MCP agent workflows barely work right now
“I would say that this stuff like barely works in fairness right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 48:18 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Not checkable as stated
Hill-Smith: Early Benchmarks Like HumanEval Are Saturated and Trivial
“Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Micah Hill-Smith Jan 9, 2026 ▶ 20:51 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Micah Hill-Smith Jan 9, 2026 ▶ 30:09 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Not publicly verifiable
Hill-Smith: Omniscience Factual Accuracy Tracks Model Parameter Count Most Closely
“If you go back up to accuracy on Omniscience, an interesting thing about this accuracy metric is that it tracks more closely than anything else that we measure, the total parameter count of models.”
Micah Hill-Smith Jan 9, 2026 ▶ 36:24 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Supported
GPT-4-level intelligence is now over 100 times cheaper than at launch
“The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 58:55 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Insight
Hill-Smith: Building LLM applications turns every component into a benchmarking problem
“The more you go into building something using LLMs, the more each bit of what you're doing ends up being a benchmarking problem.”
Micah Hill-Smith Jan 9, 2026 ▶ 4:50 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Disclosure
Artificial Analysis Intelligence Index synthesizes 10 evaluation datasets
“The artificial analysis intelligence index, which is our synthesis metric that we pulled together currently from 10 different eval data sets to give what? We're pretty confident is the best single number to look at for how smart the models are.”
Micah Hill-Smith Jan 9, 2026 ▶ 19:13 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Supported
Hill-Smith: OpenAI Was Untouchable for Well Over a Year
“If we go back even a little bit before then, we're in the era where, when you look at this chart, like, OpenAI was untouchable for well over a year.”
Micah Hill-Smith Jan 9, 2026 ▶ 23:52 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Supported
Reasoning models consume 10x more tokens on average than non-reasoning models
“So, earlier this year, and probably when you and George last spoke for the AI engineers world's fair, we had this great slide that was super easy, where we would show that the average reasoning model is using 10 times the number of tokens per query in our inte…”
Micah Hill-Smith Jan 9, 2026 ▶ 1:06:27 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith

Appearances (1)

EpisodeDateSpeaking time
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hi Jan 9, 2026 38m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.