Micah Hill-Smith, co-founder of Artificial Analysis, explains how building an LLM legal assistant highlighted the need for independent model evaluation across speed, accuracy, and cost.
“The more you go into building something using LLMs, the more each bit of what you're doing ends up being a benchmarking problem.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Micah Hill-Smith
AssertionContradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-SmithJan 9, 2026▶ 8:36Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Insight
Hill-Smith: Widely tracked AI benchmarks improve without reflecting general intelligence gains
“Once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the …”
Micah Hill-SmithJan 9, 2026▶ 15:22Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Disclosure
Artificial Analysis uses mystery shopper accounts to prevent endpoint manipulation
“We have what we call a mystery shopper policy, and so, and we're totally transparent with all the labs we work with about this, that we will register accounts not on our own domain and run both intelligence evals and performance benchmarks without them being u…”
Micah Hill-SmithJan 9, 2026▶ 13:43Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Opinion
Hill-Smith estimates Gemini 3 Pro parameter count hits 5 to 10 trillion
“You might reasonably form a view that there's a pretty good chance that Gemini three pro is bigger than that, that it could be in the five to 10 trillion parameter range. To be clear, I have absolutely no idea, but just based on this chart, like that's where y…”
Micah Hill-SmithJan 9, 2026▶ 37:42Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
PredictionNot checkable as stated
Hill-Smith: Frontier model total parameter sizes have significant room to scale up
“Chances are the last couple of years haven't seen a dramatic scaling up in the total size of these models. And so there's a lot of room to go up probably in total size of the models, especially with the upcoming hardware generations.”
Micah Hill-SmithJan 9, 2026▶ 39:05Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Opinion
Hill-Smith: Multi-tool MCP agent workflows barely work right now
“I would say that this stuff like barely works in fairness right now.”
Micah Hill-SmithJan 9, 2026▶ 48:18Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.