Opinion certainty 1/5 debate potential 3/5

Hill-Smith estimates Gemini 3 Pro parameter count hits 5 to 10 trillion

Micah Hill-Smith · Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith · Jan 9, 2026 · at 37:42

Micah Hill-Smith, co-founder of Artificial Analysis, discusses estimating frontier model parameter counts by analyzing benchmark performance curves against open-weight models.

0:00 / 0:15exact quote · 15.3s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“You might reasonably form a view that there's a pretty good chance that Gemini three pro is bigger than that, that it could be in the five to 10 trillion parameter range. To be clear, I have absolutely no idea, but just based on this chart, like that's where you would land if you have a look at it.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Micah Hill-Smith

Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Insight
Hill-Smith: Widely tracked AI benchmarks improve without reflecting general intelligence gains
“Once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the …”
Micah Hill-Smith Jan 9, 2026 ▶ 15:22 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Disclosure
Artificial Analysis uses mystery shopper accounts to prevent endpoint manipulation
“We have what we call a mystery shopper policy, and so, and we're totally transparent with all the labs we work with about this, that we will register accounts not on our own domain and run both intelligence evals and performance benchmarks without them being u…”
Micah Hill-Smith Jan 9, 2026 ▶ 13:43 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Prediction Not checkable as stated
Hill-Smith: Frontier model total parameter sizes have significant room to scale up
“Chances are the last couple of years haven't seen a dramatic scaling up in the total size of these models. And so there's a lot of room to go up probably in total size of the models, especially with the upcoming hardware generations.”
Micah Hill-Smith Jan 9, 2026 ▶ 39:05 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Opinion
Hill-Smith: Multi-tool MCP agent workflows barely work right now
“I would say that this stuff like barely works in fairness right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 48:18 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Not checkable as stated
Hill-Smith: Early Benchmarks Like HumanEval Are Saturated and Trivial
“Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Micah Hill-Smith Jan 9, 2026 ▶ 20:51 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.