Jan 9, 2026 · 1h 18m · latent-space

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith

Micah Hill-Smith · 38m spoken Shawn Wang · 17m spoken George Cameron · 14m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of Latent Space, Artificial Analysis co-founders George Cameron and Micah Hill-Smith join Swyx to discuss the creation and mechanics of their independent AI benchmarking platform, detailing their intelligence and openness indices, the economics of model inference, and the evaluation of agentic reasoning workflows.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 25.1% of the talking time here. How this is scored →

The hosts as informed peer 6.3 Guest teaching 3.5 Guest disagreement 1.6 The hosts pushing back 2.6
05100:0020:0040:001:00:004:06–7:05 · The hosts as informed peer 6/10 Origins of Artificial Analysis and the Mixtral Catalyst Swyx recounts early interactions and contextualizes the Mixtral release, while Micah and George explain the genesis of the company as an internal tooling side-project.7:05–9:14 · The hosts as informed peer 7/10 Limitations of Traditional Lab Benchmarks and Eval Discrepancies Swyx demonstrates deep domain knowledge of historical eval setups including EleutherAI harness and Stanford HELM. Micah elaborates on how labs manipulated prompts and chain-of-thought to skew early benchmark scores.9:14–16:08 · The hosts as informed peer 7/10 Benchmarking Mechanics, Eval Costs, and Mystery Shoppers Swyx brings up real practitioner issues around response parsing, regex fallback, and multiple choice ordering bias. Micah explains statistical variance, confidence intervals, and the mystery shopper mechanism to stop provider manipulation.16:09–18:58 · The hosts as informed peer 5/10 Scaling Through AI Grant and Power User Feedback Swyx questions whether frontier AI Grant batchmates were really the right target audience. Micah pushes back constructively, explaining that batch founders functioned as ideal power users testing multi-model switching.18:58–22:13 · The hosts as informed peer 6/10 Evolution and Composition of the Intelligence Index Swyx and Micah discuss benchmark saturation from V1 to V3. Micah details how datasets like HumanEval were solved by small models, forcing index redesign toward agentic workflows.22:13–27:27 · The hosts as informed peer 6/10 Charting LLM History from OpenAI Dominance to DeepSeek Swyx and the guests review the interactive index charts, charting the transition from OpenAI's monopoly to the multi-model frontier and DeepSeek's V3/R1 breakthrough over Boxing Day.27:27–36:55 · The hosts as informed peer 7/10 Omniscience Index, Hallucination Tracking, and Hard Science Evals Swyx challenges the guests on why they made their own hallucination metric instead of adopting lab cards, and notes model calibration nuances. George explains how omniscience penalizes false confidence and why Claude models perform distinctly well.36:55–39:38 · The hosts as informed peer 5/10 Estimating Frontier Model Sizes and Scaling Law Horizons Micah reveals that omniscience accuracy correlates tightly with total parameter count, using it to infer frontier model sizes. Swyx pushes back that size guessing is mostly creator hype and questions active scaling.39:38–46:43 · The hosts as informed peer 6/10 GDPval-AA and Autonomous Agent Evaluation George breaks down GDPval-AA and their agentic harness. Swyx probes the LLM-as-judge self-preference issues and suggests benchmarking against calibrated human baseline scores.46:43–51:46 · The hosts as informed peer 6/10 Enterprise Workflows, MCP Tooling, and Open-Sourcing Stirrup Micah describes MCP tooling integrations for personal data sources, while George announces open-sourcing their Stirrup agent harness. Swyx cross-references Terminal Bench and Harbor.51:47–57:56 · The hosts as informed peer 7/10 Quantifying Open Source with the Openness Index Swyx critiques the openness weighting scheme and calls out omissions like Hugging Face. Micah defends the objective scoring methodology across training code, data transparency, and license tiers.57:56–1:05:44 · The hosts as informed peer 7/10 The Smiling Curve: Cost Collapse and Inference Expansion Micah and George introduce the smiling curve dynamics. Swyx pushes back asserting a practical floor on sparsity around 5%, but Micah and George point to real-world models operating below 3% active parameters.1:05:45–1:10:47 · The hosts as informed peer 7/10 Reasoning Architectures, Token Efficiency, and Turn Count Dynamics Swyx probes the fuzzy distinction between reasoning and non-reasoning models and trade-offs between token efficiency and turn efficiency. George uses Tau-Bench data to explain why higher per-token prices can be cheaper per task.4:06–7:05 · Guest teaching 2/10 Origins of Artificial Analysis and the Mixtral Catalyst Swyx recounts early interactions and contextualizes the Mixtral release, while Micah and George explain the genesis of the company as an internal tooling side-project.7:05–9:14 · Guest teaching 3/10 Limitations of Traditional Lab Benchmarks and Eval Discrepancies Swyx demonstrates deep domain knowledge of historical eval setups including EleutherAI harness and Stanford HELM. Micah elaborates on how labs manipulated prompts and chain-of-thought to skew early benchmark scores.9:14–16:08 · Guest teaching 4/10 Benchmarking Mechanics, Eval Costs, and Mystery Shoppers Swyx brings up real practitioner issues around response parsing, regex fallback, and multiple choice ordering bias. Micah explains statistical variance, confidence intervals, and the mystery shopper mechanism to stop provider manipulation.16:09–18:58 · Guest teaching 4/10 Scaling Through AI Grant and Power User Feedback Swyx questions whether frontier AI Grant batchmates were really the right target audience. Micah pushes back constructively, explaining that batch founders functioned as ideal power users testing multi-model switching.18:58–22:13 · Guest teaching 3/10 Evolution and Composition of the Intelligence Index Swyx and Micah discuss benchmark saturation from V1 to V3. Micah details how datasets like HumanEval were solved by small models, forcing index redesign toward agentic workflows.22:13–27:27 · Guest teaching 2/10 Charting LLM History from OpenAI Dominance to DeepSeek Swyx and the guests review the interactive index charts, charting the transition from OpenAI's monopoly to the multi-model frontier and DeepSeek's V3/R1 breakthrough over Boxing Day.27:27–36:55 · Guest teaching 5/10 Omniscience Index, Hallucination Tracking, and Hard Science Evals Swyx challenges the guests on why they made their own hallucination metric instead of adopting lab cards, and notes model calibration nuances. George explains how omniscience penalizes false confidence and why Claude models perform distinctly well.36:55–39:38 · Guest teaching 4/10 Estimating Frontier Model Sizes and Scaling Law Horizons Micah reveals that omniscience accuracy correlates tightly with total parameter count, using it to infer frontier model sizes. Swyx pushes back that size guessing is mostly creator hype and questions active scaling.39:38–46:43 · Guest teaching 4/10 GDPval-AA and Autonomous Agent Evaluation George breaks down GDPval-AA and their agentic harness. Swyx probes the LLM-as-judge self-preference issues and suggests benchmarking against calibrated human baseline scores.46:43–51:46 · Guest teaching 2/10 Enterprise Workflows, MCP Tooling, and Open-Sourcing Stirrup Micah describes MCP tooling integrations for personal data sources, while George announces open-sourcing their Stirrup agent harness. Swyx cross-references Terminal Bench and Harbor.51:47–57:56 · Guest teaching 3/10 Quantifying Open Source with the Openness Index Swyx critiques the openness weighting scheme and calls out omissions like Hugging Face. Micah defends the objective scoring methodology across training code, data transparency, and license tiers.57:56–1:05:44 · Guest teaching 5/10 The Smiling Curve: Cost Collapse and Inference Expansion Micah and George introduce the smiling curve dynamics. Swyx pushes back asserting a practical floor on sparsity around 5%, but Micah and George point to real-world models operating below 3% active parameters.1:05:45–1:10:47 · Guest teaching 4/10 Reasoning Architectures, Token Efficiency, and Turn Count Dynamics Swyx probes the fuzzy distinction between reasoning and non-reasoning models and trade-offs between token efficiency and turn efficiency. George uses Tau-Bench data to explain why higher per-token prices can be cheaper per task.4:06–7:05 · Guest disagreement 1/10 Origins of Artificial Analysis and the Mixtral Catalyst Swyx recounts early interactions and contextualizes the Mixtral release, while Micah and George explain the genesis of the company as an internal tooling side-project.7:05–9:14 · Guest disagreement 1/10 Limitations of Traditional Lab Benchmarks and Eval Discrepancies Swyx demonstrates deep domain knowledge of historical eval setups including EleutherAI harness and Stanford HELM. Micah elaborates on how labs manipulated prompts and chain-of-thought to skew early benchmark scores.9:14–16:08 · Guest disagreement 1/10 Benchmarking Mechanics, Eval Costs, and Mystery Shoppers Swyx brings up real practitioner issues around response parsing, regex fallback, and multiple choice ordering bias. Micah explains statistical variance, confidence intervals, and the mystery shopper mechanism to stop provider manipulation.16:09–18:58 · Guest disagreement 2/10 Scaling Through AI Grant and Power User Feedback Swyx questions whether frontier AI Grant batchmates were really the right target audience. Micah pushes back constructively, explaining that batch founders functioned as ideal power users testing multi-model switching.18:58–22:13 · Guest disagreement 1/10 Evolution and Composition of the Intelligence Index Swyx and Micah discuss benchmark saturation from V1 to V3. Micah details how datasets like HumanEval were solved by small models, forcing index redesign toward agentic workflows.22:13–27:27 · Guest disagreement 1/10 Charting LLM History from OpenAI Dominance to DeepSeek Swyx and the guests review the interactive index charts, charting the transition from OpenAI's monopoly to the multi-model frontier and DeepSeek's V3/R1 breakthrough over Boxing Day.27:27–36:55 · Guest disagreement 2/10 Omniscience Index, Hallucination Tracking, and Hard Science Evals Swyx challenges the guests on why they made their own hallucination metric instead of adopting lab cards, and notes model calibration nuances. George explains how omniscience penalizes false confidence and why Claude models perform distinctly well.36:55–39:38 · Guest disagreement 2/10 Estimating Frontier Model Sizes and Scaling Law Horizons Micah reveals that omniscience accuracy correlates tightly with total parameter count, using it to infer frontier model sizes. Swyx pushes back that size guessing is mostly creator hype and questions active scaling.39:38–46:43 · Guest disagreement 1/10 GDPval-AA and Autonomous Agent Evaluation George breaks down GDPval-AA and their agentic harness. Swyx probes the LLM-as-judge self-preference issues and suggests benchmarking against calibrated human baseline scores.46:43–51:46 · Guest disagreement 1/10 Enterprise Workflows, MCP Tooling, and Open-Sourcing Stirrup Micah describes MCP tooling integrations for personal data sources, while George announces open-sourcing their Stirrup agent harness. Swyx cross-references Terminal Bench and Harbor.51:47–57:56 · Guest disagreement 2/10 Quantifying Open Source with the Openness Index Swyx critiques the openness weighting scheme and calls out omissions like Hugging Face. Micah defends the objective scoring methodology across training code, data transparency, and license tiers.57:56–1:05:44 · Guest disagreement 4/10 The Smiling Curve: Cost Collapse and Inference Expansion Micah and George introduce the smiling curve dynamics. Swyx pushes back asserting a practical floor on sparsity around 5%, but Micah and George point to real-world models operating below 3% active parameters.1:05:45–1:10:47 · Guest disagreement 2/10 Reasoning Architectures, Token Efficiency, and Turn Count Dynamics Swyx probes the fuzzy distinction between reasoning and non-reasoning models and trade-offs between token efficiency and turn efficiency. George uses Tau-Bench data to explain why higher per-token prices can be cheaper per task.4:06–7:05 · The hosts pushing back 2/10 Origins of Artificial Analysis and the Mixtral Catalyst Swyx recounts early interactions and contextualizes the Mixtral release, while Micah and George explain the genesis of the company as an internal tooling side-project.7:05–9:14 · The hosts pushing back 1/10 Limitations of Traditional Lab Benchmarks and Eval Discrepancies Swyx demonstrates deep domain knowledge of historical eval setups including EleutherAI harness and Stanford HELM. Micah elaborates on how labs manipulated prompts and chain-of-thought to skew early benchmark scores.9:14–16:08 · The hosts pushing back 2/10 Benchmarking Mechanics, Eval Costs, and Mystery Shoppers Swyx brings up real practitioner issues around response parsing, regex fallback, and multiple choice ordering bias. Micah explains statistical variance, confidence intervals, and the mystery shopper mechanism to stop provider manipulation.16:09–18:58 · The hosts pushing back 3/10 Scaling Through AI Grant and Power User Feedback Swyx questions whether frontier AI Grant batchmates were really the right target audience. Micah pushes back constructively, explaining that batch founders functioned as ideal power users testing multi-model switching.18:58–22:13 · The hosts pushing back 1/10 Evolution and Composition of the Intelligence Index Swyx and Micah discuss benchmark saturation from V1 to V3. Micah details how datasets like HumanEval were solved by small models, forcing index redesign toward agentic workflows.22:13–27:27 · The hosts pushing back 1/10 Charting LLM History from OpenAI Dominance to DeepSeek Swyx and the guests review the interactive index charts, charting the transition from OpenAI's monopoly to the multi-model frontier and DeepSeek's V3/R1 breakthrough over Boxing Day.27:27–36:55 · The hosts pushing back 4/10 Omniscience Index, Hallucination Tracking, and Hard Science Evals Swyx challenges the guests on why they made their own hallucination metric instead of adopting lab cards, and notes model calibration nuances. George explains how omniscience penalizes false confidence and why Claude models perform distinctly well.36:55–39:38 · The hosts pushing back 3/10 Estimating Frontier Model Sizes and Scaling Law Horizons Micah reveals that omniscience accuracy correlates tightly with total parameter count, using it to infer frontier model sizes. Swyx pushes back that size guessing is mostly creator hype and questions active scaling.39:38–46:43 · The hosts pushing back 3/10 GDPval-AA and Autonomous Agent Evaluation George breaks down GDPval-AA and their agentic harness. Swyx probes the LLM-as-judge self-preference issues and suggests benchmarking against calibrated human baseline scores.46:43–51:46 · The hosts pushing back 2/10 Enterprise Workflows, MCP Tooling, and Open-Sourcing Stirrup Micah describes MCP tooling integrations for personal data sources, while George announces open-sourcing their Stirrup agent harness. Swyx cross-references Terminal Bench and Harbor.51:47–57:56 · The hosts pushing back 4/10 Quantifying Open Source with the Openness Index Swyx critiques the openness weighting scheme and calls out omissions like Hugging Face. Micah defends the objective scoring methodology across training code, data transparency, and license tiers.57:56–1:05:44 · The hosts pushing back 5/10 The Smiling Curve: Cost Collapse and Inference Expansion Micah and George introduce the smiling curve dynamics. Swyx pushes back asserting a practical floor on sparsity around 5%, but Micah and George point to real-world models operating below 3% active parameters.1:05:45–1:10:47 · The hosts pushing back 3/10 Reasoning Architectures, Token Efficiency, and Turn Count Dynamics Swyx probes the fuzzy distinction between reasoning and non-reasoning models and trade-offs between token efficiency and turn efficiency. George uses Tau-Bench data to explain why higher per-token prices can be cheaper per task.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 40.7% · guest 59.3%0:00 · the hosts 40.7% · guest 59.3%3:00 · the hosts 12.3% · guest 87.7%3:00 · the hosts 12.3% · guest 87.7%6:00 · the hosts 35.8% · guest 64.2%6:00 · the hosts 35.8% · guest 64.2%9:00 · the hosts 29.9% · guest 70.1%9:00 · the hosts 29.9% · guest 70.1%12:00 · the hosts 15.4% · guest 84.6%12:00 · the hosts 15.4% · guest 84.6%15:00 · the hosts 23.9% · guest 76.1%15:00 · the hosts 23.9% · guest 76.1%18:00 · the hosts 13.6% · guest 86.4%18:00 · the hosts 13.6% · guest 86.4%21:00 · the hosts 16.4% · guest 83.6%21:00 · the hosts 16.4% · guest 83.6%24:00 · the hosts 17% · guest 83%24:00 · the hosts 17% · guest 83%27:00 · the hosts 18.5% · guest 81.5%27:00 · the hosts 18.5% · guest 81.5%30:00 · the hosts 19.1% · guest 80.9%30:00 · the hosts 19.1% · guest 80.9%33:00 · the hosts 26.7% · guest 73.3%33:00 · the hosts 26.7% · guest 73.3%36:00 · the hosts 20.5% · guest 79.5%36:00 · the hosts 20.5% · guest 79.5%39:00 · the hosts 31.9% · guest 68.1%39:00 · the hosts 31.9% · guest 68.1%42:00 · the hosts 16.3% · guest 83.7%42:00 · the hosts 16.3% · guest 83.7%45:00 · the hosts 20.2% · guest 79.8%45:00 · the hosts 20.2% · guest 79.8%48:00 · the hosts 25.7% · guest 74.3%48:00 · the hosts 25.7% · guest 74.3%51:00 · the hosts 19.1% · guest 80.9%51:00 · the hosts 19.1% · guest 80.9%54:00 · the hosts 47.6% · guest 52.4%54:00 · the hosts 47.6% · guest 52.4%57:00 · the hosts 32.5% · guest 67.5%57:00 · the hosts 32.5% · guest 67.5%1:00:00 · the hosts 22.7% · guest 77.3%1:00:00 · the hosts 22.7% · guest 77.3%1:03:00 · the hosts 27.5% · guest 72.5%1:03:00 · the hosts 27.5% · guest 72.5%1:06:00 · the hosts 17.9% · guest 82.1%1:06:00 · the hosts 17.9% · guest 82.1%1:09:00 · the hosts 24% · guest 76%1:09:00 · the hosts 24% · guest 76%1:12:00 · the hosts 48.9% · guest 51.1%1:12:00 · the hosts 48.9% · guest 51.1%1:15:00 · the hosts 29.6% · guest 70.4%1:15:00 · the hosts 29.6% · guest 70.4%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 1:04:29 Micah rejecting Swyx's sparsity lower bound

When Swyx claims models cannot practically drop below 5% active parameter sparsity, Micah flatly challenges the assertion and cites active models running at 3%.

Hardest push from the hosts ▶ 1:04:02 Swyx pushing back on sparsity expansion limits

Swyx directly intervenes on the smiling curve premise to argue that fine-grained expert sparsity has reached its mathematical and architectural limits.

Biggest teaching moment ▶ 13:00 Micah explaining statistical variance in reasoning evals

Micah educates Swyx on the necessity of high repeat counts and confidence intervals in 4-option evals to prevent noisy leaderboards.

The host holds their own ▶ 6:52 Swyx detailing eval harness tooling history

Swyx lays out the exact historical lineage of open-source eval tooling, citing Stanford HELM and EleutherAI's harness to frame the benchmarking problem space.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Origins of Artificial Analysis and the Mixtral Catalyst 6212 Swyx recounts early interactions and contextualizes the Mixtral release, while Micah and George explain the genesis of the company as an internal tooling side-project.
Limitations of Traditional Lab Benchmarks and Eval Discrepancies 7311 Swyx demonstrates deep domain knowledge of historical eval setups including EleutherAI harness and Stanford HELM. Micah elaborates on how labs manipulated prompts and chain-of-thought to skew early benchmark scores.
Benchmarking Mechanics, Eval Costs, and Mystery Shoppers 7412 Swyx brings up real practitioner issues around response parsing, regex fallback, and multiple choice ordering bias. Micah explains statistical variance, confidence intervals, and the mystery shopper mechanism to stop provider manipulation.
Scaling Through AI Grant and Power User Feedback 5423 Swyx questions whether frontier AI Grant batchmates were really the right target audience. Micah pushes back constructively, explaining that batch founders functioned as ideal power users testing multi-model switching.
Evolution and Composition of the Intelligence Index 6311 Swyx and Micah discuss benchmark saturation from V1 to V3. Micah details how datasets like HumanEval were solved by small models, forcing index redesign toward agentic workflows.
Charting LLM History from OpenAI Dominance to DeepSeek 6211 Swyx and the guests review the interactive index charts, charting the transition from OpenAI's monopoly to the multi-model frontier and DeepSeek's V3/R1 breakthrough over Boxing Day.
Omniscience Index, Hallucination Tracking, and Hard Science Evals 7524 Swyx challenges the guests on why they made their own hallucination metric instead of adopting lab cards, and notes model calibration nuances. George explains how omniscience penalizes false confidence and why Claude models perform distinctly well.
Estimating Frontier Model Sizes and Scaling Law Horizons 5423 Micah reveals that omniscience accuracy correlates tightly with total parameter count, using it to infer frontier model sizes. Swyx pushes back that size guessing is mostly creator hype and questions active scaling.
GDPval-AA and Autonomous Agent Evaluation 6413 George breaks down GDPval-AA and their agentic harness. Swyx probes the LLM-as-judge self-preference issues and suggests benchmarking against calibrated human baseline scores.
Enterprise Workflows, MCP Tooling, and Open-Sourcing Stirrup 6212 Micah describes MCP tooling integrations for personal data sources, while George announces open-sourcing their Stirrup agent harness. Swyx cross-references Terminal Bench and Harbor.
Quantifying Open Source with the Openness Index 7324 Swyx critiques the openness weighting scheme and calls out omissions like Hugging Face. Micah defends the objective scoring methodology across training code, data transparency, and license tiers.
The Smiling Curve: Cost Collapse and Inference Expansion 7545 Micah and George introduce the smiling curve dynamics. Swyx pushes back asserting a practical floor on sparsity around 5%, but Micah and George point to real-world models operating below 3% active parameters.
Reasoning Architectures, Token Efficiency, and Turn Count Dynamics 7423 Swyx probes the fuzzy distinction between reasoning and non-reasoning models and trade-offs between token efficiency and turn efficiency. George uses Tau-Bench data to explain why higher per-token prices can be cheaper per task.

Statements from this episode (25)

Insight
Hill-Smith: Building LLM applications turns every component into a benchmarking problem
“The more you go into building something using LLMs, the more each bit of what you're doing ends up being a benchmarking problem.”
Micah Hill-Smith Jan 9, 2026 ▶ 4:50
Opinion
George Cameron: Mixtral 8x7B transformed the landscape for serverless inference providers
“We had Mixtrel A times seven B and it was a key. Like a open source model that really changed the landscape and opened up people's eyes to other serverless inference providers and thinking about speed, thinking about cost.”
George Cameron Jan 9, 2026 ▶ 6:32
Assertion Contradicted
Hill-Smith: Google used unpublished 32-shot CoT to claim Gemini beat GPT-4
“Back when I'm Googled a Gemini one when I ultra and needed a number that would say it was better than GPT four. And Like, constructed I think never published, like, chain of thought examples, 32 of them in every topic in MLU to run it, to get the score.”
Micah Hill-Smith Jan 9, 2026 ▶ 8:36
Disclosure
Artificial Analysis uses mystery shopper accounts to prevent endpoint manipulation
“We have what we call a mystery shopper policy, and so, and we're totally transparent with all the labs we work with about this, that we will register accounts not on our own domain and run both intelligence evals and performance benchmarks without them being u…”
Micah Hill-Smith Jan 9, 2026 ▶ 13:43
Insight
Hill-Smith: Widely tracked AI benchmarks improve without reflecting general intelligence gains
“Once an eval becomes the thing that everyone's looking at, schools can get better on it without there being a reflection of overall generalized intelligence of these models getting better. That has been true for the last couple of years. It'll be true for the …”
Micah Hill-Smith Jan 9, 2026 ▶ 15:22
Disclosure
Artificial Analysis Intelligence Index synthesizes 10 evaluation datasets
“The artificial analysis intelligence index, which is our synthesis metric that we pulled together currently from 10 different eval data sets to give what? We're pretty confident is the best single number to look at for how smart the models are.”
Micah Hill-Smith Jan 9, 2026 ▶ 19:13
Assertion Not checkable as stated
Hill-Smith: Early Benchmarks Like HumanEval Are Saturated and Trivial
“Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Micah Hill-Smith Jan 9, 2026 ▶ 20:51
Assertion Supported
Hill-Smith: OpenAI Was Untouchable for Well Over a Year
“If we go back even a little bit before then, we're in the era where, when you look at this chart, like, OpenAI was untouchable for well over a year.”
Micah Hill-Smith Jan 9, 2026 ▶ 23:52
Assertion Supported
Anthropic Claude models have lowest hallucination rates on Omniscience benchmark
“Like, one of the things that we saw in the hallucination rate is that Anthropoc's Claude models at the very left-hand side here with the lowest hallucination rates out of the models that we've evaluated Amnesians on.”
Micah Hill-Smith Jan 9, 2026 ▶ 30:09
Assertion Supported
Cameron: General model intelligence does not correlate with hallucination rates
“One interesting aspect is that we've found that there's not really a, not a strong correlation between intelligence and hallucination rate. That's to say that the smarter the models are in a generalist sense isn't correlated with their ability to, when they do…”
George Cameron Jan 9, 2026 ▶ 31:28
Assertion Not publicly verifiable
Hill-Smith: Omniscience Factual Accuracy Tracks Model Parameter Count Most Closely
“If you go back up to accuracy on Omniscience, an interesting thing about this accuracy metric is that it tracks more closely than anything else that we measure, the total parameter count of models.”
Micah Hill-Smith Jan 9, 2026 ▶ 36:24
Opinion
Hill-Smith estimates Gemini 3 Pro parameter count hits 5 to 10 trillion
“You might reasonably form a view that there's a pretty good chance that Gemini three pro is bigger than that, that it could be in the five to 10 trillion parameter range. To be clear, I have absolutely no idea, but just based on this chart, like that's where y…”
Micah Hill-Smith Jan 9, 2026 ▶ 37:42
Insight
Cameron: Inference economics incentivize larger, sparser AI models over dense architectures
“It's, I think, less about total parameters in many cases when thinking about inference costs and more around number of active parameters, and so there's a bit of an incentive towards larger, sparser models.”
George Cameron Jan 9, 2026 ▶ 38:16
Prediction Not checkable as stated
Hill-Smith: Frontier model total parameter sizes have significant room to scale up
“Chances are the last couple of years haven't seen a dramatic scaling up in the total size of these models. And so there's a lot of room to go up probably in total size of the models, especially with the upcoming hardware generations.”
Micah Hill-Smith Jan 9, 2026 ▶ 39:05
Assertion Not checkable as stated
Gemini 3 Pro performs poorly on GDPval-AA benchmark evaluator tasks
“One data point there is that even as the, as an evaluator, Gemini three pro interestingly doesn't do actually that well in GDP val AA.”
George Cameron Jan 9, 2026 ▶ 41:41
Assertion Not publicly verifiable
Models perform better in custom agent harnesses than native web chatbots
“And what's really interesting is that if you compare, for instance, Claude, 4.5 Opus using the Claude web chatbot, it performs worse than the model in our Agentic harness. And so in every case, the model performs better in our agentic harness than its web chat…”
George Cameron Jan 9, 2026 ▶ 45:40
Opinion
Hill-Smith: Multi-tool MCP agent workflows barely work right now
“I would say that this stuff like barely works in fairness right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 48:18
Disclosure
Artificial Analysis open-sources minimalist agent harness Stirrup on GitHub
“We released that on, on GitHub yesterday. It's called Stirrup, so if people want to check it out, and it's a great you know, base for, you know, generalist building a generalist agent.”
George Cameron Jan 9, 2026 ▶ 50:12
Insight
Cameron: Models perform better with minimal tools than rigid frameworks
“I think where we're getting to is that these models have gotten smart enough, they've gotten better, better tools that they can perform better when just given a minimalist set of tools and let them run, let the model Control the agentic workflow rather than us…”
George Cameron Jan 9, 2026 ▶ 51:27
Assertion Supported
AI2's OLMo 3 32B leads Artificial Analysis's 18-point Openness Index
“It's out of 18 currently. And so we've got an openness index page, but essentially these are points. You get points for being more open across these different categories and the maximum you can achieve is 18. So AI two with their extremely open OMO three, 32 B…”
George Cameron Jan 9, 2026 ▶ 53:35
Assertion Supported
GPT-4-level intelligence is now over 100 times cheaper than at launch
“The, like, one fact on that is that you can get intelligence at the level of GPT-IV for over a hundred times cheaper than GPT-IV was at launch right now.”
Micah Hill-Smith Jan 9, 2026 ▶ 58:55
Assertion Supported
Cameron: Model performance correlates with total parameters, not active parameters
“We, in our benchmark, see a lot of performance correlated more with total parameters than active, and not that correlated with how sparse like the models are. Our accuracy benchmark is part of a omniscience. It's very correlated with total. It's not correlated…”
George Cameron Jan 9, 2026 ▶ 1:05:08
Assertion Supported
Reasoning models consume 10x more tokens on average than non-reasoning models
“So, earlier this year, and probably when you and George last spoke for the AI engineers world's fair, we had this great slide that was super easy, where we would show that the average reasoning model is using 10 times the number of tokens per query in our inte…”
Micah Hill-Smith Jan 9, 2026 ▶ 1:06:27
Assertion Contradicted
Frontier models are cheaper for agentic tasks because they require fewer turns
“Interestingly, in Tau Tau Two Bench Telecom, it's cheaper to run, you know, on a per token basis, more expensive models, like a GBD five, compared to some smaller open source models, because the some of the GBD five, for instance got to the answer faster. And …”
George Cameron Jan 9, 2026 ▶ 1:09:14
Prediction Not checkable as stated
Cameron: Turn count will become a major AI benchmarking metric
“I think number of turns is, is going to be a metric that we're going to be talking about a lot more. And I think it'll be something that people want to really start to think about a lot more.”
George Cameron Jan 9, 2026 ▶ 1:09:46
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.