SimpleQA

product on 2 shows · 2 statements across 2 episodes · said 8 times in 3 episodes since 2025

Latent Space 4 the a16z Podcast 4

Mentions by year, every show

tap a year for its mentions
0042832025episodesmentions
0232025episodes it came up in
001.51.5332025episodesmentions per episode

the a16z Podcast 4Latent Space 4

2025 8 mentions in 3 episodes 3 per episode

every mention on every show, scene by scene, with the transcript →

2 statements about SimpleQA, every show

a16z Assertion Supported
Labenz: GPT-4.5 achieved 65% accuracy on SimpleQA versus o3's 50%
“The O-three class of models got about a 50% on that benchmark, and GPT 4.5 popped up to like 65%. So, in other words, it basically, of the things that were not known to the previous generation of models, it picked up a third of them.”
Nathan Labenz Oct 14, 2025 ▶ 8:58 Is AI Slowing Down? Nathan Labenz Says We're Asking the Wrong Question
LATENT SPACE Assertion Partly supported
Lambert: SimpleQA benchmark scores drop across reasoning models tested without tools
“You look at all the evals from reasoning models, and one of the trends is that, like simple QA numbers all drop. It's like DeepSeq R-one to the new R-one, it goes down. It's like all the new, like, QN-II to QN-III, simple QA goes down, at least when you're eva…”
Nathan Lambert Jul 31, 2025 ▶ 22:59 The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.