HumanEval

product on 3 shows · 12 statements across 11 episodes · said 40 times in 16 episodes since 2023

Latent Space 39 BG2 Pod 1 Lenny's Podcast

Mentions by year, every show

tap a year for its mentions
0013525102023202420252026episodesmentions
05102023202420252026episodes it came up in
00254102023202420252026episodesmentions per episode

Latent Space 39BG2 Pod 1

2026 2 mentions in 1 episode
2025 11 mentions in 3 episodes 4 per episode
2024 24 mentions in 10 episodes 2 per episode
2023 3 mentions in 2 episodes 2 per episode

every mention on every show, scene by scene, with the transcript →

12 statements about HumanEval, every show

LATENT SPACE Assertion Not checkable as stated
Hill-Smith: Early Benchmarks Like HumanEval Are Saturated and Trivial
“Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Micah Hill-Smith Jan 9, 2026 ▶ 20:51 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Husain: General LLM benchmarks do not correlate with product-specific evals
“Up until now, a lot of the big labs understandably focused on general benchmarks, like MMLU score, human eval, things like that, which are very important for foundation models. And, you know, those not very related to product specific evals, like the ones we t…”
Hamel Husain Sep 25, 2025 ▶ 1:21:35 Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar
LATENT SPACE Assertion Open · timeframe Sep 2028
Bachman: PowerCoder-3B reaches 35% HumanEval accuracy versus StarCoder's 30%
“In the end, this converges to, I believe, about 35% accuracy on human eval, whereas the star coder baseline was about 30%.”
Diego Bachman Sep 23, 2025 ▶ 17:33 ⚡️ Beyond Transformers with Power Retention
Mohan: HumanEval benchmark scores are inflated due to GitHub training contamination
“One of the issues that ends up coming up with things like human eval is contamination, because a lot of these things that train models end up training on all of GitHub. GitHub itself has human eval. So they end up Training on that, and then the numbers are arb…”
Varun Mohan Jul 28, 2025 ▶ 40:03 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Ramachandran: SWE-bench and HumanEval do not reflect real professional software engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have SweeBench, that's cool, no actual Professional work looks like Sweebench, like human eval, same thing.”
Anshul Ramachandran Jul 28, 2025 ▶ 1:50:27 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Ramachandran: SWE-bench and HumanEval do not reflect real professional engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have sweet bench. That's cool. No actual Professional work looks like Sweebench. Like, human eval, same thi…”
Anshul Ramachandran Dec 13, 2024 ▶ 13:17 Windsurf: The Enterprise AI IDE
Schluntz: Traditional coding evals remain useful alongside SWE-bench
“I think there's definitely a space for these more traditional coding evals that are sort of easy to implement, quick to run and do get you some signal. And maybe hopefully there's just sort of harder versions of human eval that get created.”
Erik Schluntz Nov 28, 2024 ▶ 9:01 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
LATENT SPACE Assertion Supported
Jesse Hu: Frontier models have effectively capped out HumanEval performance
“I think through multiple techniques and through the most recent models, you can actually basically cap out on the performance on human eval.”
Jesse Hu Oct 19, 2024 ▶ 1:21 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Training on pull request diffs yields code generators, not software engineers
“What the, very crudely, what the pre-trained models are reading is they're reading those final diffs and they're Emulating that and then being able to output it. Right. But of course it's a super lossy thing, a PR. You have no idea why or how, for the most par…”
Alistair Pullen Aug 22, 2024 ▶ 17:14 Is finetuning GPT4o worth it?
Yi Tay: GSM8K and HumanEval are saturated, contaminated, and uninformative
“I mean, like, you know, the things like GSMK human eval, the coding human eval, they're all, like Contaminated. Like, not, not, I wouldn't say, they're all, like, saturated, contaminated, you know, like, you know, GSMK, whether you're a 92, 91, like, no one ca…”
Yi Tay Jul 5, 2024 ▶ 1:23:30 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
BG2 Assertion Supported
Gerstner: GPT-4o Dramatically Lowers Inference Pricing Over Benchmark Gains
“And we plotted chat GPT for Omni on this chart, and you can see how it barely improved in terms of human level, the human eval score, but it dramatically improved in terms of pricing, you know, in terms of inference pricing.”
Brad Gerstner May 16, 2024 ▶ 10:58 Ep9. GPT-4o, Astra, Multi Modal, China Tariffs, Tech Earnings | BG2 with Bill Gurley & Brad Gerstner · Bg2 Pod
LATENT SPACE Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.