HumanEval, every mention

19 scenes · ← back to HumanEval

tap a year for its mentions
0013525102023202420252026episodesmentions
05102023202420252026episodes it came up in
00254102023202420252026episodesmentions per episode

every year anyone Varun Mohan 5Anshul Ramachandran 5Alessio Fanelli 5Shawn Wang 3Shunyu Yao 2Michael Royzen 2Jesse Hu 2Diego Bachman 2Alistair Pullen 2Yi Tay 1

Verbatim, from the transcripts: the passages where HumanEval comes up

loading…

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith Jan 9, 2026 · 2 mentions

  • ▶ 20:51 Micah Hill-Smith Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.
  • ▶ 33:23 Shawn Wang Is it sort of like a human eval type or something different?

⚡️ Beyond Transformers with Power Retention Sep 23, 2025 · 2 mentions

  • ▶ 17:15 Diego Bachman This is the, um, human evaluation, the human eval, negative likelihood, and basically corresponds to the performance. 2 times in the scene

🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R) Jul 28, 2025 · 8 mentions

Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex Mar 19, 2025 · 1 mention

  • ▶ 16:27 Shawn Wang Um, do you see the, for convex evals, which is kind of sounds like just human eval for convex.

Windsurf: The Enterprise AI IDE Dec 13, 2024 · 2 mentions

The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic Nov 28, 2024 · 4 mentions

[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu Oct 19, 2024 · 6 mentions

  • ▶ 0:43 Jesse Hu So there's been these benchmarks that are out there before, um, such as human eval, 2 times in the scene
  • ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 4 times in the scene

Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph Sep 27, 2024 · 2 mentions

  • ▶ 36:50 Shunyu Yao And the people that were working on coding are, you know, trying to solve human evil. 2 times in the scene

Is finetuning GPT4o worth it? Aug 22, 2024 · 2 mentions

  • ▶ 17:58 Alistair Pullen You have something that's good at human eval, but, but not very good at SweetBenge, essentially.
  • ▶ 31:53 Alistair Pullen And like, even the things that get really good scores on human evil agents as well, cause they have these loops, right?

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models Jul 29, 2024 · 3 mentions

  • ▶ 27:03 unnamed speaker The coding evaluation, um, it seems like they, so there's, I can't share my screen, um, but basically human eval, um, human eval is a kind of, 1:02, let me try and share it, um, can you see? 3 times in the scene

Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI Jul 23, 2024 · 1 mention

  • ▶ 38:21 Alessio Fanelli So the April, 15 checkpoint, MMLU on Instruct is like, 86, GPUA, 48, Human Eval, 84, GSMEK, 94, MAT, 57.8.

The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka Jul 5, 2024 · 1 mention

  • ▶ 1:23:30 Yi Tay Uh, so, I think, like, uh, yeah, I think analysis is probably the most legit one, like, out of all the, I mean, like, you know, the things like GSMK human eval, the coding human eval, they're all, like Contaminated.

Breaking down the OG GPT Paper by Alec Radford Apr 23, 2024 · 2 mentions

  • ▶ 46:53 unnamed speaker It's like saying, uh, LAMA-II has a human eval of 70 and GPT-IV has a human eval of maybe 90. 2 times in the scene

Beating GPT-4 with Open Source Models - with Michael Royzen of Phind Nov 3, 2023 · 2 mentions

  • ▶ 41:26 Michael Royzen Whole experience releasing those models is that human eval doesn't really matter. 2 times in the scene

Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue Oct 21, 2023 · 1 mention

  • ▶ 43:22 Shawn Wang Beat GPT-IV and Human Eval by role-playing a software agent, development agency, instead of having a sort of single shot, a single role, you have multiple roles, and having all of them criticize each other as agents communicating with…
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.