HumanEval, every mention
19 scenes · ← back to HumanEval
tap a year for its mentions
every year anyone Varun Mohan 5Anshul Ramachandran 5Alessio Fanelli 5Shawn Wang 3Shunyu Yao 2Michael Royzen 2Jesse Hu 2Diego Bachman 2Alistair Pullen 2Yi Tay 1
Verbatim, from the transcripts: the passages where HumanEval comes up
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 20:51 Micah Hill-Smith Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.
- ▶ 33:23 Shawn Wang Is it sort of like a human eval type or something different?
⚡️ Beyond Transformers with Power Retention
- ▶ 17:15 Diego Bachman This is the, um, human evaluation, the human eval, negative likelihood, and basically corresponds to the performance. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 39:37 Varun Mohan And there are things online, like human eval, right? 5 times in the scene
- ▶ 1:50:38 Anshul Ramachandran Professional work looks like Sweebench, like human eval, same thing. 3 times in the scene
Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
- ▶ 16:27 Shawn Wang Um, do you see the, for convex evals, which is kind of sounds like just human eval for convex.
Windsurf: The Enterprise AI IDE
- ▶ 13:30 Anshul Ramachandran Like, human eval, same thing. 2 times in the scene
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 7:52 Alessio Fanelli Why do we still use human eval? 4 times in the scene
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- ▶ 0:43 Jesse Hu So there's been these benchmarks that are out there before, um, such as human eval, 2 times in the scene
- ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 4 times in the scene
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 36:50 Shunyu Yao And the people that were working on coding are, you know, trying to solve human evil. 2 times in the scene
Is finetuning GPT4o worth it?
- ▶ 17:58 Alistair Pullen You have something that's good at human eval, but, but not very good at SweetBenge, essentially.
- ▶ 31:53 Alistair Pullen And like, even the things that get really good scores on human evil agents as well, cause they have these loops, right?
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 27:03 unnamed speaker The coding evaluation, um, it seems like they, so there's, I can't share my screen, um, but basically human eval, um, human eval is a kind of, 1:02, let me try and share it, um, can you see? 3 times in the scene
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 38:21 Alessio Fanelli So the April, 15 checkpoint, MMLU on Instruct is like, 86, GPUA, 48, Human Eval, 84, GSMEK, 94, MAT, 57.8.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Breaking down the OG GPT Paper by Alec Radford
- ▶ 46:53 unnamed speaker It's like saying, uh, LAMA-II has a human eval of 70 and GPT-IV has a human eval of maybe 90. 2 times in the scene
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 41:26 Michael Royzen Whole experience releasing those models is that human eval doesn't really matter. 2 times in the scene
Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
- ▶ 43:22 Shawn Wang Beat GPT-IV and Human Eval by role-playing a software agent, development agency, instead of having a sort of single shot, a single role, you have multiple roles, and having all of them criticize each other as agents communicating with…