HumanEval, every mention
20 scenes across 2 shows · ← back to HumanEval
tap a year for its mentions
Latent Space 39
BG2 Pod 1
every year every show
Latent Space 39
BG2 Pod 1
Verbatim, from the transcripts: passages where HumanEval comes up on Latent Space, BG2 Pod
Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
- ▶ 20:51 Micah Hill-Smith Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.
- ▶ 33:23 Shawn Wang Is it sort of like a human eval type or something different?
⚡️ Beyond Transformers with Power Retention
- ▶ 17:15 Diego Bachman This is the, um, human evaluation, the human eval, negative likelihood, and basically corresponds to the performance. 2 times in the scene
🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
- ▶ 39:37 Varun Mohan And there are things online, like human eval, right? 5 times in the scene
- ▶ 1:50:38 Anshul Ramachandran Professional work looks like Sweebench, like human eval, same thing. 3 times in the scene
Fullstack-Bench: The Eval for Coding Agents — with Sujay Jayakar, Chief Scientist, Convex
- ▶ 16:27 Shawn Wang Um, do you see the, for convex evals, which is kind of sounds like just human eval for convex.
Windsurf: The Enterprise AI IDE
- ▶ 13:30 Anshul Ramachandran Like, human eval, same thing. 2 times in the scene
The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- ▶ 7:52 Alessio Fanelli Why do we still use human eval? 4 times in the scene
[Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- ▶ 0:43 Jesse Hu So there's been these benchmarks that are out there before, um, such as human eval, 2 times in the scene
- ▶ 8:40 unnamed speaker Yeah, uh, they see for the, uh, SWE Bench Verified, like, like, um, as you also, SWE Bench is very different, um, compared to the human evolve. 4 times in the scene
Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- ▶ 36:50 Shunyu Yao And the people that were working on coding are, you know, trying to solve human evil. 2 times in the scene
Is finetuning GPT4o worth it?
- ▶ 17:58 Alistair Pullen You have something that's good at human eval, but, but not very good at SweetBenge, essentially.
- ▶ 31:53 Alistair Pullen And like, even the things that get really good scores on human evil agents as well, cause they have these loops, right?
[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- ▶ 27:03 unnamed speaker The coding evaluation, um, it seems like they, so there's, I can't share my screen, um, but basically human eval, um, human eval is a kind of, 1:02, let me try and share it, um, can you see? 3 times in the scene
Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- ▶ 38:21 Alessio Fanelli So the April, 15 checkpoint, MMLU on Instruct is like, 86, GPUA, 48, Human Eval, 84, GSMEK, 94, MAT, 57.8.
The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Ep9. GPT-4o, Astra, Multi Modal, China Tariffs, Tech Earnings | BG2 with Bill Gurley & Brad Gerstner · Bg2 Pod
- ▶ 10:58 Brad Gerstner And we plotted chat GPT for Omni on this chart, and you can see how it barely improved in terms of human level, the human eval score, but it dramatically improved in terms of pricing, you know, in terms of inference pricing.
Breaking down the OG GPT Paper by Alec Radford
- ▶ 46:53 unnamed speaker It's like saying, uh, LAMA-II has a human eval of 70 and GPT-IV has a human eval of maybe 90. 2 times in the scene
Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
- ▶ 41:26 Michael Royzen Whole experience releasing those models is that human eval doesn't really matter. 2 times in the scene
Why AI Agents Don't Work (yet) - with Kanjun Qiu of Imbue
- ▶ 43:22 Shawn Wang Beat GPT-IV and Human Eval by role-playing a software agent, development agency, instead of having a sort of single shot, a single role, you have multiple roles, and having all of them criticize each other as agents communicating with…