HumanEval
also referred to as: human eval
10 statements across 9 episodes · 2 bullish · 7 bearish · 9 people on the record · first statement Nov 3, 2023 by Michael Royzen · said 39 times in 15 episodes since 2023 · across every show →
Mentions by year
brought up most by Varun Mohan (5), Anshul Ramachandran (5), Alessio Fanelli (5), Shawn Wang (3), Shunyu Yao (2), Michael Royzen (2), Jesse Hu (2), Diego Bachman (2)
2026 2 mentions in 1 episode
2025 11 mentions in 3 episodes 4 per episode
2024 23 mentions in 9 episodes 3 per episode
- [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
- The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
- [LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models
- Windsurf: The Enterprise AI IDE
- Language Agents: From Reasoning to Acting — with Shunyu Yao of OpenAI, Harrison Chase of LangGraph
- Is finetuning GPT4o worth it?
- Breaking down the OG GPT Paper by Alec Radford
- Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
- 1 more episode that year, every mention in 2024 →