The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 10 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 1 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional software engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have SweeBench, that's cool, no actual Professional work looks like Sweebench, like human eval, same thing.”
Anshul Ramachandran Jul 28, 2025 ▶ 1:50:27 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Opinion
Ramachandran: SWE-bench and HumanEval do not reflect real professional engineering
“Most evals and benchmarks that exist out there for software development is kind of bogus. There's not really a better way of putting it. Like, okay, you have sweet bench. That's cool. No actual Professional work looks like Sweebench. Like, human eval, same thi…”
Anshul Ramachandran Dec 13, 2024 ▶ 13:17 Windsurf: The Enterprise AI IDE
Assertion Not checkable as stated
Royzen: GPT-4 was trained on HumanEval, proving data contamination
“GPT-IV itself has been trained on human eval, and we know this because GPT-IV is able to predict the exact doc string in many of the problems. I've seen it predict, like, the specific example values in the doc string, which is extremely improbable for it to ju…”
Michael Royzen Nov 3, 2023 ▶ 41:31 Beating GPT-4 with Open Source Models - with Michael Royzen of Phind
Insight
Mohan: HumanEval benchmark scores are inflated due to GitHub training contamination
“One of the issues that ends up coming up with things like human eval is contamination, because a lot of these things that train models end up training on all of GitHub. GitHub itself has human eval. So they end up Training on that, and then the numbers are arb…”
Varun Mohan Jul 28, 2025 ▶ 40:03 🕰️ The Oral History of Windsurf (ft. Varun Mohan, Scott Wu, Jeff Wang, Kevin Hou, Anshul R)
Insight
Training on pull request diffs yields code generators, not software engineers
“What the, very crudely, what the pre-trained models are reading is they're reading those final diffs and they're Emulating that and then being able to output it. Right. But of course it's a super lossy thing, a PR. You have no idea why or how, for the most par…”
Alistair Pullen Aug 22, 2024 ▶ 17:14 Is finetuning GPT4o worth it?
Assertion Not checkable as stated
Hill-Smith: Early Benchmarks Like HumanEval Are Saturated and Trivial
“Well, V one would be completely saturated right now by almost every model coming out because doing things like writing the Python functions and human evil is now pretty trivial.”
Micah Hill-Smith Jan 9, 2026 ▶ 20:51 Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith
Assertion Open · timeframe Sep 2028
Bachman: PowerCoder-3B reaches 35% HumanEval accuracy versus StarCoder's 30%
“In the end, this converges to, I believe, about 35% accuracy on human eval, whereas the star coder baseline was about 30%.”
Diego Bachman Sep 23, 2025 ▶ 17:33 ⚡️ Beyond Transformers with Power Retention
Opinion
Schluntz: Traditional coding evals remain useful alongside SWE-bench
“I think there's definitely a space for these more traditional coding evals that are sort of easy to implement, quick to run and do get you some signal. And maybe hopefully there's just sort of harder versions of human eval that get created.”
Erik Schluntz Nov 28, 2024 ▶ 9:01 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Assertion Supported
Jesse Hu: Frontier models have effectively capped out HumanEval performance
“I think through multiple techniques and through the most recent models, you can actually basically cap out on the performance on human eval.”
Jesse Hu Oct 19, 2024 ▶ 1:21 [Paper Club] SWE-Bench [OpenAI Verified/Multimodal] + MLE-Bench with Jesse Hu
Opinion
Yi Tay: GSM8K and HumanEval are saturated, contaminated, and uninformative
“I mean, like, you know, the things like GSMK human eval, the coding human eval, they're all, like Contaminated. Like, not, not, I wouldn't say, they're all, like, saturated, contaminated, you know, like, you know, GSMK, whether you're a 92, 91, like, no one ca…”
Yi Tay Jul 5, 2024 ▶ 1:23:30 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.