The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 6 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Assertion Supported
Gur-Ari: Augment Code Achieved #1 on SWE-Bench Verified
“We just made number one on Sweepbench. So for us Sweepbench has been a useful tool for exploring how can we get the most out of agents. And so we were able to get the best result on Sweepbench verified right now.”
Guy Gur-Ari Apr 2, 2025 ▶ 1:24 The #1 SWE-Bench Verified Agent
Assertion Supported
Cosine's Genie achieved a state-of-the-art 43.8% on SWE-bench Verified
“We got 219 out of 500, which is 43.8%, which is To my knowledge, at least right now, state of the art also”
Alistair Pullen Aug 22, 2024 ▶ 52:38 Is finetuning GPT4o worth it?
Assertion Supported
Watkins: About 90% of SWE-bench Verified tasks take under an hour
“For Sweep Edge Verified I think that's something like, 90% of the problems are things that were estimated to take, like, an expert software engineer like, less than an hour.”
Olivia Watkins Feb 23, 2026 ▶ 10:59 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Contradicted
Merrill: 60% to 70% of SWE-bench Verified tasks come from Django
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”
Mike Merrill Oct 18, 2025 ▶ 19:06 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Supported
Watkins: OpenAI hired nearly 100 engineers to curate 500 SWE-bench tasks
“So folks at OpenAI did a pretty extensive human data campaign, hiring like almost a hundred real-world software engineers to go through the problems and figure out, like, are the tasks well-specified? Are the tests actually fair and kind of created a curated s…”
Olivia Watkins Feb 23, 2026 ▶ 2:56 The End of SWE-Bench Verified — Mia Glaese & Olivia Watkins, OpenAI Frontier Evals
Assertion Supported
Schluntz: SWE-bench Verified was created in partnership with OpenAI
“SweetBench Verified was actually made in partnership with OpenAI, and they hired humans to go review all these tasks and pick out a subset to try to remove any obstacle like this that would make the tasks impossible.”
Erik Schluntz Nov 28, 2024 ▶ 10:03 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.