Assertion Contradicted AI assessment confidence: 95% certainty 3/5 debate potential 2/5

Merrill: 60% to 70% of SWE-bench Verified tasks come from Django

Mike Merrill · Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits · Oct 18, 2025 · at 19:06

Terminal-Bench co-creator Mike Merrill critiques the repository diversity in SWE-bench Verified compared to Terminal-Bench.

0:00 / 0:06exact quote · 6.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“If you go look at sweet bench verified, I think like 60, 70% of the tasks in there are from Django.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Mike Merrill

Insight
Merrill: AI terminal interfaces will outpace GUI computer use development
“The bet here that we made was that GUI based computer use was going to be much slower to come into its own. Than terminal based use. And there's a few reasons for this. I mean, as I mentioned earlier, text is just the modality that works best with these models…”
Mike Merrill Oct 18, 2025 ▶ 6:29 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Prediction Not checkable as stated
Merrill: AI evals will shift to observing real jobs over 2-3 years
“I think like the future of evals does look much more like this is like observing people who are doing their real jobs and then translating those real jobs into a format that allows you to evaluate language models and harnesses on them. And it's probably going …”
Mike Merrill Oct 18, 2025 ▶ 18:41 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Prediction Not checkable as stated
Merrill: Frontier AI labs will center operations around vertical products
“And with the Claude codes and the codec CLIs and the deep researchers, researchers, you starting to see some evidence that the products are going to be a much more central part of how these frontier labs operate.”
Mike Merrill Oct 18, 2025 ▶ 26:19 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Assertion Supported
Merrill: Terminus-2 harness yields highest agent performance on Terminal-Bench 2.0
“Terminus-II, which is our harness for measuring language models, is the highest performing agent for many of these models.”
Mike Merrill Nov 8, 2025 ▶ 6:10 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Opinion
Merrill: Text-based agents are currently the frontier of performance
“We like text-based agents. We think right now they're the frontier of agentic performance. This very well may change.”
Mike Merrill Nov 8, 2025 ▶ 33:58 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders
Prediction Not checkable as stated
Merrill: AI agents could ace Terminal-Bench with $1,000 compute per task
“Like, I think there's probably a way in which you spend a thousand dollars on tokens for each terminal bench task. And like get them all right just by scaling inference time compute.”
Mike Merrill Oct 18, 2025 ▶ 30:33 Terminal-Bench: Pushing Claude Code, OpenAI Codex, Factory Droid, et al to the limits
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.