The Ledger

Every statement that passed quotation and attribution checks. Mix any filter with any other: certainty 1/5, debate potential 5/5, or both at once.

clear all ✕

why aren't all 5 resolved? a statement only gets an assessment when the public record can support or contradict it. opinions and what-ifs never can, and 0 checkable ones are still open, waiting for their date. predictions held up or didn't; assertions are supported or contradicted. on every card: ▮▮▮▮▮ certainty · ▮▮▮▮▮ debate potential. speakers are clickable

Opinion
Feinberg: Gemini is obviously worse at coding despite benchmark wins
“So Gemini does pretty well on Sweebench. Sometimes Gemini publishes models that win on some of those software benchmarks. Raise your hand if you're using Gemini to write code right now instead of, you know, the obvious other name competitors. No one. Like, why…”
Evan Feinberg Jun 30, 2026 ▶ 55:32 🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Opinion
SWE-bench tasks do not reflect actual enterprise software engineering use cases
“That, and also like just in the enterprise, the use cases are pretty different than those represented in something like Sweebench.”
Matan Grinberg May 29, 2025 ▶ 26:27 The AI Coding Factory
Insight
Schluntz: SWE-bench reflects real engineering by requiring repository navigation
“Sweebench, you're starting in the context of an entire repository. And so it adds this entirely new dimension to the problem of finding the relevant files. And, you know, this is a huge part of real engineering”
Erik Schluntz Nov 28, 2024 ▶ 7:25 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Prediction Not checkable as stated
Schluntz: Real-World Coding Agent Workflows Will Be Interactive, Not One-Shot
“So I think that like real tasks are going to be much more interactive with the agent rather than this kind of like one shot system.”
Erik Schluntz Nov 28, 2024 ▶ 32:37 The new Claude 3.5 Sonnet, Computer Use, and Building SOTA Agents — with Erik Schluntz, Anthropic
Assertion Not checkable as stated
Running a 100-problem SWE-bench evaluation takes one to two hours
“So a Sweebench eval for me takes about an hour to two hours to run on like a subset of a hundred problems.”
Shawn Lewis Jan 28, 2025 ▶ 10:56 Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.