John Yang

PhD Student, Stanford University · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

academicscientistengineer@jyangballin ↗john-b-yang.github.io ↗

John Yang is the lead co-creator of SWE-bench, an industry benchmark for evaluating autonomous AI software engineering agents. He is a computer science PhD student at Stanford University whose research focuses on language model evaluation and interaction environments.

9statements → 3claims → 2claims resolved → 3.67/5average certainty → 1.67/5average debate potential → 1said about them ↓

2 supported 0 partly supported 0 contradicted 1 not checkable as stated how the 3 claims stand · each chip opens the sources

3 assertions · 1 opinion · 3 insights · 2 disclosures · every statement was checked. The predictions and assertions are the 3 claims: statements the public record can support or contradict. 2 are resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how John argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Yang: AI models falsely claim completion on impossible coding tasks
“I think they're all, the models are all kind of attempting and saying, like, oh, I did it, you know, so maybe not great.”
John Yang Dec 31, 2025 ▶ 10:54 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

Expressed certainty vs assessment result

none yet certainty 1
none yet certainty 2
100% certainty 3
100% certainty 4
none yet certainty 5

weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet

Everything John Yang said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Opinion
Yang: Unit tests and independent task instances limit SWE-bench's effectiveness
“I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other.”
John Yang Dec 31, 2025 ▶ 3:27 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Insight
Yang: Models should pass completion benchmarks before expensive multi-turn agentic evals
“Like, you know, you can do well on those first, and then sort of graduate to the multi-turn expensive stuff.”
John Yang Dec 31, 2025 ▶ 7:14 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Assertion Supported
Yang: AI models falsely claim completion on impossible coding tasks
“I think they're all, the models are all kind of attempting and saying, like, oh, I did it, you know, so maybe not great.”
John Yang Dec 31, 2025 ▶ 10:54 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Insight
Yang: TerminalBench enables more creative eval design than SWE-bench
“As we bench, you're confined in some sense to the domain of issues and PRs that already exist which I think has its benefits of being close to reality and natural, but I think with Terminal Bench, there's a lot of creativity that you can infuse into that”
John Yang Dec 31, 2025 ▶ 11:14 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Insight
Yang: Evaluating Human-AI Code Interaction Requires Compelling Products or Simulators
“From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build, like, a really compelling product, like Elmarina, that people have people use consistently, which is, I mean, really tricky in and of itself. Or yo…”
John Yang Dec 31, 2025 ▶ 14:48 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Assertion Not checkable as stated
Yang: SWE-bench adoption took off only after Cognition's Devin launch
“You know, we put it out October, 22, 23, and then people didn't really touch it too much. And then, of course, like, Cognition came on the scene, and Devon was an amazing release, and I think after that, it kind of kicked off the arms.”
John Yang Dec 31, 2025 ▶ 0:37 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Disclosure
Yang: Walden Yan emailed benchmark results two weeks before Devin launched
“I got an email about, like, two weeks ago. I think it was from, I think it was from Walden. It was like, hey, you know, we have a good number on it.”
John Yang Dec 31, 2025 ▶ 0:51 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Assertion Supported
Yang: SWE-bench Multilingual spans nine languages across roughly 40 repositories
“Yeah, multilingual, it's like nine languages across, like, 40 repos, but yeah, you got them, like, JavaScript, Rust, Java, C, you know, Ruby.”
John Yang Dec 31, 2025 ▶ 1:57 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Disclosure
Yang: Code Clash is building economically valuable competitive coding arenas
“The current ongoing effort is, you know, to build economically valuable arenas.”
John Yang Dec 31, 2025 ▶ 5:27 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

The other half of the tape: John Yang's own voice is left out of every number here. Other people bring the name up 1 time in 1 episode on Latent Space. every mention, with the transcript →

Who brings them up most Shunyu Yao 1

Every mention by year

tap a year for its mentions
0011112024episodesmentions
0112024episodes it came up in
000.50.5112024episodesmentions per episode

Appearances (1)

EpisodeDateSpeaking time
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Ya Dec 31, 2025 11m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.