John Yang
PhD Student, Stanford University · 1 appearance on the record.
computed by AI from the episodes · how this works → · full disclaimer →
academicscientistengineer@jyangballin ↗john-b-yang.github.io ↗
John Yang is the lead co-creator of SWE-bench, an industry benchmark for evaluating autonomous AI software engineering agents. He is a computer science PhD student at Stanford University whose research focuses on language model evaluation and interaction environments.
2 supported 0 partly supported 0 contradicted 1 not checkable as stated how the 3 claims stand · each chip opens the sources
3 assertions · 1 opinion · 3 insights · 2 disclosures · every statement was checked. The predictions and assertions are the 3 claims: statements the public record can support or contradict. 2 are resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.
The record, in short
What the tape says about how John argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.
Expressed certainty vs assessment result
weighted support: a fully supported claim counts one, a partly supported claim counts half. Each filled bar is clickable and opens exactly those claims; "none yet" means nothing said at that certainty level has resolved yet
Everything John Yang said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →
The other half of the tape: John Yang's own voice is left out of every number here. Other people bring the name up 1 time in 1 episode on Latent Space. every mention, with the transcript →
Who brings them up most Shunyu Yao 1
Every mention by year
Appearances (1)
| Episode | Date | Speaking time |
|---|---|---|
| [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Ya | Dec 31, 2025 | 11m |