People, every show

John Yang

PhD Student, Stanford University. On 1 show, 1 appearance. The Shows tab opens the full record on each.

academicscientistengineer@jyangballin ↗john-b-yang.github.io ↗

John Yang is the lead co-creator of SWE-bench, an industry benchmark for evaluating autonomous AI software engineering agents. He is a computer science PhD student at Stanford University whose research focuses on language model evaluation and interaction environments.

1shows
1appearances
9statements
2resolved
2supported
0contradicted
100%fully supported
1said about them ↓

Everything John Yang said on any show that made the record, most notable first. Each card names its show and opens the statement there.

Yang: Unit tests and independent task instances limit SWE-bench's effectiveness
“I don't like unit tests as a form of verification, and I also think there's an issue with Sweepbench where all of the task instances are independent of each other.”
John Yang Dec 31, 2025 ▶ 3:27 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Yang: Models should pass completion benchmarks before expensive multi-turn agentic evals
“Like, you know, you can do well on those first, and then sort of graduate to the multi-turn expensive stuff.”
John Yang Dec 31, 2025 ▶ 7:14 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
LATENT SPACE Assertion Supported
Yang: AI models falsely claim completion on impossible coding tasks
“I think they're all, the models are all kind of attempting and saying, like, oh, I did it, you know, so maybe not great.”
John Yang Dec 31, 2025 ▶ 10:54 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Yang: TerminalBench enables more creative eval design than SWE-bench
“As we bench, you're confined in some sense to the domain of issues and PRs that already exist which I think has its benefits of being close to reality and natural, but I think with Terminal Bench, there's a lot of creativity that you can infuse into that”
John Yang Dec 31, 2025 ▶ 11:14 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
Yang: Evaluating Human-AI Code Interaction Requires Compelling Products or Simulators
“From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build, like, a really compelling product, like Elmarina, that people have people use consistently, which is, I mean, really tricky in and of itself. Or yo…”
John Yang Dec 31, 2025 ▶ 14:48 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
LATENT SPACE Assertion Not checkable as stated
Yang: SWE-bench adoption took off only after Cognition's Devin launch
“You know, we put it out October, 22, 23, and then people didn't really touch it too much. And then, of course, like, Cognition came on the scene, and Devon was an amazing release, and I think after that, it kind of kicked off the arms.”
John Yang Dec 31, 2025 ▶ 0:37 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
LATENT SPACE Disclosure
Yang: Walden Yan emailed benchmark results two weeks before Devin launched
“I got an email about, like, two weeks ago. I think it was from, I think it was from Walden. It was like, hey, you know, we have a good number on it.”
John Yang Dec 31, 2025 ▶ 0:51 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
LATENT SPACE Assertion Supported
Yang: SWE-bench Multilingual spans nine languages across roughly 40 repositories
“Yeah, multilingual, it's like nine languages across, like, 40 repos, but yeah, you got them, like, JavaScript, Rust, Java, C, you know, Ruby.”
John Yang Dec 31, 2025 ▶ 1:57 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang
LATENT SPACE Disclosure
Yang: Code Clash is building economically valuable competitive coding arenas
“The current ongoing effort is, you know, to build economically valuable arenas.”
John Yang Dec 31, 2025 ▶ 5:27 [State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

The other half of the tape: John Yang's own voice is left out of every number here. Other people bring the name up 1 time in 1 episode across the shows. every mention, with the transcript →

Who brings them up most Shunyu Yao 1

Every mention by year

tap a year for its mentions
0011112024episodesmentions
0112024episodes it came up in
000.50.5112024episodesmentions per episode

Latent Space 1

One line per show, most statements first. The link opens John's full record on that show: the calibration, argument clarity, speaking style and every statement made there.

ShowRole thereEpsStatementsRecord
LATENT SPACELEDGER PhD Student, Stanford University 1 9 100% 2/2 full record on Latent Space →
Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.