Ludwig Schmidt

1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

He is a researcher involved in work surrounding AI evaluations and benchmarks. He participated in a fireside chat focused on Terminal-Bench 2.0 and the Harbor evaluation harness.

1statements → 0claims → 0claims resolved → 4/5average certainty → 2/5average debate potential → 13said about them ↓

1 insight · every statement was checked. none of them is a claim the record can settle: opinions, insights and disclosures never carry an assessment.

Everything Ludwig Schmidt said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Insight
Schmidt: Co-authorship functions like startup equity for crowdsourced research
“If you make it a paper and you give everyone a share in the project, this is a little bit like in the startup ecosystem, right? Everyone gets a little bit of equity. Everyone gets a little bit of co-authorship.”
Ludwig Schmidt Nov 8, 2025 ▶ 28:47 Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w/ founders

The other half of the tape: Ludwig Schmidt's own voice is left out of every number here. Other people bring the name up 12 times in 2 episodes on Latent Space. 1 statement on the record names them. every mention, with the transcript →

Who brings them up most Mike Merrill 8Alex Shaw 3Ari Morcos 1

Statements about Ludwig Schmidt, by other people (1)

Assertion Contradicted
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Ari Morcos Aug 29, 2025 ▶ 18:06 Better Data is All You Need — Ari Morcos, Datology

Every mention by year

tap a year for its mentions
00811522025episodesmentions
0122025episodes it came up in
0031622025episodesmentions per episode
2025 12 mentions in 2 episodes 6 per episode

Appearances (1)

EpisodeDateSpeaking time
Terminal-Bench 2.0: the most impt coding agent benchmark of 2025 gets a v2! Launch + Q&A w Nov 8, 2025 1m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.