Wei-Lin Chiang

Co-Founder & CTO, Arena · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutivescientistacademic@infwinston ↗LinkedIn ↗infwinston.github.io ↗

Wei-Lin Chiang is the co-founder and CTO of Arena, formerly known as LMSYS and Chatbot Arena. He previously conducted AI systems research at UC Berkeley's Sky Computing Lab, where he co-developed Vicuna and FastChat.

3statements → 2claims → 1claims resolved → 3.67/5average certainty → 1.67/5average debate potential →

1 supported 0 partly supported 0 contradicted 1 not checkable as stated how the 2 claims stand · each chip opens the sources

1 prediction · 1 assertion · 1 disclosure · every statement was checked. The prediction and assertion are the 2 claims: statements the public record can support or contradict. 1 is resolved, and 1 names no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Wei-Lin argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

Their most notable supported claim

Assertion Supported
Chiang: Red Team Arena maintains a leaderboard for AI jailbreakers
“So in Red Team Arena, we have a leaderboard, not just for model, but for user, for job breakers, who is the best job breakers that can like identify issues. For all different models.”
Wei-Lin Chiang May 29, 2025 ▶ 1:40:59 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable

How they sound: speaking style how? →

248 words/min while actually speaking · 13.4 um and uh per 1k words

No argument clarity score for Wei-Lin Chiang: no usable question→answer exchanges on raw tape (a fair score needs 8+). We do not score a sample that small. Roundtable and news formats yield far fewer direct exchanges than interviews.

Measured by listening to the audio itself: 3,213 words across 1 episode of raw-level tape, transcribed verbatim with every um and uh kept, each one attributed only where the alignment onto our timed stream is unambiguous. These are measurements of speaking style. We do not rank them: across this corpus, fluency and argument quality are nearly uncorrelated (ρ≈0.2), and smooth talking does not signal clear thinking. How it's measured →

Everything Wei-Lin Chiang said on the a16z Podcast that made the record, most notable first. Filter by type, assessment or year in the ledger →

Prediction Not checkable as stated
Chiang: Personalized leaderboards will improve LMArena data quality
“And in that case, we align the interests of individuals and the platform as a whole, because you don't want to mess up your personal leaderboard. Just like how people these days, when they use social media, They don't like a random post because if they do that…”
Wei-Lin Chiang May 29, 2025 ▶ 1:32:22 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Assertion Supported
Chiang: Red Team Arena maintains a leaderboard for AI jailbreakers
“So in Red Team Arena, we have a leaderboard, not just for model, but for user, for job breakers, who is the best job breakers that can like identify issues. For all different models.”
Wei-Lin Chiang May 29, 2025 ▶ 1:40:59 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Disclosure
Chiang: LMArena open-sources all code, infrastructure, and models
“All the code infrastructures that we process the data is published as open source and also research blog, paper, and then including prompt leaderboard, we publish the paper. Open source, the models, the code, and everything.”
Wei-Lin Chiang May 29, 2025 ▶ 1:35:01 Beyond Leaderboards: LMArena’s Mission to Make AI Reliable

Appearances (1)

EpisodeDateSpeaking time
Beyond Leaderboards: LMArena’s Mission to Make AI Reliable May 29, 2025 18m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.