Wei-Lin Chiang

Co-Founder & CTO, Arena · 1 appearance on the record.

computed by AI from the episodes · how this works → · full disclaimer →

founderexecutivescientistacademic@infwinston ↗LinkedIn ↗infwinston.github.io ↗

Wei-Lin Chiang is the co-founder and CTO of Arena, formerly known as LMSYS and Chatbot Arena. He previously conducted AI systems research at UC Berkeley's Sky Computing Lab, where he co-developed Vicuna and FastChat.

5statements → 4claims → 0claims resolved → 3.6/5average certainty → 1.4/5average debate potential → 2said about them ↓

4 not checkable as stated how the 4 claims stand · each chip opens the sources

4 assertions · 1 insight · every statement was checked. The predictions and assertions are the 4 claims: statements the public record can support or contradict. 0 are resolved, and 4 name no date, number or outcome precise enough to check. Everything else (opinions, insights, what ifs, disclosures) can never be settled by the record, so it carries no assessment.

The record, in short

What the tape says about how Wei-Lin argues and how the claims held up. Everything they said, and everything said about them, is in the tabs below.

How they sound: not measured why? →

We measure speaking style by listening to the audio itself, and a fair number needs at least 2,000 words from one person on tape we have measured. There is too little of Wei-Lin Chiang on measured tape to publish a rate. This says nothing about how they speak.

Everything Wei-Lin Chiang said on Latent Space that made the record, most notable first. Filter by type, assessment or year in the ledger →

Assertion Not checkable as stated
Chiang: Vicuna Showed Open-Weight Models Could Match ChatGPT Quality
“And people were very excited about it because it kind of like demonstrate open way model can reach this conversation capability similar to ChatGPT.”
Wei-Lin Chiang Nov 1, 2024 ▶ 3:01 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Not checkable as stated
Chiang: There are currently no good benchmarks for evaluating LLM routers
“Right now, currently, there seems to be the, one of the end point when we developed this project was like, there's just no good benchmark for a router.”
Wei-Lin Chiang Nov 1, 2024 ▶ 35:57 In the Arena: How LMSys changed LLM Benchmarking Forever
Insight
Chiang: Online dynamic benchmarks are slower and more expensive than offline benchmarks
“This kind of like online dynamic benchmark is slow, is more expensive than Standing benchmark, offline benchmark, where people still need it, like when they build models, they need static benchmark to track.”
Wei-Lin Chiang Nov 1, 2024 ▶ 7:48 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Not checkable as stated
Chiang: Coding questions drive 20% to 30% of Chatbot Arena usage
“We do see a lot of like developers come to the site asking polling questions. Only 30%.”
Wei-Lin Chiang Nov 1, 2024 ▶ 10:36 In the Arena: How LMSys changed LLM Benchmarking Forever
Assertion Not checkable as stated
Chiang: Chatbot Arena almost died after launch due to low engagement
“At some point, almost died. Because as you can imagine, this leaderboard depends on user, like part of, like community engagement participation. If no one comes to vote, Tomorrow then no deal.”
Wei-Lin Chiang Nov 1, 2024 ▶ 11:10 In the Arena: How LMSys changed LLM Benchmarking Forever

The other half of the tape: Wei-Lin Chiang's own voice is left out of every number here. Other people bring the name up 2 times in 1 episode on Latent Space. every mention, with the transcript →

Who brings them up most Anjney Midha 2

Every mention by year

tap a year for its mentions
0011212026episodesmentions
0112026episodes it came up in
0010.5212026episodesmentions per episode
2026 2 mentions in 1 episode

Appearances (1)

EpisodeDateSpeaking time
In the Arena: How LMSys changed LLM Benchmarking Forever Nov 1, 2024 10m
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.