The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Wei-Lin Chiang no published score: only 2 usable exchanges on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
2exchanges match
2on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 4 4.85

Q And what was the origin of it? How did you come up with the idea? Uh, how did you get people to buy in? And then maybe what were one or two of the pivotal moments early on that kind of made it the standard for, for these things?

A Yeah. Yeah. Chatbot Arena project was started last year in April, May, around that. Before that, we were basically experimenting in the lab how to fine-tune a chatbot open source based on the Lama-one model that had released. At that time, Lama-one was like a base model, and people didn't really know how to fine-tune it, so we were doing some explorations. We were inspired by Stanford's Alpaca project. So, we basically, yeah, grow a data set from the internet, which is called shared gpt data set, which is, like, a dialog data set between user and chat gpt conversation. And it turns out to be, like, pretty high-quality data, dialog data. So, we fine-tune on it, and then we try to, and release the model called wikunia. And people were very excited about it because it kind of like demonstrate open way model can reach this conversation capability similar to ChatGPT. And then we basically released the model ways and also do the demo website. The model. That, people were very excited about it, but during the development, the biggest challenge to us at the time was, like, how do we even evaluate it? How do we even argue this model we trend is better than others? And, like, what's the gap between this open source model and other proprietary offering? At that time, it was, like, GPT-FOR was just announced. It's, like, cloud one, right? What's the difference between them? And then after …

AI assessment note: “we quickly realized that people need a tool to compare between different models.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q benchmarks are really not what we need to evaluate whether or not a model is good. Why did you not make a benchmark? Maybe at the time, you know, it was just like, Hey, let's just put together a whole bunch of data again, run a, Make a score that seems much easier than coming out with a whole website where, like, users need to vote. Any thoughts behind that?

A I think it's more like fundamentally we don't know how to automate this kind of benchmarks when it's more like, you know, conversational, multi-turn, and more open-ended tasks that may not come with a ground truth in. So let's say if you ask a model to help you write an email for you for whatever purpose, there's no ground truth in how you score them. Or write a story, or a creative story, or many other things. Like, how we use strategy these days, it's oftentimes more like open-ended. You know, we need human in a loop to give us feedback. Which one is better? And I think nuance here is like, sometimes it's also hard for human to give the absolute rating. So that's why we have this kind of pairwise comparison, easier for people to choose which one is better. So, from that we, you know, use these pairwise comparisons and those to, to calculate the leadable.

AI assessment note: “fundamentally we don't know how to automate this kind of benchmarks”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.