The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Anastasios Angelopoulos no published score: only 2 usable exchanges on raw tape, and a fair score needs 8+ record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
2exchanges match
2on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Was there a moment for you where you were, uh, I'm sure you were debating it yourself. You had other opportunities. What was the deciding factor for you?

A It became clear that the only way to scale what we were building was to build a company out of it. That the world really needed something like Arena. Arena being really a place to sort of measure, understand, and, and advance the frontier AI capabilities in, on real world users, on real world usage. Based on organic feedback and that in order to achieve the scale and, you know, distribution necessary and the quality of course of the platform necessary to do this effectively, we would need to start a company out of it. You know, we considered other options. Are we going to keep doing this as an academic project? Are we going to do it as a nonprofit? Blah, blah, blah. But ultimately under those constructs, we didn't feel like we'd have the resources necessary to accomplish our mission.

AI assessment note: “It became clear that the only way to scale what we were building was to build a company”

Partly raw tape D 3 · C 5 · P 4 · Cm 4 4.00

Q Expert arena. So basically like w what is in the critical path for you, let's say for next year and what, what have you decided you will never do?

A So let me first Talk about things that I'll never do. The platform, integrity comes first to the platform. The, basically the public leaderboard that we show on Ellen Marina, I think of as a charity. It's a loss leader for us. We don't really make money on the public leaderboard. You can't pay to get on the public leaderboard. It's not like a Gartner in that sense. It's not like any of these, like, uh, you know, pay to play systems, never going to be like that. Models are going to be listed on the leaderboard, whether or not the providers pay. And whether or not they're getting a good score. They can't pay to take it off either. And so what that means, that that's very important. And so what that means is that the leaderboard has a certain integrity that will never be compromised, of course.

AI assessment note: “So let me first Talk about things that I'll never do.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.