The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Edwin Chen no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q These, you know, course, uh, as you said, 5:02 reaction, academic benchmarks, or even non-academic industrial benchmarks are, uh, easily hacked or like not the right gauge of performance against any given task. They are very popular. What is the alternative, um, for somebody who's trying to, like, choose the right model or understand model capability?

A So the alternative that I think all the Frontier Labs view as a gold standard is basically human evaluation. So again, proper human evaluation where you're actually taking the time to look at the response. You're going to fact check it. You're going to see whether or not it followed all the instructions. You have good pace so you know whether or not the model has good writing quality. Like this concept of, like, doing all that and spending all the time to do that as opposed to just vibing for five seconds. I think actually is really, really important because if you don't do this, you basically, you're basically just training your models on the analog of clickbait. Um, so I, I think it actually is really, really important for model progress.

AI assessment note: “the alternative that I think all the Frontier Labs view as a gold standard is basically human evaluation”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Okay. Yeah. I mean, so Sarah brought up earlier, um, how everybody kind of wants high quality data. What does that mean? How do you think about that? How do you generate it? Can you tell us a bit more about your thoughts on that?

A So let's, let's say you wanted to train a model to write an eight line poem about the moon. And so the way most companies think about it is, well, let's just hire a bunch of people from Craigslist or through some recruiting agency and let's ask them to write poems. And then the way they think about quality is, well, is this a poem? Is it eight lines? Does it contain the word moon? If so, like, okay, yeah, I hit these three checkboxes. So yeah, sure. This is a great poem because it follows all these instructions. But if you think about it, like the reality is you'd get these terrible poems, like sure it's eight lines and it has the word moon. But they feel like they're written by kids from high school. And so other companies be like, okay, sure. These people on Craigslist don't have any poetry experience. So I'm going to do instead is hire a bunch of people with PhDs in English literature. But this is also terrible. Like a lot of PhDs, they are actually not good writers or poets. Like if you think, like think of Hemingway or Emily Dickinson, they definitely didn't have a PhD. I don't think they even completed college. And like, one of the things I will say is like, yeah, I, I, I went to MIT. I think Eli, you went, you went there too. And a lot of people I knew from MIT who graduated with a CS degree, they're terrible coders. And so we think about quality completely differently. …

AI assessment note: “we think about quality completely differently. Like what we want isn't poetry that checks”

Answered raw tape D 4 · C 5 · P 4 · Cm 4 4.30

Q Is the measurement through human evaluation? Is it Through a model based evaluation. I'm a little bit curious, like how you create that feedback loop since to some extent, it's a little bit of this question of how do you have enough evaluators to evaluate the output relative to the people generating the output or do you use models or how, how do you approach it?

A Like, I think one analogy that we often make is think about something like Google search or think about something like YouTube. Like you have, you know, millions of search results. You have millions of web pages. You have millions of videos. How do you evaluate the qualities of these videos? Like, is this a high quality, like, is this a high quality webpage? Is it informative? Or is it really spammy? Like in the way you do this is like you just need, I mean, you gather so many signals, you gather like page dependent signals, you gather like user dependent signals, you gather activity-based signals, and all of these feed into, you know, a giant ML algorithm at the end of the day. And so in the same way, we, we gather all these signals about our annotators, about the work that they're performing, about like their activity on the site, and we just feed it into, um, a lot of these different, uh, like we basically have an ML team internally that builds a lot of these algorithms to measure all of this.

AI assessment note: “we basically have an ML team internally that builds a lot of these algorithms”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q And you guys are also known for having bootstrap the company versus raising a lot of external venture money or things like that. Do you want to talk about that choice in terms of going profitable early and then scaling off of that?

A In terms of why we didn't raise. So I think, I mean, a big part of it was obviously just that we didn't need the money. I think we were very, very lucky to be profitable from, from the start. So we didn't need the money. It always felt weird to give up control. And like, one of the things I've always hated about Silicon Valley is that you always see so many people raising for the sake of raising. Like, I think one of the things that I often see is that a lot of founders that I know, they don't, they don't have some big dream of building a product that solves some idea that they really believe in. Like, if you talk to a bunch of YC founders or whoever it is, like, what, what is their goal? It really is to tell all their friends that they raised ten million dollars and show their parents they got a headline on TechCrunch. Like, that is their goal. Like, I, I think of like my friends at Google, they, they often tell me, oh yeah, I, you know, I've been at Google or Facebook for 10 years and I want to start a company. I'm like, okay, so what problem do you want to solve? They don't know. They're like, yeah, I just want to start something new. I'm bored. And it's weird because they can like pay their own salaries for a couple of months. Again, they've been on Google and Facebook for 10 years. They're not just like fresh out of school. They, they can pay their own salaries, but the fi…

AI assessment note: “a big part of it was obviously just that we didn't need the money.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q If you were to make a five or 10 year bet on like what scales most in terms of demand from people training AI models and types of data, is it RL environments or is it Traces on types of, like, expert reasoning, or what other areas do you think there's going to be a really large demand for?

A I mean, I think it will be all of the above. Like, I don't think our environments alone will suffice just because, I mean, it depends on how you think about our environments, but oftentimes these are very, very rich trajectories are very, very long. And so it's almost like inconceivable that a single reward, um, I mean, I think even today, we often think about things in terms of multiple rewards, not just a single reward, but a thing like a single reward and we just may not be like rich enough to capture, um, all the, all the work that goes into like the model solving some like very, very complicated goal. Um, so I think it'll probably be a combination of, of all those.

AI assessment note: “I mean, I think it will be all of the above.”

Partly raw tape D 3 · C 4 · P 3 · Cm 3 3.30

Q I guess maybe another, um, sort of broader question is, do you think there's three competitive frontier models, 10 competitive frontier models, uh, a couple years from now? And is any of those open source?

A Yeah. So I actually see more and more frontier models, uh, opening up over time because I actually don't think that the models will be commodities. Like I think one of the things that we've, I mean, I think one of the things that has actually been surprising the past couple of years is that you actually see all of their models have their own focuses that give them unique strengths. Like for example, I think it's obviously been really, really amazing at coding and enterprise and open AI has this big consumer focus because of chat activity. Like I actually really love it. It's model's personality. And then Croc, you know, just has a different set of things that's willing to say and to build. And so it's almost like every company has, it's almost like a different set of principles that they care about. Like they're like, some will just never do one thing. Others are totally willing to do it. Others had just had different, like models will just have so many different facets to their personality. So many different facets to the type of skills that they will be good at. And sure, like eventually AGI will maybe encompass this, this all, but in the meantime, you just kind of need to focus. Like there's only so many focuses that you can have as a company. And so I think that just will be to, uh, like different strengths for all of the model providers. So, I mean, I think today, you know…

AI assessment note: “I actually see more and more frontier models, uh, opening up over time”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.