The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Edwin Chen no published score: only 6 usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.0/5 from 6 raw tape exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
6exchanges match
6on raw tape
0redirected or not addressed
Answered raw tape D 5 · C 5 · P 5 · Cm 5 5.00

Q This connects to, uh, reinforcement learning. That's something that you're, you're big on and something I'm hearing more and more is just becoming a big deal in the world of post-training. Can you just help people understand what is reinforcement learning and reinforcement learning environments and why they're so, they're going to be more and more important in the future?

A Reinforcement learning is essentially training your model to reach a certain reward. And let me explain what an R environment is. An R environment is essentially a simulation of real world. So think of it like building a video game with a fully fleshed out universe. Every character has a real story. Every business has tools and data you can call, and you have all these different entities interacting with each other. So for example, we might build a world where you have a startup with Gmail messages and Slack threads and Jira tickets and GitHub ERs and a whole code base. And then suddenly AWS goes down and Slack goes down. And so, okay, model, what do you do? Like the model needs to figure it, figure it out. So we give the models tasks in these environments. We design interesting challenges for them, and then we run them to see how they perform. And then we teach them, we give them these rewards when they're doing a good job or a bad job. And I think one of the interesting things is that these environments really showcase where models are end to end, are weak at end to end tasks in the real world. You have all these models that seem really smart on isolated benchmarks. Like they're good at single step tool calling. They're good at single step instruction following. But suddenly you dump them into these messy worlds where you have confusing Slack messages and tools they've never …

AI assessment note: “Reinforcement learning is essentially training your model to reach a certain reward.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So what I'm hearing is there's kind of a, just Going much deeper in, uh, understanding what quality is within the verticals that you are selling data around. So you, and is this like a person you hire that is incredibly talented at poetry plus, uh, evals that they, I guess, help write that tell them that this is great. How, what's the mechanics of that?

A The way it works is we essentially gather thousands of signals about everything that you're doing when you're working on a platform. So we are looking at your keyboard strokes. We are looking how fast you answer things. We are using reviews. We are using code standards. We are using, like, we're training models ourselves on the outputs that you create, and then we're seeing whether they improve the model's performance. And so in a very similar way to how Google search, like when Google search is trying to determine what is a good webpage, there's almost two aspects of it. One is you want to remove all of the worst of the worst web pages. So you want to remove all the spam, all the, uh, just like low quality content, all the pages that don't load. And so just like a, it's almost like a content moderation problem. You just want to remove the worst of the worst. But then you also want to discover the best of the best. Okay, like this is the best webpage or, you know, this is the best person for this job. They are not just somebody who writes the equivalent of high school level poetry. Again, like they're not just for biology writing poetry that checks all these boxes, checks all these explicit instructions, but rather, yeah, they're writing poetry that makes you emotional. And so we have all these signals as well that again, like completely differently from moving the worst to the…

AI assessment note: “The way it works is we essentially gather thousands of signals about everything”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Knowing that, with that in mind, how do you kind of get a sense of if we're heading towards AGI? How do you measure progress?

A Yeah, so the way we really care about measuring model progress is by running all these human evaluations. So, for example, what we do is, yeah, we will take Or human annotators, and we'll ask them, okay, go have a conversation with the model. Maybe you're having a small conversation with the model across all of these different topics. So, okay, you are a Nobel Prize winning physicist, so you go have a conversation about pushing different tier of your own research. You are a teacher, and you're trying to create lesson plans for your students, so go talk to the model about these things. Or you are a, yeah, you're, you're a coder, and you're working at One of these big tech companies, and you have these problems every day, so go talk to the model and see how much it helps you. And because, or surgers or annotators, they are experts at the top of their fields, and they are not just skimming the responses, they're actually working through the responses deeply themselves. They are, yeah, they're going to evaluate the code that it writes. They're going to double check the physics equations that it writes. They're going to evaluate the models in a very, very deep way. So they're going to pay attention to accuracy and instruction following and all these things that casual users don't when you suddenly get a pop-up on your ChatGPT response asking you to compare these two different respon…

AI assessment note: “measuring model progress is by running all these human evaluations.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Awesome. I want to ask a couple broad AI kind of market questions. What else do you think is coming in the next couple of years that people are maybe not thinking enough about or not expecting in terms of where AI is heading? What's going to matter?

A I think one of the things that's going to happen in the next few years is that the models are actually going to become increasingly differentiated because of the personalities and behaviors That the different labs have and the kind of objective functions that they are optimizing their models for. Like, I think it's one thing I didn't appreciate a year or so ago. Like a year or so ago, I thought that all of the AI models would essentially become very, very commoditized. They would all behave like each other. And sure, one of them might be slightly more intelligent in one way today, but sure, the other ones would catch up in the next few months. But I think over the past year, I've realized that the values that the companies have will shape the, that the model. So let me give an example. So I was asking Claude to help me drop an email the other day, and it went through 30 different versions. And after 30 minutes, yeah, I think it really crafted me the perfect email, and I sent it. But then I realized that I spent 30 minutes doing something that didn't matter at all. Like sure, now I got the perfect email, but I spent 30 minutes doing something I wouldn't have worried at all before, and this email probably didn't even move the needle on anything anyways. So I think there's a deep question here, which is, if you could choose the perfect model behavior, which model would you want? D…

AI assessment note: “models are actually going to become increasingly differentiated because of the personalities and behaviors”

Partly raw tape D 3 · C 5 · P 3 · Cm 4 3.75

Q 70 people. You're completely bootstrapped, haven't raised any VC money. I don't believe anyone has ever done this before. So you guys are actually achieving the dream of what people are describing will happen with AI. I'm curious, just do you think this will happen more and more as a result of AI? And also just where has AI most helped you find leverage to be able to do this?

A Yeah, so we hit over a billion of revenue last year with under a hundred people. And I think we're going to see companies with even crazier ratios, like a hundred billion per employee in the next few years. AI is just going to get better and better and make things more efficient. So that ratio just becomes inevitable. Like I used to work at a bunch of the big tech companies, and I always felt that we could fire 90% of people and we would move faster because the best people wouldn't have all these distractions. And so when we started Surge, we wanted to build it completely differently with a super small, super elite team. And yeah, what's crazy is that we actually succeeded. And so I think two things are colliding. One is that people are realizing that you don't have to build giant organizations in order to win. And two, yeah, all these efficiencies from AI. And they're just going to lead to a really amazing time in company building. Like, the thing I'm excited about is that the types of companies are going to change too. It won't just be that they're smaller. We're going to see fundamentally different companies emerging. Like, if you think about it, fewer employees means less capital. Less capital means you don't need a raise. So instead of companies started by founders who are great at pitching and great at hyping, you'll get founders who are really great at technology and pro…

AI assessment note: “I think we're going to see companies with even crazier ratios”

Answered raw tape D 4 · C 4 · P 3 · Cm 3 3.60

Q Well, let me follow that thread to enlightening this answer. Uh, do you have any advice for how to Build those sorts of experiences that help lead to that. Is it, you know, follow things that are interesting to you? Because, you know, it's easy to say that it's hard to actually acquire these really unique sets of experiences that allow you to create something really important.

A Yeah, so I think it would always be to really follow your interests and do what you love. And it's almost like a lot of decisions I make about Surge. Like, I think one of the things that I didn't think about a couple of years ago, but then someone said it to me, it's that companies in a sense are an embodiment of their CEO. And it's kind of funny. I hadn't thought about that because I never quite knew what a CEO did. I always thought a CEO was kind of generic and like, okay, you're just doing whatever VPs and your board and whatever tell you to do and just saying yes to decisions. But instead it's this idea where when I think about certain big hard decisions we have to make, I don't think what would company do? I don't think what metrics are we trying to optimize? I just think what do I personally care about? Like what are my values and what do I want to see happen in the world? And so I think following that idea about, okay, so ask yourself, what are, what are the values you care about? What are the things you're trying to shape and not what will look good on a dashboard? I think that that's pretty important.

AI assessment note: “I think it would always be to really follow your interests and do what you love.”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.