Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. Have the thread models been updated a lot, or do you feel like you're still using the same thread models as GPT-II of like, you know, paperclip factory, blah, blah, blah, you know, but like, how much are you rising, you know, increasing the bar?
A Yeah, so I'm not an expert in, in the threat modeling piece, more in the, more in the capabilities piece. Um, I, I do think they've been changing to, to some extent. So, something like the autonomous replication threat model, that is being able to set yourself up and, and control resources, something like that, has been deprioritized relative to, um, AR and D acceleration. That is, you know, the possibility there could be some capabilities explosion inside of, inside of a lab, and that could be destabilizing for, for all sorts of reasons that we could talk about. Um, so it's mainly, mainly we're focusing on that latter one, although, although we do think about a, a, a number of trend models.
AI assessment note: “I do think they've been changing to, to some extent.”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q How did you pick the tasks? I would say that's one question that people have. You have some labels kind of like Trin classifier, fixed bugs, and small Python library. They all seem kind of arbitrary, you know, like what's the process of task selection?
A People are right to be worried about, uh, task selection, um, or, or there are many, um, many finicky details in here. I would say the aspiration was to pick economically valuable tasks Relevant, especially to sort of general autonomy and R&D, the, the, the threat models that we're primarily, um, primarily interested in. So, you know, what one misreading of the time horizon graph is this is referring to, you know, the full distribution of like any, any tasks that you might give AIs. And I think that's, that's clearly not right. You know, in particular tasks that are requiring of vision capabilities, they're probably, um, to take one example, they're probably much less capable today, um, as measured by time horizon, as, as for, for these tasks that are typically not requiring vision, vision capabilities that we give them. So we, we try and, you know, we, we sample these tasks by having people inside of meter create the tasks and by, um, having a bounty so that people from, uh, from outside of meter can, can, can provide us with these tasks, stuff, stuff like this, you know, that, that's not a sort of perfectly random selection process. In particular, it's, it's a process that has a bunch of constraints, you know, in order to be able to, uh, scalably run our revals, it's helpful, not necessary, but helpful for the, for success on the tasks to be automatically gradable. And that m…
AI assessment note: “we sample these tasks by having people inside of meter create the tasks”
Answered raw tape
D 4 · C 5 · P 4 · Cm 4 4.30
Q How do you validate previous research? So like take the developer productivity study, right? You had AI slowed people down. If you were to redo it with Opus 4.5, would you expect the results to be dramatically different? And should we redo the study? Like, should we stop coding the study? Like, how do you think about that?
A We have been redoing it in the background. I think it's, um, and I won't comment on exact results, but, but I think it is, it is much harder to do it today than it was in the past for all sorts of reasons. You know, the first is as AIs get better at coding, it's harder and harder to find, you know, developers submitting tasks who are willing to, to be randomized to, to AI disallowed. There's a quote unquote selection issue where, you know, maybe we end up only observing the tasks that they, that they thought AI wouldn't greatly uplift them on. Uh, ahead of time, because they're, you know, those are the tasks that they're willing to be, to be paid for, to, to be flipped into AI disallowed. There are other issues, like I think today a common workflow is to work on multiple issues or multiple lines of work at the same time concurrently, and that wasn't, wasn't really true before. It's, it's difficult to know how to capture that in our, in our study design. If you flip a single task to be AI allowed or AI disallowed, you know, you're sort of supposed to work on that single task, but actually that's not how developers are working today. I think basically these weren't Threats to the previous study design or in, you know, approximately like March, 20, 25, people weren't really working concurrently or not, not nearly to the same degree. They basically were giving, giving us, giving us…
AI assessment note: “We have been redoing it in the background. I think it's, um, and I won't comment”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Yeah, we have not made, we would have made a lot more money trading on embargo news than we've made on anything else. Um, what else? What are other interesting model evaluation trajectories or like anything that you're not doing at Meter that you've maybe seen other people do that you find interesting or like you would like more people to, to do?
A Yeah, one project that I think is interesting is, um, AI Village that I think possibly both of you would have come across. These are these, um, very open-ended goals given to a village of agents, and they try to accomplish them. It's like, um, I think set up a merchandise shop is maybe one of them. Organize an event in a park, build a human subjects experiments, this, this sort of thing. I think I have a number of questions about the, um, uh, about exactly what I should learn. You know, they're using old models as well as new models in this quote-unquote village. Um, the models are relying a lot on vision capabilities, which we, we spoke about models, um, not being so capable of today, this, this sort of thing, but the vibe of models trying to achieve open-ended things instead of benchmark-like tasks, you know, the vibe that's a bit more like that, um, a vending machine bench, um, in some ways, it seems like a very interesting direction to me for the, for the science to go, or seems like, you know, something that comes with a lot of cons, but attacks some of the, some of the cons of benchmarks in a pretty interesting way. I think seeing the, the ways in which these models trip up, seeing the ways in which they're, in which they're derpy is a, is an important source of information. Uh, I'd be interested in, in, in more work like that coming about. I think that's one of them. Ano…
AI assessment note: “one project that I think is interesting is, um, AI Village”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q What's the, what's the thing you're looking for?
A So there are lots of different shapes of people we can look for. So, so different, different stuff for, for, for, for different folks. One thing is, is, um, you know, good kind of basic research intuitions, like, Um, like checking your data. You know, we don't work on free training at Meta, but if you're working on free training, you should look at the corpus to get some sense of what's going into the models, even working on this uplift RCT. Um, that was, that was, that was pretty important. You know, really, really having a shape of, of these issues in your head. I think people who are communicating in writing with, with sort of a lot of transparency, not overstating their results. Um, you know, my hope is that your sense of, of Meta work in the past is that it's, um, it's trying to be level-headed, not, not to, Um, not to understate, not to overstate what the, what the, what the science says. That's important internally. It's, it's, it's important. It's important externally, you know, and then I think productivity or something that there are a lot of, um, a lot of people with, um, uh, great talents who, who are not going to, um, work quite as well in a, um, scrappier environment, working on, working on sort of frontier science, and, um, and that's, and that's the thing we do.
AI assessment note: “One thing is, is, um, you know, good kind of basic research intuitions”
Answered raw tape
D 4 · C 4 · P 4 · Cm 4 4.00
Q Yeah. How did you get the two together and then some of the findings that you had?
A Yeah. Maybe for a second, let's take time horizon very literally. We don't have the, the qualms about it that we've, that we've just been discussing. Um, it makes sense to continue extrapolating it into the future. What are some important forces that might cause it to rise more quickly? Some of the things we've just been talking about, automated R&D versus, um, versus go more, go more slowly. One of the most obvious forces that might cause it to go more slowly is if inputs slow. Um, one important input is compute. You know, I think, I think we all have the intuition that to some extent if, if compute growth slows, um, which we expect it to at some point in the not so distant future, then capabilities will slow. But by how much? It's a big, it's a big question. The suggestion in this, in this paper is that if you think that algorithmic progress, you know, that, that is coming up with the transformer, coming up with RLHF, you know, MOEs, all of, all of this stuff, better learning rate schedules is, is, um, uh, is itself a function of compute because, you know, you, you need to compute to, to discover it. You know, the, the transformer, the gains from, from transformers, um, show up much better with scale. If you don't, if you don't put in those resources, you know, you'll never find out that this is, um, this is the, uh, superior algorithm. You need to run a ton of experiments, y…
AI assessment note: “The suggestion in this, in this paper is that if you think that algorithmic progress”