Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Given the model evaluation cycle and the fact that performance does not Asymptote for many tasks over, um, quite a long period of time. What do you do about that issue? The fact that some of the evals that you would want to run are both beyond the scope of budget or time that's reasonable given the current model release cycle?
A I mean, I think for things like cyber, we've seen, and actually the AISI, um, in their evaluations has shown that the models continue to improve, um, at A hundred million tokens. You know, if you run them for a hundred million tokens, they're still improving at beyond that point. And that can take a very long time to run, but you also do see that like the performance is, is it's not just like a discontinuous jump. It's actually like, you can see the slope of improvement over those hundred million tokens. And so you could, you could probably do some kind of, uh, evaluation up to a certain budget and then just say, okay, well, this is, What we project the performance to look like. And I think this, this, there hasn't been a lot of research on this yet. I actually think this would be a great paper to publish if there's any academics out there looking for something to research. Can you predict what the performance looks like at an inference budget of, let's say, 10,000 dollars only using inference budgets up to 10 or a hundred, or 10 or a hundred dollars?
AI assessment note: “you could probably do some kind of, uh, evaluation up to a certain budget and then”
Answered raw tape
D 5 · C 5 · P 5 · Cm 4 4.85
Q Let's talk about the larger implications of You know, needing to evaluate these models relative to, um, let's say, like, speed of their reasoning or efficiency versus, you know, token volume, right, um, or dollar budget or whatever, whatever your scaler is. Can you describe some of the larger implications in your essay, including around, um, like, safety evaluations?
A Yeah, the safety evaluations thing, um, it's, it's a bit of an inconvenient truth thing, where Okay, so I guess for background, a lot of the, all the labs have these things called either responsible scaling policies, preparedness frameworks, they go by various names. But the idea is that whenever a model is released, they go through a series of evaluations to measure, are there dangerous capabilities? Um, could these models do things that we're, we, we wouldn't want, um, a bad actor to do? And if the model isn't very capable, then it's no big deal. But if it is very capable, if it could be used, for example, to make bioweapons, then you want to put in mitigations against that. But the question is, okay, well, how do you evaluate whether the model is capable of that? And they have, like, various protocols about, like, how they do these valuations. But a lot of these frameworks were developed around the era of ChatGPT, either before or after, when test time compute scaling was not really, uh, as much of a thing. And it made sense. Like, with GPT-III, you couldn't scale test time compute. Like, if you gave it a budget of ten million dollars and said, okay, well, let's see what GPT-III can do, it really can't do that much, more than what you could do with, like, 10 dollars or one dollar. The preparedness frameworks and responsible scaling policies, they don't really account for the…
AI assessment note: “preparedness frameworks and responsible scaling policies, they don't really account for”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q our end goal is AGI and it's always been, and I think sometimes that's sort of invented later as sort of an interesting story for what they're doing. Did you view this as just doing primary research and it's just personal interest? Did you view it as like, there's a path leading to agents that function on behalf of people or, or was there some other sort of driving motivator?
A Well, so I started grad school in, in, in, and it was a very different time in, uh, you know, The idea of AGI was, was really science fiction. Um, there were, there were some people that were, you know, uh, serious about it, but, but very few, the majority opinion was that AI was, if anything, it was kind of a dead field. Um, I actually remember like emailing a professor and having this conversation where I was like, look, I'm really interested in AI, but I'm, I'm kind of worried to pursue a PhD in this because, you know, I, I, I get the impression that it's just a dead field and I'm not, I'm worried I'll be able to get, if I'll be able to get a job afterwards. Conveniently, like a couple of years into grad school, Things changed pretty drastically and, um, and I happened to be in the right place at the right time. I think I was really fortunate in that respect. So the, the original intention wasn't to, to pursue AGI. The original intention was, you know, you learn interesting things about, um, AI and game theory and you, you build slowly. Um, and it was really only a, a couple of years into grad school that it became clear that the pace of progress was, was quite dramatic.
AI assessment note: “So the, the original intention wasn't to, to pursue AGI.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Do you think we're going to get bots that, um, negotiate with humans soon? Let me promise that is we are eventually going to get them. What do you think the timeline is or the use case?
A That, that seems doable. It depends on how How constrained the domain is. I think if you were to look at constrained domains, uh, certain negotiation tasks, I think that AIs could probably do better than humans in that today. I mean, I'm trying to think of like specific examples, but things like, um, you know, if you wanted to negotiate over the price of a good, um, it could probably do better than, than a human in a lot of those, in a lot of those situations. I think if there's things like salary negotiations, um, It might be better than humans at that also. Um, I think it depends on how much you need to know about the world. I think contract negotiations, for example, would still be difficult because there's so much subtlety. There's so much nuance to like every contract and it's not gonna replace the professional negotiator for that kind of task just yet. Um, but kind of the things that are more constrained don't require as much like outside knowledge about the world. I think AIs are probably up to the task already.
AI assessment note: “I think that AIs could probably do better than humans in that today.”
Partly raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q What, what measure do you think, or measures do you think makes sense to use? And then also, what do you think is missing on sort of the road to general intelligence?
A Um, I think there's a few things that are missing. The big, the big thing that I'm interested in particular is reasoning capabilities. You have these bots and they can, they can, they basically, they're all doing next word prediction, right? Cicero is a bit different actually in that it's actually conditioning its dialogue generation on a plan. Um, and I think that's one of the really interesting things that, that distinguishes Cicero from a lot of the work that's happening in language models today. But a lot of the research that is happening is, is using next word prediction. And when it's trying to do something that's like, More sophisticated in terms of reasoning capabilities. It's a lot of chain of thought where it's just like rolling out, you know, the kind of reasoning that it's observed humans do and they're in, in its training data and seeing where that leads. I think there's a general recognition among AI researchers that, um, this is a big weakness in the bots today. And that if we want truly general artificial general intelligence, then this, uh, this needs to be addressed. Now there's a big question about how to address it. And that's actually why I really like this direction because it's still an open question. About how, how to actually fix this problem. There's been some progress, but I think there's a lot of room for improvement.
AI assessment note: “I think there's a few things that are missing. The big, the big thing”
Partly raw tape
D 3 · C 4 · P 4 · Cm 3 3.55
Q How did that, if you look at a lot of other games, the, those sorts of, uh, big shifts in performance from a bot relative to people then shifts how people play, right? They learn from the bot or they adapt their game from watching games that the bots play. How did that play out in terms of poker?
A Oh, that's, yeah, that's a great question. So, you know, the competition, it was really, it was really interesting because, you know, so kind of like as a last minute thing, we, uh, we added this ability. So, okay, the way the bot works, We give it different bet sizes that it can use. Like there's the game that we were playing. There's 20,000 chips, 100 dollar, 200 dollar blinds, um, or 5100 dollar blinds actually. Um, and so it can bet any, it can bet any amount it wants from like a hundred dollars up to 20,000 dollars. Um, and so. There's not much value in like being able to bet both 5000 dollars and 5001 dollars. And so we would discretize that action space to constrain it to like only considering a few different options. And so there's a question of like, okay, well, what sizes do you give it the choice between, you know, towards the end when we were developing this bot, like we just had room for extra competition. And so we just like threw in some extra sizes, um, like four, four X, the pot, 10 X, the pot, like it, it doesn't cost that much more. So why not just give it the option? Um, I didn't think it would actually use those sizes. And then during the competition. It, it actually ended up using those sizes a lot. Um, and it would sometimes bet like, you know, 20,000 dollars into a 100 dollar pot, which was completely unheard of in a professional poker play. And, um, You…
AI assessment note: “completely unheard of in a professional poker play... thought it was a mistake at first”