Every argument clarity score on this site is built from rows on this page. Each
question and answer was assessed with names hidden, the host's own answers included, on
four things from 1 to 5:
directness (does it answer the question asked), coherence (do the ideas follow),
precision (concrete details and clear references), compression (says a lot per word). The weighted
mix (30/30/25/15) is the exchange score. A person's published score averages their exchange
scores on raw tape only, at least 8 of them, shrunk toward the cohort mean.
Full method →
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q But maybe it just happens super, super fast, 24 seven. Or, or is there like a new machine language that emerges? Yeah.
A Um, I mean, I think one thing that's really helped us so far in AI development is, uh, to come in with some priors for, um, you know, how humans do things. And that's actually, um, you know, if you bake those priors in, they typically are great starting points. So I could imagine, like, maybe you start with something that's Slack-like and give it enough flexibility that it can kind of develop beyond that and really figure out the way that's most effective For it to communicate. Um, one important thing though is, uh, you know, we want interpretability too, right? I think it's, it's very helpful for us today that what the agents do is, you know, uh, easy for us to read and interpret. And I don't think you want that to go away as well. So I think there's a lot of benefits just even from a pure, like debug the whole system perspective, or just let the models, you know, speak in a way that it's familiar with us. And, you know, you, you can also imagine like we might want to plug in To the system too, right? So, you know, um, whatever interfaces we're familiar with, we would ideally like our model to be familiar with as well.
AI assessment note: “maybe you start with something that's Slack-like and give it enough flexibility”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q dropping a buzzword on it is like program synthesis. Is there a world where, uh, I know that, I know the tokens, like the images, we see them as, as renderings of squares and different colors, but, uh, the, when they're fed into the LLM, they're typically, uh, just a stream of, of numbers effectively. Is there a world where actually adding a screenshot is what's important? Like visual reasoning.
A Yeah. Yeah. So I think, I think that could be important. It's just like kind of, uh, you know, whenever it comes to like textual representation of grids, um, models today just don't really do that well, right? And I think it's just kind of because humans don't really ever write down textual representations of grids, right? Yeah. You know, we have a chess board, like no one really kind of just like types it out in a grid. Um, and, um, and so the models are kind of like, Undertrained a little bit on, on what that looks like and what that means. So, um, you know, I, I think with more reasoning, we'll, we'll, we'll just bridge the gap. Um, I think with better visual perception, we'll just bridge that gap.
AI assessment note: “So I think, I think that could be important.”
Answered raw tape
D 5 · C 5 · P 4 · Cm 4 4.60
Q Yeah. And I guess like when I think of agents, it's, it's deep research is like, you know, curating Information on which you can take action on, but it's like, at what point is action a part of that sort of loop, right? Where you can not only curate a list of flights that you want, but then, you know, actually go out and, and, and have agency.
A I think one of our explorations in that space is operator, right? It's where you kind of just feed in raw pixels from your, your, your laptop, um, into, or, you know, from some virtual machine into the model and it, it produces, you know, either a click or some keyboard actions, right? And Um, so there it's taking action. Um, and I think the trouble is, you know, it, you don't ever want to mess up when you're taking action, right? I think the cost of that is super high. Um, uh, you, you only have to get it wrong once to lose trust in, in a user. Um, and so we want to make sure that that feels super robust, uh, before we get to the point where we're like, Hey, look, here's a tool. Um,
AI assessment note: “we want to make sure that that feels super robust, uh, before we get”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q Because, yeah, I mean, it's like if, if the reasoning tokens are inference tokens and, and, but they're what lead to higher intelligent, more intelligent models, like it's almost back in the training bucket again. Um, what bucket should we be thinking about and, and, uh, and, or, or are we, how firmly are we in the, the, the, uh, the, uh, the applied AI era versus the research era?
A Well, I think research is here to stay, and it's for all the reasons I mentioned above, right? It's such a, like, a rich time to be doing research, but I do think, you know, inference is going to be increasingly important as well, right? It's such a core part of RL, um, that you're doing rollouts, and I think, you know, we see 25 as this year of agents, right? Um, we think of it as a year where models are going to do a lot more autonomous work. You can let them Kind of be unsupervised for much longer periods of time. And, um, that is just going to put big demands on inference, right? When you think about kind of our overall vision, right? We, we lay it out as a series of steps and levels on the way to AGI, right? And I think the pinnacle really that last level is organizational AI, right? Like you can imagine a bunch of AIs all interacting. Um, and yeah, I think that's just going to put huge demands on inference, right?
AI assessment note: “Well, I think research is here to stay, and it's for all the reasons”
Answered raw tape
D 5 · C 4 · P 4 · Cm 4 4.30
Q new Harry Potter. Everyone has to read it. It's amazing. And it's fully aged and it's fully generated or, uh, or this image, the images, they do go viral, but they go viral because they're AI move. 37 in the context of go did not go viral because it was AI. It felt like it was Actual innovation. So, uh, is that the right frame? Does that make any sense?
A Yeah. Um, I think it's not the wrong frame. So I think some, some quick thoughts on, on, on that. Um, I, I think kind of, um, when you have something that's, you know, very measurable, like win or lose, right? Something like, uh, like go, um, yeah, it's like very easy for us to kind of just judge, right? Like did, did the model do something right here? Um, and I think the more fuzzy you get, um, you know, it, it is just harder, right? Like, um, when it comes to, you know, is this the next Harry Potter, right? Like, you know, it's not a universally loved book. I think it's fairly universal, but you know, there's, there's some haters. Um, and yeah, I, I think it's, it is just kind of hard when it comes to these human subjective things where, um, it's really hard to put down in words, like what makes you like Harry Potter, right? And, um, and so, Um, I think those are always gonna lag a little bit, but, you know, I, I think, you know, we're, we're developing more and more techniques to attack kind of these more open-ended, um, uh, domains. And I don't know, I, I wouldn't say that we're not at an innovative stage today. So, um, I think my biggest touch with this was when we had the models compete on the IOI last year. So, uh, IOI, it's like the, the international, basically Olympics for, for computer science, um, Basically, the, the top four kids from, from each country go and comp…
AI assessment note: “I think it's not the wrong frame... when you have something that's... very measurable”
Answered raw tape
D 4 · C 4 · P 4 · Cm 3 3.85
Q got it correct, but you actually took the correct path. Kind of like, you know, you're graded for your work, not just the answer if you're in grade school. Um, And, and, you know, Dario said that, uh, interpreting interpret interpretability research will actually contribute to capabilities and even give a decisive lead. Do you agree with that? What's your reaction to that concept of interpretability research being very important?
A Yeah. I mean, we care a lot about it here at opening as well. So, um, one thing that we care a lot about is interpreting how the model reasons, right? Um, because I think, um, We've had a very kind of specific and strong view on this, um, in that we don't want to apply optimization pressure to how the model thinks so that it can be faithful in the way it thinks and to expose that to us, you know, without any kind of incentives to cater to what the user wants, right? I think it's actually very important to have that unfiltered view, um, because, you know, uh, Oftentimes, like, if, if the model isn't sure, you don't want to hide that fact, right? Just for, for it to kind of please the user. And sometimes it really isn't sure, right? And, and so we've, Really done a lot of work to try to promote this norm of chain of thought faithfulness and, and interpretability. Um, and I think it, it gives you a lot of, uh, sense into what the model's thinking and, you know, what are the pitfalls that it can go off into if it's not reasoning correctly.
AI assessment note: “we care a lot about it here at opening as well.”
Answered raw tape
D 4 · C 4 · P 3 · Cm 3 3.60
Q IOI sample problems, I think this would be a 20 year process for me to figure out how to achieve that, and I can do the Arc AGI on my phone. Uh, is this the spiky intelligence concept? Is this something that a small tweak in, in algorithmic design just one shots AGI, Arc AGI, or, or is there something else going on there that we should be aware of?
A Yeah. I mean, I think, um, part of this is the beauty of RKGI as well, right? Like, um, I think I'm not sure if there's another kind of like human intuitive, simpler benchmark, which is for the models. Um, and I think really that's one of the things they really optimize for on that benchmark. Um, I do think when it comes to models though, like there's just a little bit of a perception gap as well. Like, you know, uh, models aren't used to this kind of native, um, you know, like Just screen type input. Um, I think there's a lot we can bridge there. Actually, um, even O for many, um, it's a state of the art multimodal model in many ways, including visual reasoning. And I think, uh, you know, you're starting to kind of build up the capacity for the models to take images, manipulate and reason about them, um, or generate new images, write code on images. And, um, I think it's just been kind of under focused, but, um, I think When I talk to researchers in the field, they all see this as a part of intelligence too, and we're going to continue to focus there.
AI assessment note: “when it comes to models though, like there's just a little bit of a perception gap”