The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Zico Kolter argument clarity score 4.1/5 from 14 exchanges on raw tape · average scores: directness 4.5 · coherence 4.3 · precision 3.9 · compression 3.3 record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
14exchanges match
14on raw tape
2redirected or not addressed
Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q So I'm curious from your, your vantage points, the, the, the whole acceleration is to versus doomerism, uh, debate that has been raging for the last couple of years that seem to, you know, come and go depending on the moment. Uh, is that, is that at all helpful? Is that how you, you think about it?

A I, I, I, I dislike those labels. A lot on both sides. I think they're oddly enough used as largely, uh, pejoratively by both sides, right? People will dismiss someone as a doomer if they express too much concern about risks of AI systems, or if someone's trying to release models, they'll be called an accelerationist. It's, it's all, I mean, people, some people then, you know, use the terms of pride, I guess, but they're sort of inherently kind of dismissive terms, I think. Um, I believe I am on, I, I, I, I, I have never expressed a P. Doom and things like this. I just think it's a very weird concept as if the world is some stochastic set of dice that you can roll multiple times and that we don't have direct influence over this. Um, so I, I think that, um, I think that the reality is, and the, the, these sort of labels tend to, um, it tends to sort of dismiss a lot of the, the, the reality of the situation right now, which is that AI is not a technology that is, that is wholly bad, in my view, and it's not a technology that has no risks either, that just, we can just, you know, develop however, with no constraints whatsoever. Um, and I would say that, I think. 95% of all researchers, maybe 99% of all researchers feel probably a very similar way that, you know, this technology has great promise. There are massive opportunities, but we have to be mindful of the risks. It's sort of…

AI assessment note: “I dislike those labels. A lot on both sides.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q Mm-hmm. And speaking of the Chinese, is, is, is safety a global movement? Like the, the way you have some level of, uh, uh, cooperation in conferences.

A Yeah. There are certainly efforts in many different countries. Um, uh, I'm less familiar with the Chinese efforts, but there are efforts in China certainly, but there's lots of safety in AI safety institutes or AI security institutes in many different countries. So. The UK obviously was the first AI safety, now AI security Institute. Uh, but Singapore has one as well. The US has the Casey, which, which, uh, uh, does similar function. Um, And many other countries have sort of burgeoning institutes as well. There's definitely global understanding of this problem. Now I do think that, um, these things are subject to some degree of political headwind and the fact that the, you know, AI safety, uh, conference was, or AI safety summit was renamed the AI action summit or something is, has some significance actually in terms of the sort of taking temperature of, of where the, Where the world is politically. But, at the same time, I also think a lot of the work being done Is, is a very similar nature. The, the actual researchers and what they're doing, um, they've people, these organizations have continued to do great work, continue to push the frontier and understanding how to assess, how to evaluate systems, how to safeguard them, all these things, they are happening in an ongoing fashion. And I think, um, you know, the good, good work is being done by researchers at companies in acad…

AI assessment note: “Yeah. There are certainly efforts in many different countries.”

Answered raw tape D 5 · C 5 · P 4 · Cm 4 4.60

Q What happened then? Like how did the labs react?

A When the models were constrained to just be the models themselves, this is not that easy to patch. I mean, you can patch single strings. A lot of labs sort of blocked individual strings, um, that we had published just Which is fine, right? But if you ran the whole process again, you could find another string that would actually actually circumvent it. It wasn't until the development, A, of additional safety classifiers that people started to really kind of be able to detect and stop these things. But then also reasoning models. Reasoning models were much more effective because you can't really do the same trick of optimizing for a probability with a reasoning model that has a whole trace of reasoning that happens in the middle and kind of reflect a bit more. So it's much harder to break reasoning models in the same way. But yeah, the, the, the short is that there, there was certainly some work done to address these things, but it took additional layers of security and, and, and security and the advent of reasoning models before they really became ineffective.

AI assessment note: “A lot of labs sort of blocked individual strings”

Answered raw tape D 5 · C 5 · P 4 · Cm 3 4.45

Q Where do safety issues come from? Is that Uh, the models, um, uh, get better at reasoning. Therefore they can come up with good or bad ideas. The data set.

A Yeah. So, so I think, I think to answer this question, you have to unpack a little bit about AI safety. It's a, it's an extremely broad term. And I would actually argue that it has to be a broad term because the truth is there are fundamentally different questions related to AI safety that all kind of go under this moniker. And, and frankly, a, a, a, A challenge that sometimes people use this same term to refer to very different problems. I typically kind of think of four categories of risks of AI, and this is a, I, I hate, and all ontologies are wrong, to be clear, and this is, or maybe some are useful, but that's debatable actually. Um, this one's very much wrong and incomplete, but I sort of think about AI risk as spanning kind of a, a spectrum from basically risks that come from just mistakes of the model. And on the sort of category one, this is includes hallucinations. It includes the model, just making silly mistakes sometimes not knowing what to do and just getting things wrong. Right. Um, prompt injections actually an aspect of this. We can talk about prompt injections more, but they're basically other people being able to fool the model just cause the models a little bit, doesn't really understand the full context, doesn't understand things. So that's, that's sort of number one. So model must kind of silly mistakes. I know I don't want to use the word silly click on t…

AI assessment note: “one side of safety issues come from the model making mistakes”

Answered raw tape D 5 · C 4 · P 4 · Cm 4 4.30

Q Was there like any sense that, uh, this was going to become what it is, uh, today? Was, the, the ambition was always there, right?

A The ambition was always there, and Ilya was always an ambitious person, and, you know, many of the people there, um, were, Always extremely ambitious. Frankly, you know, they saw things that I did not see at the time. I, I remained continually surprised, not just by stuff that happened in AI, but things happening kind of broadly in the field, right? I eventually started to just felt like, man, I gotta stop being so surprised. That's when I kind of got a little bit more, you know, AI-pilled, right? Um, but I think that, that, um, the interesting thing that I remember About, about opening up early on is that they always had this bet on scale. In a time where I think that was looked upon very suspiciously, um, that, oh, if you, the thought somehow that we had all the methods already, and all you had to do was scale them up, uh, that mindset had not pervaded academia. Um, academia was still obsessed with, we need new methods. We need new approaches. That's what's going to lead to Breakthroughs in AI systems. Cause for a long time, it kind of arguably had. I mean, um, Rich Sutton has this great, this very famous essay called the bitter lesson that kind of argues this, um, though he doesn't love LLMs either. He thinks LLMs are actually not bitter lesson enough. Um, so, so, uh, I remember the real, that real philosophy on scale that I think folks probably like, well, I didn't know, I …

AI assessment note: “The ambition was always there, and Ilya was always an ambitious person”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q And in the cat-and-mouse game between attackers and defenders, uh, so the flip side, what, what is the state-of-the-art of, um, attacks? Is it like a new kind of, like, prompt injection?

A Right. So the state-of-the-art, and, and I think it's actually some things, for example, that develops, I mean, I'll actually say things that are outside of the group. I, uh, outside of, outside of my work, you know, I think, for example, some of Gray Swan's work in, in, in our automated red teaming methods is, is some of the state of the art. Um, the, the, some of the state of the art techniques, what they do, and I think the, the UK AC published one, um, uh, these recently, uh, what you do is you sort of, You use many, many queries to the classifier, to these guardrail classifiers, or I shouldn't say guardrail classifiers, but these, these sort of, uh, uh, input and output classifiers to find their boundaries. Kind of actually in a very similar attack is very similar to, um, to GCG, but you sort of probe their boundaries. You also then include a jailbreak for the underlying model, and you also include a jailbreak, a similar sort of jailbreak for the output model. So you have to kind of Develop simultaneously jail breaks for each of these. Um, and it is doable. Now it takes many, many queries as far as we know how to do it to these safety classifiers. So you need a lot of data, uh, from the models to really do that well. And it's something that, you know, again, you, your accounts will be flagged if you try to do this in the wild. So it's, it's, it's this kind of thing where t…

AI assessment note: “some of Gray Swan's work in, in, in our automated red teaming methods is”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q What do you advise your PhD students to Focus on what are some of the exciting directions that you recommend?

A Yeah, I've, so I've mentioned before, right, about sort of the trends and talking about, uh, you know, doing research in academia on AI safety, working on fields like robotics, where I think there's really need for fundamental new methods before, before we're quite at the pure scaling phase, uh, and then science, basic science. So those are, those are, those are, I mean, we just had our visit days for new PhD, newly admitted PhD students, so I can, Talk very confidently about. This is what I sort of talk about. Um, but the actual, the bigger thing I would say is, um, you should actually just work on what you're excited about, is the real advice for PhD students. Um, if you are excited about something that, that I think is just completely wrong, you should go and work on it, because progress will be made by people. This is like, this is a famous statement, right? I mean, there's so many statements that, you know, uh, I don't want to use the more morbid such statements of them, but, but, uh, basically progress happens when The current crop of young researchers ignores the things they've been taught that the, the old guard believes. And, and look, I, I mean, I, I think I'm adaptive to new technologies and, you know, fairly malleable, but I'm sure I'm not as malleable. And actually, uh, I'm more stuck in my ways than I ever want to admit. And so you should ignore everything I'm say…

AI assessment note: “work on what you're excited about, is the real advice for PhD students.”

Answered raw tape D 5 · C 4 · P 4 · Cm 3 4.15

Q But should they be in production from a security standpoint?

A Yes, I think so, actually. Um, I think if you run with proper guardrails, you know, we, we, we release guardrails for, um, for coding agents, for example. Um, if you're on proper guardrails with proper sandboxing, and right now, yes, you probably also take some care to be a little bit careful in terms of what control authority you give to your agents. They can clearly do a whole lot. They can clearly be beneficial. And again, it's a risk reward kind of thing, right? So do the benefits outweigh the risks? I think so. I mean, I certainly use them. I don't write code anymore. I do all my work now, and I do lots of, you know, I still do some research, right? It's entirely telling Codex what to do. So yes, we should be using agents.

AI assessment note: “Yes, I think so, actually. Um, I think if you run with proper guardrails”

Answered raw tape D 4 · C 4 · P 4 · Cm 4 4.00

Q Yes. What's in the water? And, um, and, uh, as a related question, like, how do you Fair in a, in a, in a world where, uh, you know, so much is going on in industry and the gravitational pull of industry is so strong.

A Yeah, it's a great, it's a great question. So first of all, CMU, I mean, look, I think, I think CMU and a few other institutions to be clear, you know, have, have emerged kind of as, have been fortunate to emerge kind of as global leaders in driving the field forward. Um, you know, since the inception of the field, right? Uh, when, when, um, Newell and Simon were, were building a logical theorist back in the, back in the fifties. Um, I think I'm getting the name of that wrong. I think it's called logical theorist, but it might be something, it might be something a little different. Um, I think in some sense what's enabled CM places like CMU, um, but CMU in particular, I think is a bit of a willingness to take risks. So CMU has a structure where we have a whole school of computer science. So we're not in an engineering school. We're not in some of those, we have a school of computer science. We've had that for a very long time and it sort of enabled a degree of experimentation and, you know, forming something like a machine learning department. And that's more than 25 years old now. There weren't a lot of people thinking you have a whole department in machine learning 25 years ago, and Tom Mitchell was one of the people that did. Uh, and so I, I think that this ability to sort of take risks because you have a bit more autonomy is something that really has driven at least the his…

AI assessment note: “what's enabled places like CMU in particular, I think is a bit of a willingness to take risks.”

Answered raw tape D 5 · C 4 · P 3 · Cm 3 3.90

Q Great. Taking a step back on, um, this whole, um, safety and security discussion, do you think in two years from now we are more secure and more safe as an industry or less?

A I think, I mean, I, I think we're definitely gonna be more secure and more safe. Uh, I, I, I think that, I mean, look, in some sense, I expect the trajectory that we are on right now will continue. And when I say that, what I mean is it's kind of mind boggling to actually think that when you, when you realize what the trajectory has been the last three years, right? Um, I think that there's going to be massive advances and just widespread deployment of these things. They'll be act much more longer term, much more autonomously, all this kind of things that will, will, will happen. The models will be, but again, the, the, the challenge, what the, what the challenge is, is not sort of just to make that more safe because it will be more safe. But the question is, is the, is the, Safety and the safety work that we're doing going to be commensurate with the increase in, in control surface, in actuation surface, and all those kind of things, right? Um, and that's what I work, that's, that's what I work on, right? Is, is ensuring, ensuring that we are on the trajectory to match the increase in capabilities.

AI assessment note: “I think we're definitely gonna be more secure and more safe.”

Answered raw tape D 4 · C 4 · P 4 · Cm 3 3.85

Q What do you think, um, happens in the next, um, year in terms of like likely breakthrough? I mean, I guess everybody's talking about continual learning. Is that something that's happening?

A I mean, look, there are going to be breakthroughs, right? So yes, uh, so continual learning, um, it's not clear to me that we don't already know how to do this to a certain extent. I mean, if we really did the serious thing of taking, you know, your data, um, your interactions, Generate synthetic data from those. Retraining on that. Having some sort of LoRa model, which would be your, your model, your memory. Even just having some amount of sort of, ah, compressed KV cache. This is sort of the cache that stores context for, for these models. Um, it's really unclear to me we don't get a lot of this already. It hasn't really been deployed in production yet, but it's not clear to me we don't have the technology already for a lot of these things. However, could there be more breakthroughs? Absolutely. And, and, On a small scale, I mean, I think you have sort of a major advance, like, like models in general, and maybe I would say reasoning models were the next big breakthrough. Those are, those are rare. They, they do take kind of a, you know, both, both a massive scale and kind of a bit of, uh, of, of luck to get there. Um, but are there going to be breakthroughs? Absolutely. And maybe one of them will, will be the one that we look back and say, yeah, that kind of just, you know, that was continual learning there. There's no, there's no more issues.

AI assessment note: “continual learning, um, it's not clear to me that we don't already know how to do this”

Redirected raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q Extremely important. You mentioned, uh, the various teams, uh, at OpenAI around, um, safety, security, uh, can you, can you provide a bit more color about, like, how that's organized internally?

A Yeah, I mean, so, so the safety systems, I mean, there, there are different groups there, and the organization is a little bit, uh, I shouldn't say changing, but it is sometimes, it is a little bit flexible, the precise organization. But the main point I want to highlight is not the precise sort of structure of those teams, but what the different teams do. Um, so one example would be the preparedness team at OpenAI. So preparedness is a public, public framework. OpenAI's released the preparedness frameworks. I think the first one was released in February of 2024, actually before I joined the board, and then we've updated it a few times since then. What preparedness is, is essentially a document that lays out kind of, uh, certain conditions that have to be met when models reach certain capabilities. And this is a nice way I think of thinking about kind of safety from a, from a model release perspective, right? Uh, to be very clear, not all safety issues fit into this framework. This is more about things like catastrophic harms that models may be capable of, but The idea of preparedness is that when models reach a certain level of capability, right? This can be used positively for many situations, of course, but it also can be used by bad actors in a harmful manner. So as models get better in basic biological knowledge, they can be used by malicious actors that want to misuse tha…

AI assessment note: “main point I want to highlight is not the precise sort of structure of those teams”

Redirected raw tape D 3 · C 4 · P 4 · Cm 3 3.55

Q Well, thanks for this. Um, let's switch to the, um, let's actually go into the, the, the, the substance of the, um, sort of safety and security, uh, field. So, uh, you provided, uh, upfront a, a bit of a taxonomy, uh, maybe to double click on some of this, uh, Uh, what's the difference between safety and security?

A Right. So, so, okay. Security. So I, I laid out this, this sort of four pillars of AI safety, right? You know, mistakes and, and, uh, and harms and, and, and, um, societal effects, loss of control. Um, security is more. Is, is, is a slightly separate term. And I, and I want to actually, the, the real thing I want to differentiate actually is AI security from between AI security, as I think about it, which is the security of AI systems themselves. You know, what new security issues do AI models and agents introduced by way of being AI systems? Um, and AI for security, which is sort of also on very much top of mind right now, which is basically how can we use AI to, Address, to address or exacerbate, uh, traditional security concerns. Uh, what, what I work on and what we, for example, at Gray Swan, but really most of my research works on is AI security. So how can we make AI models themselves fundamentally more robust to manipulation? Security fundamentally is about how well do models or systems react to adverse pressure, to adversarial pressure to the systems. So most evaluations are done kind of in a, They measure expected value, basically. They measure sort of how well does it work on average, and security measures how well does it work in the worst case. That's what security is. Um, and so AI security is basically how well do models work in the worst case. Uh, especially when…

AI assessment note: “the real thing I want to differentiate actually is AI security”

Answered raw tape D 4 · C 4 · P 3 · Cm 2 3.45

Q Is that, is that something that happens routinely or, or is that completely exceptional?

A We, we, we don't want to talk too much about sort of the details of the sort of the details of sort of how it happens there, but, We have these meetings for every release, and we actually have them for, for every major model release, and we actually have them a lot also for, um, just prior to a release, we'll of course be in a lot of touch with researchers, understanding that the nature, so there aren't surprises usually, right? Um, really it is an oversight role. So again, I know corporate governance is just thrilling to talk about, but for those that know corporate governance, uh, it's, it's, it's not dissimilar to the role of an audit Audit committee, right? So an audit committee is sort of overseas finances, talks with the CFO a lot, kind of views a lot of things the company's producing for reports to the SEC and stuff like that. Um, and I, I think it's actually very important that AI companies start to establish similar governance policies, because this is something that requires that level of just oversight and of assurance. It is a, it is a becoming You know, a massive industry, and just like there are audit committees of boards, I think it's very important, and I would hope to see more of these going forward, for AI companies in particular to have things like safety and security committees, by whatever name they have, that oversee the sort of the model release and gover…

AI assessment note: “We have these meetings for every release”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.