The Exchanges

Every argument clarity score on this site is built from rows on this page. Each question and answer was assessed with names hidden, the host's own answers included, on four things from 1 to 5: directness (does it answer the question asked), coherence (do the ideas follow), precision (concrete details and clear references), compression (says a lot per word). The weighted mix (30/30/25/15) is the exchange score. A person's published score averages their exchange scores on raw tape only, at least 8 of them, shrunk toward the cohort mean. Full method →

Ryan Greenblatt no published score: no usable exchanges on raw tape, and a fair score needs 8+ · coarse estimate ≈4.5/5 from 13 produced feed exchanges record → ← everyone

Every exchange below was scored with names hidden, four dimensions each from 1 to 5. An exchange's score is 0.30·directness + 0.30·coherence + 0.25·precision + 0.15·compression. The published score averages the raw tape exchange scores and shrinks small samples toward the cohort mean, so five great answers can't beat twenty good ones. Produced feed rows count only toward coarse estimates, never toward a full score.

clear all ✕
13exchanges match
0on raw tape
0redirected or not addressed
Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q And just to verbalize, uh, the question early in this conversation, uh, what is so bad about super intelligence? Obviously, uh, there's a lot of talk about, uh, scientific progress and curing cancer, um, and your, uh, documents, uh, effectively recommends pausing the rates towards the super intelligence. So why, why is it so bad?

A Yeah, I wouldn't say super intelligence is bad. I would say it's dangerous. Like it's a very dangerous Thing to create. So the most straightforward story for why, um, it's dangerous is that it seems pretty likely that on sort of a trajectory similar to the trajectory we seem to be finding ourselves on, um, you end up with AI takeover as a result of building super intelligence because the AIs are in a position where they can take over due to being highly capable, widely deployed, and, um, you know, uh, building basically, uh, the huge amount of industrial capacity potentially. Um, and we could talk more about what a takeover would look like. And then if they're in the position where they could take over, then there's a question of like, would they want to, or how would the motives shake out? And it looks like we don't really have that much control over the motivations of AIs. And it seems like that problem gets harder as they're, you know, much more capable and built via a process where AIs are automating AIR and D and we maybe are losing our understanding of how that process works. There's sort of this, like, we don't necessarily control the technology, misland AI takeover. Another concern is that Historically, like, you know, um, at least in recent times, uh, the distribution of power among humans has been reasonably distributed, though not necessarily, like, super, super dist…

AI assessment note: “I wouldn't say super intelligence is bad. I would say it's dangerous.”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q Yeah, reading Redwood research stuff over the years, it seems to, that there has been an evolution from being focused largely on interpretability to much more AI control. Is that, is that fair?

A Yeah, that's fair. So I would say that like our arc as an organization was we, when I joined the organization, I just finished up a project on adversarial training and was interested in getting into like doing interpretability and what we would call like model internals work where it's like, can we take advantage of the fact that we have white box access to these models to do something, you know, better than just the naive methods of sort of prompting and training when like trying to like align these models, understand their motives, like know what's going on. And we explored that area for a while and then for a mix of reasons decided it was like quite a bit less promising than we'd initially hoped and decided to move on to other things. And one of the things we moved on to shortly after that was, uh, AI control, um, which is the idea that maybe it would be a good idea to prevent AIs from being capable of accomplishing problematic things, or basically make it so the AIs aren't able to cause huge problems, even if the AIs wanted to. With, you know, there's a bunch of different stories for why this is a good idea, but basically the idea is, like, there may be some intermediate period, an intermediate period that I would say we're currently in, where the AIs are maybe capable enough to cause at least moderate problems, and then I think increasingly able to cause quite large proble…

AI assessment note: “Yeah, that's fair. So I would say that like our arc as an organization”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q And as an aside, by the way, there's, uh, one of the parts of the write-up that I find the most fascinating is, uh, uh, your model says, uh, world GDP could grow roughly 200 X, uh, during the 20 thirties under the restrain plan. Can you, can you talk to that? Like the, the number is staggering. I mean, we're in a world of, uh, three percent GDP growth.

A Yeah, yeah, yeah. So, so a key part of our Perspective is like, even, you know, even non super intelligent AI systems at the level of capability we discussed would be radically transformative across the, you know, across the world and like, and just like for all kinds of different things. And we are imagining sort of slowing down AI development some, or like going at a more cautious pace for some period, and then eventually hitting a level of capability where the AI can basically automate basically everything that humans can do. And staying at that level of capability for a while, where we work on safety and security and at that level of capability, those AIs would be capable enough to be doing huge amounts of autonomous R and D. And in addition, It would be totally possible to have robots that are basically like, you know, more capable than humans at manufacturing and industrialization and so on. And in particular, you could have robots build robots. And so you can end up in a situation where you have huge amounts of robotic industrial capacity that is itself building more robotic industrial capacity that can then produce downstream goods. And that total capacity can basically grow very fast. I think we propose like Limiting that growth somewhat for various reasons with like various types of taxes, but like, uh, we can, we're imagining sort of the robot population or like, you…

AI assessment note: “robot population... doubling or quadrupling every year, which... means the economy would double”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q were a couple of parts to the discussion that I found particularly interesting. In particular, uh, there was this argument that, uh, RSI may be very good at automating the process of creating the next generation of AI, but it may or may not have the, uh, the kind of intuition, uh, that one needs For a scientific breakthrough and truly novel ideas. So what's the rebuttal to that argument?

A Yeah. So, um, the way I put this argument is AIs might be very good at sort of the more mechanistic or nitty gritty parts of AI development, like writing code, running experiments, but not as good at sort of the broader conceptual leaps. So first I would say that like AIs seem Significantly better at engineering and grungy stuff and sort of just keeping trying than they seem to be at conceptual breakthroughs, but their ability to do sort of these. Breakthroughs, especially in easy to verify domains are improving. And like, you know, one example, um, is like their ability in math, but even in like, you know, ML, their taste has been improving. I think it continues to improve. Um, and so it's not, it's not so clear to me that this will lag super far behind. Another thing is that, uh, you can measure how good these AIs are at intuition or research taste or having breakthroughs, especially in domains that are relatively easier to verify. And if you can measure it, then you can take your grungy AI Um, labor or even just your human laborers and try to optimize that. So, you know, if it can be measured, it can be hill climbed on very roughly speaking, at least. And I think this is a case where you could just hill climb on how good the AIs are at making these sorts of breakthroughs in a wide variety of different settings. And then I expect that would transfer to making the actual break…

AI assessment note: “if it can be measured, it can be hill climbed on very roughly speaking”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q All right. So, uh, going back to AI, 2040, give us a quick version of, uh, what it is, who wrote it and what is the main thesis? And then we'll go into some details.

A Um, there's sort of two components. AI-TV is a scenario, um, focused on, like, what the authors, including me, think is, like, a plausible good route for things to go, or, like, a reasonable plan, at least in some circumstances. Um, and it's written by Thomas Larson, Daniel Cucutello, me, uh, Eli Lifland, um, Brendan and Romeo. And I would say, like, The, the, the basic story is like, how would you do a deal with China to make AI development both be safer and also so that we can sort of like hang around at a point that's short of super intelligence, but where the AIs are still really, really capable for a long time so that we can study those systems and have a longer time to sort of integrate them into the economy, understand how things will go and so on. Where I think a concern that we have is like on the default trajectory, you maybe go straight from like AI systems that are like competitive with humans to AI systems that are wildly superhuman in a very short period of time. And that seems like. Uh, quite scary in a variety of ways. Um, in addition to that, we worry about like AI development being insufficiently transparent for sort of third parties to provide a reasonable check to AI companies and whether their plans will work. And we sort of have a unified proposal that solves a bunch of these different problems and makes it so that, for example, you can pay a huge amount o…

AI assessment note: “it's written by Thomas Larson, Daniel Cucutello, me... the basic story is”

Answered produced feed D 5 · C 5 · P 4 · Cm 4 4.60

Q And you know, as you were describing this, uh, we were talking about continual learning earlier. So whether that's continual learning or something else, if there was a technique that appeared that, um, made the compute a lot more efficient and, uh, you were able to, uh, just vastly reduce the, the compute effort, especially for those pre-training runs, how would that impact the plan?

A Yeah. So if, let me give a hypothetical and then talk about what I think the realistic case is. So if hypothetically it was the case that tomorrow a recipe for training super, super intelligence on like, you know, 64 H 100 or like some small amount of compute just dropped. I think we'd have super intelligence very fast or like that would be my sense. And it would not be possible to do this sort of deal, but I think there's, there's, it's not, um, but, but the hope with restricting compute Isn't just that like, you know, AI development is currently very compute hungry. It's that also that R and D is very compute hungry. And so even in a regime where you were doing continual learning, before you had the version of continual learning where you could do everything on a single, you know, on like some tiny amount of compute, you're going to have a shittier version that can do it on a moderate amount of compute. In order to develop the version that can do it on a tiny amount of compute, based on the history of error progress you need, you would need a lot of compute to do that research. Or I shouldn't say need, but in practice that, that would come about through a lot of compute. Now, if it was the case that there was some alternative research direction, Which in practice was a lot less compute hungry, and which got quickly very developed, quickly developed during this period, and whi…

AI assessment note: “it would not be possible to do this sort of deal”

Answered produced feed D 5 · C 4 · P 4 · Cm 3 4.15

Q Very helpful. So what's your latest prediction on timing for RSI to happen?

A Yeah. Um, so maybe my median for, let's just say like full automation of AIR and D by which I mean basically like, even if humans left the picture, things wouldn't slow down by that much would be like. Maybe end of year, or maybe like early, not that much precision in these numbers, but like something, I don't know, like, like, or like, you know, not that much stability, like these numbers fluctuates on, but that'd be my guess for median. But then I think that's sort of the, the central scenario I planned for, which is maybe more like my. 35th percentile would be like end of year, 20, 28 slash beginning of 20, 29, which I think is like very, very likely. I think that seems super plausible. And I think that sort of, if I just like extrapolate out the current Trajectory. It looks like more like you've got that mile, that trajectory. Um, and the reason why I don't think that's my median is there's just a bunch of factors that might kick in to push things back. Like maybe there's some like key bottleneck I'm not seeing. Um, maybe there's some. Like something that I think will work to overcome some obstacle won't actually work, or maybe there's going to be, um, significant government slowdowns because people, you know, freak the out about this technology, which seems kind of plausible. Um, and so, so because of that, I push later, but I think that like, in terms of what I would reco…

AI assessment note: “35th percentile would be like end of year, 20, 28 slash beginning of 20, 29”

Answered produced feed D 5 · C 4 · P 4 · Cm 3 4.15

Q Obviously early 2028 is like tomorrow morning effectively. Um, is there a scenario where all of this is already too late?

A Yeah. Um, I mean, it depends on what you mean by too late or like in, in what all of this means. I think I am worried that specifically like AI 20 40 plan A is assuming like too much like Effective, like government time, or like it's assuming the government has like more time than it actually does to take all these actions. And I think it's pretty realistic that like the actual plan we should go for is going to be, should be in practice, like, uh, quite a bit, let's just say like sloppier and like, you know, less well organized and faster, just because like, we just don't actually have time to do something quite that elaborate. I'm not confident in that. I mean, I think there's a bunch of different options, but like in general, I would say to like, Yeah, we, it might be too late for some interventions. I think like various policy windows, if you follow the normal timeline or closing, that said, I think that there's a long history, at least in the U S of like, in times of crisis, things can happen much faster. And there's a lot of different levers for that. And so I think that if there's, if we got to a position where everyone is like, holy shit, we need to take this specific action that could happen very quickly, but we might not get to a position where there's that much consensus. And also it might be that the government just isn't tracking AI or isn't aware of AI To a suffici…

AI assessment note: “Yeah, we, it might be too late for some interventions.”

Answered produced feed D 5 · C 4 · P 4 · Cm 3 4.15

Q And you also make the argument, uh, that, um, for purposes, purposes of, of a potential discontinuity, whether it transfers or whether it generalizes may not even be, uh, really a question. And that if it was super good, just at the industrial part of accelerating AI, that would be enough. Is that fair?

A Yeah, I would say that, like, if AIs could just automate AIR&D and automate, like, sort of the industrial process of building more computers, then you could quickly end up in a In a process where sort of like robots are building robots and the whole world is greatly transformed. And that very quickly can get you to a point where AIs could take over, you know, the economy has been radically changed, even if it's hard to train AIs at some other tasks. But I do think that my sense is that once the AIs are, you know, once the situation is like there are robots building robots that build computers and the full feedback loop is closed at that level, my sense is by that point, the AIs will be good at You know, basically all human work, or at least not, not too deep into that. Maybe there's some period where, you know, Robotics is a big deal, but the eyes can't quite automate a bunch of like, there's a bunch of stuff they can't automate, but it seems to me like that that's how it's going to go. But just in general, sort of automating R and D seems like it's enough to radically transform the world. And if you look at sort of why is human, why, why is like humanity a big deal? Like why, why have we been able to accomplish so much collectively? A lot of it is because of just like, you know, having technology, having organizations, being able to organize ourselves in various ways and accom…

AI assessment note: “Yeah, I would say that, like, if AIs could just automate AIR&D”

Answered produced feed D 5 · C 4 · P 4 · Cm 3 4.15

Q You know, up until recently, my general sense is that, um, all those ideas about, like, pausing and stuff here were kind of, um, perceived, at least by a portion of the tech world, as, um, kind of, like, decel, cosplay, and were, like, doomerism that was, uh, based on poor understanding of what AI actually does. Is there, is, is the, the, the general mood, uh, turning for good?

A Yeah, I would say that, like, overall, there have been, you know, some positive elements here, though. I think I would say, like, it's not obvious that people are reacting to events as much as they should, but they are reacting some. And I think that, like, there has been quite a bit of, I think there has been quite a bit of evidence that, you know, there's some reasons to be worried and some reasons that we might need to, like, get our shit together to handle some of these safety and security problems. And that might require Shifting a bunch of resources or slowing down. So we have time to manage various things or whatever, um, pacing the frontier as, as they say, or whatever. I think that like different of these things seem like they're going differently. Well, so like, I think, um, my sense is that there's been quite a bit of buy-in from AI company employees to, you know, take, take this stuff more seriously and do something about, um, about misalignment risk. But I'm not necessarily so sure that like, that's actually like amounted to that much yet. Um, other than companies sort of putting in a decent amount of effort. Um, but it hasn't like amounted to any sort of very durable long run thing. Um, and then there's, uh, I think the government got very freaked out about cyber capabilities and then just more generally was like, well, we need to like be overseeing this technolog…

AI assessment note: “overall, there have been, you know, some positive elements here”

Answered produced feed D 5 · C 4 · P 4 · Cm 3 4.15

Q We will go back to, um, in a minute, um, but I wanted to do a quick segue about you and your story. What first pulled you into AI safety?

A Yeah, so, um, I was, uh, in my junior year in college, um, and I was sort of alone in my apartment because it was COVID, and I was listening to a lot of podcasts, and I was sort of thinking a bit about what I should do with my life, and I ended up Through some somewhat twisted path, I ended up thinking I should like be way more interested in like helping other people and being altruistic than I, than I was at the time. And I should be very focused on like, how can I make. You know, just sort of like other people's lives as good as possible and make things go as well as possible. And then from there, I considered a bunch of different routes and was looking into a bunch of different things and was researching different possibilities and eventually decided that the best thing I could do with sort of my career and my life was try to make AI go better and in particular avoid AI takeover, but also more generally sort of, you know, try, try to make that go better. And then I applied to a bunch of places. I ended up working at Redwood. This was about, um, You know, about five years ago at this point, um, and I was just been, I've been working there since, and I've done a bunch of different work there. And, you know, sort of the field has really evolved a lot since then, because, you know, five years ago, it was like, GPD, 3.5 hasn't, it wasn't released yet. Uh, I remember when like tex…

AI assessment note: “eventually decided that the best thing I could do... was try to make AI go better”

Answered produced feed D 4 · C 4 · P 3 · Cm 3 3.60

Q And, uh, to get a concrete about, uh, the dealing plan A, what exactly would the U S give China? What would China give the U S and, um, how do you maintain the balance so that, um, you know, nobody, uh, cheats?

A So the core of the deal or like the start of the deal is understanding where all the computers, because compute is this really important driver of AI progress, where if you like sort of stopped the flow of compute, You would probably stop the flow or mostly stop the flow of AI progress, or these things would slow at some point, um, after, you know, the, the existing amount of compute, uh, diminished, or if you cut off all the, uh, if you like turned off all the computers, things would certainly stop. And so first we sort of try to find all the compute, then in order to make sure that the deal is stable and that each side isn't sort of racing to secure as much advantage as they can, uh, you, you stop training. And switch to just doing inference, um, and basically stop most of the R and D. Um, and then you try to track down as much of the computer as possible. And this is like, you know, both in the U S and China, but also in other places where compute resides, like, uh, you know, various like countries in Southeast Asia, Europe, Australia, whatever. And you have to get everywhere where there's enough compute in on the deal and make sure you track it down enough. And if you don't do that, then I think probably you have to pursue some option that's less ambitious, or at least you, you do something less ambitious than plan A in terms of the level of slowdown and probably do less tr…

AI assessment note: “to make sure that the deal is stable... you stop training. And switch to just doing inference”

Partly produced feed D 3 · C 4 · P 4 · Cm 3 3.55

Q Uh, as we get, uh, towards the end of this conversation, the question in my mind, uh, is What do we do in the meantime as a policy gets decided and plan A, B, or C happens? What, uh, what can we do? We started talking about AI control. Uh, so in terms of the reality today, state of the art of AI control, what works and what doesn't work?

A I'll talk specifically about AI control. So the thing that needs to happen on AI control is we need to get to the point where We basically understand all of the AI traffic within at least AI companies. We have some ability to over, like, to look, like, look at that traffic, monitor it. And then we have some pipeline for flagging particular examples to be further investigated. And that eventually escalates to humans actually looking into particular examples and seeing how concerning they are. And that pipeline has the ability to also block traffic in cases where we're like, whoa, something weird is going on that we don't understand, or that looks obviously concerning. We should like stop these AIs from proceeding and potentially also stop some other similar AIs from proceeding until someone has the time to check it out and adjust things. And in addition to that, we also want to just sort of greatly improve the security, right? So like, you know, if these AIs are, you know, really broadly deployed, it would be very nice if we could sort of make it so that we have very precise invariance about they're supposed to be able to do and not supposed to be able to do and have a permissioning system that allows for that. And then if AIs need to request some escalated permissions, then we could carefully track that rather than just giving AIs all the permissions by default. We sort of woul…

AI assessment note: “the thing that needs to happen on AI control is we need to get”

page 1
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.