Bostrom: AI misuse is a governance challenge, not a technical one
“You're focusing there on the misuse potential that this people might choose to do bad things with AI technology. And that certainly is one big category of risk, right? But that's not primarily a technical challenge. It's more ultimately a governance challenge …”
Bostrom: Weak aligned superintelligence could help align stronger superintelligence
“If you get a kind of weak super intelligence that is For the most part aligned, we might then be able to use that to make a more powerful form of super intelligence that is more reliably aligned.”
Bostrom: Natural language AI interfaces provide a safer alignment buffer before superintelligence
“This gives us more sort of surface area to work with. Like you can more easily understand and interact with these systems because they have human double concepts and you can talk with them.”
Ries: Human alignment, not technical alignment, is AI's top unsolved problem
“This is the number one unsolved problem in AI. It's not the tech, we're making great progress on the technical alignment problem, but we haven't made jack progress on the human alignment problem, which is that we've known since the development of Conway's laws…”
Midha: Human alignment is a bigger challenge than technical AI alignment
“AI alignment, don't get me wrong, is hard, but not the hardest problem. Human alignment is really the problem right now.”
Emmett Shear: Goal-based alignment covers only a tiny fraction of human experience
“Goals are one level of alignment. You can align something around goals. The kind of goals we're talking about here are one level of alignment. You can align something around goals by like if you can explicitly articulate in concept and in description, the stat…”
Johnson: Project Blueprint is fundamentally about AI alignment
“What a lot of people don't realize is this endeavor is entirely about AI alignment.”
Pre-training aids AI alignment by implicitly instilling human values
“I definitely think we would keep using pre-training data, not just from an efficiency point of view as well, but also I think there is interesting safety angles, because by pre-training and, you know, all this human knowledge, we're implicitly creating an agen…”
Schrittwieser: AI safety must span the entire stack, not just RL
“Yeah, I wouldn't view it alignment adjust like an RL problem. I think it sort of, it goes throughout the whole stack. You might, you know, for example, filter the pre-training data in some way. You might, after training, you might have classifiers that, you kn…”
Tworek: AI alignment is a never-ending pursuit as human goals evolve
“And it's I think it's a never ending pursuit because like, even, even for humans, it's not super easy to define what's, what do we consider a light? And I think as our civilization will evolve, it will, the notion of alignment and the goals of humanity will, K…”
Fisher: Economic pressure for long-horizon agents will drive AI alignment progress
“I'm actually really positive and bullish that there is this economic pressure in a good way To make progress on alignment because long horizon agents require it.”
Anthropic's Joseph: Certain AI Alignment Pieces Will Move to Pre-Training
“I do think at some point there will be, like, some pieces of alignment that, like, you do want to export back into pre-training because that might be a way to, like, Put them in with more strength, like, more robustness, kind of, or more core to the intelligen…”
Mann: Language Models Understand Human Values in a Core Way
“And since then, my estimation of how hard the problem would be has gone down significantly actually because things like language models actually do really understand human values in a core way. The problem is definitely not solved, but I'm more hopeful than I …”
Hendrycks: AI alignment is only a subset of AI safety
“So I view the distinction between alignment and safety as alignment as being a sort of subset of safety. Obviously you want the value systems of the AIs to be in keeping with or compatible with say the US public for USAIs or for you as an individual, but that …”
Steinberger: AI alignment is only solvable via recursive automated models
“The only way to sort of reasonably approach this is to iteratively ask your model to solve alignment and safety at that stage, not, not, you know, surely you can also ask it to solve your product level problems, but like that, that's nice, but that's not the f…”
Bach: Coercing LLMs into good behavior is unsustainable
“At the moment, the idea that we build LLMs that are being coerced with good behavior is not really sustainable. Because if they cannot prove that the behavior is actually good I think we are doomed.”
Polosukhin: Society needs human alignment rather than AI alignment
“So I have this view that we need human alignment instead of AI alignment. So right now, kind of when we talk about, you know, hey, we need to align AIs with like human values, but the reality is that, you know, all the problems that exist, they all exist becau…”
Andreessen: AI alignment has fundamentally become social engineering and politics
“And for that second risk expressed as AI alignment, what you're dealing with fundamentally is social engineering. And when you're dealing with social engineering, you're necessarily dealing with politics.”
Mayya: Rule-based AI alignment fails because edge cases cannot be enumerated
“Alignment is about figuring out all these edge cases and saying, don't do this, basically putting it in a sheet. And the problem is we just don't know what it'll end up doing even this world. So nobody can write all the edge cases.”
Masad: AI alignment reflects Silicon Valley sensibilities, not average humans
“I think a lot of what's called AI alignment today is not really aligning with what the average human being wants. It's aligning with like what the sort of Silicon Valley average sensibility is, which I don't think it's good.”