Joseph: Training safety into models is more robust than using system prompts
“I think you get different amounts of robustness if it's trained into the model versus if it's in a prompt that you can, like, add or remove or tell, like, ignore all previous instructions, that sort of thing.”
Mann: Anthropic bases Constitutional AI on UN Human Rights and Apple terms
“And then another piece that's come out is constitutional AI, where we have this list of natural language principles that leads the model to learn how we think a model should behave. And they've been taken from things like the UN declaration of human rights and…”
Clark: Dario Amodei's early Constitutional AI concept sounded completely crazy
“On, I think, like, week four we were talking about RL and language models, and Jared was like, oh, Dario says we're just gonna write a constitution for the AI and it'll just follow that. And I remember being like, that's completely crazy. Why would this ever w…”
Byun: Anthropic's Constitutional AI slashed Elicit's query costs tenfold in days
“At the start of twenty-twenty-three, Anthropik kind of launched their constitutional AI paper and within a few days, I think four days, he had basically implemented that in production, and then we had it in-app, like, a week or so after that, and he has since …”
Lambert: Practitioners Will Adopt Constitutional AI for Preferences in 2024
“I think in twenty-twenty-four at some point people will start doing things like constitutional AI for preferences.”
Lambert: Anthropic Constitutional AI and OpenAI Superalignment share intellectual roots
“The constitutional AI and the super alignment is, like, very conceptually linked. It's like a group of people that has, like, a very similar intellectual upbringing, and they work together for a long time, like, coming to the same conclusions in different ways…”
Amodei: Anthropic's core AI safety constitution is five pages long
“Our set of principles is in our constitution. It's very short. It's five pages.”