Saunders: AI-assisted critique helps humans detect model errors
William Saunders · What The Ex-OpenAI Safety Employees Are Worried About · Jul 3, 2024 · at 8:13
Former OpenAI Superalignment researcher William Saunders explains technical oversight methods developed at OpenAI to help humans evaluate AI answers beyond their domain expertise.
“What some of the research that I did at OpenAI was, you know, trying to develop techniques for this, and the simple technique that we tried and showed that worked was just ask a different AI system, or even the same AI system, is there a problem with this answer you gave me? And then the AI will sometimes put out a list of problems and say like, hey, this part of the answer was like made up. It's not supported by any evidence, or you're leaving out this important piece of information. And then we showed that if you just, if you show people both an answer and a set of possible problems, people are better at spotting when the AI system has made a mistake.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →