tree enforcement learning with human feedback
1 statements across 1 episodes · 1 bullish · 0 bearish · 1 people on the record · first statement Sep 25, 2023 by Mira Murati · across every show →
Everything said about tree enforcement learning with human feedback, oldest first
Sep 25, 2023 positive
Murati: Scaling RLHF alone might be sufficient to solve AI hallucinations
“And we also wanted to figure out the issue of hallucinations, which is always an extremely hard problem. But I do think that with this method of tree enforcement learning with human feedback, maybe that is all we need if we push this hard enough.”