Skip to content

Confronting Reward Model Overoptimization with Constrained RLHF.

Ted Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm, Ruslan Salakhutdinov, Anca D. Dragan, Stephen Marcus McAleer

VenueA*ICLR
Year2024
ProceedingsICLR

Browse the full ICLR paper archive.