Skip to content

Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint.

Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, Tong Zhang

VenueA*ICML
Year2024
ProceedingsICML

Browse the full ICML paper archive.