Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint.
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, Tong Zhang
Browse the full ICML paper archive.
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, Tong Zhang
Browse the full ICML paper archive.