PILAF: Optimal Human Preference Sampling for Reward Modeling.
Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng, Julia Kempe, Yaqi Duan
Browse the full ICML paper archive.
Yunzhen Feng, Ariel Kwiatkowski, Kunhao Zheng, Julia Kempe, Yaqi Duan
Browse the full ICML paper archive.