Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model.
Wenhong Zhu, Zhiwei He, Xiaofeng Wang, Pengfei Liu, Rui Wang
Browse the full ICLR paper archive.
Wenhong Zhu, Zhiwei He, Xiaofeng Wang, Pengfei Liu, Rui Wang
Browse the full ICLR paper archive.