Skip to content

On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization.

Yong Lin, Skyler Seto, Maartje ter Hoeve, Katherine Metcalf, Barry-John Theobald, Xuan Wang, Yizhe Zhang, Chen Huang, Tong Zhang

VenueA*EMNLP
Year2024
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.