Skip to content

Reward Difference Optimization For Sample Reweighting In Offline RLHF.

Shiqi Wang, Zhengze Zhang, Rui Zhao, Fei Tan, Cam-Tu Nguyen

VenueA*EMNLP
Year2024
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.