Skip to content

A Comprehensive Survey on Learning from Rewards for Large Language Models: Reward Models and Learning Strategies.

Xiaobao Wu

VenueA*EMNLP
Year2025
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.