Skip to content

Joint Reward and Policy Learning with Demonstrations and Human Feedback Improves Alignment.

Chenliang Li, Siliang Zeng, Zeyi Liao, Jiaxiang Li, Dongyeop Kang, Alfredo Garca, Mingyi Hong

VenueA*ICLR
Year2025
ProceedingsICLR

Browse the full ICLR paper archive.