Skip to content

Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons.

Banghua Zhu, Michael I. Jordan, Jiantao Jiao

VenueA*ICML
Year2023
ProceedingsICML

Browse the full ICML paper archive.