Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons.
Banghua Zhu, Michael I. Jordan, Jiantao Jiao
Browse the full ICML paper archive.
Banghua Zhu, Michael I. Jordan, Jiantao Jiao
Browse the full ICML paper archive.