A Minimaximalist Approach to Reinforcement Learning from Human Feedback.
Gokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu, Alekh Agarwal
Browse the full ICML paper archive.
Gokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu, Alekh Agarwal
Browse the full ICML paper archive.