Provable Offline Preference-Based Reinforcement Learning.
Wenhao Zhan, Masatoshi Uehara, Nathan Kallus, Jason D. Lee, Wen Sun
Browse the full ICLR paper archive.
Wenhao Zhan, Masatoshi Uehara, Nathan Kallus, Jason D. Lee, Wen Sun
Browse the full ICLR paper archive.