Human-in-the-loop: Provably Efficient Preference-based Reinforcement Learning with General Function Approximation.
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, Liwei Wang
Browse the full ICML paper archive.
Xiaoyu Chen, Han Zhong, Zhuoran Yang, Zhaoran Wang, Liwei Wang
Browse the full ICML paper archive.