In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement Learning.
Songjun Tu, Jingbo Sun, Qichao Zhang, Yaocheng Zhang, Jia Liu, Ke Chen, Dongbin Zhao
Browse the full AAAI paper archive.
Songjun Tu, Jingbo Sun, Qichao Zhang, Yaocheng Zhang, Jia Liu, Ke Chen, Dongbin Zhao
Browse the full AAAI paper archive.