WildReward: Learning Reward Models from In-the-Wild Human Interactions.
Hao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao, Lei Hou, Juanzi Li
Browse the full ACL paper archive.
Hao Peng, Yunjia Qi, Xiaozhi Wang, Zijun Yao, Lei Hou, Juanzi Li
Browse the full ACL paper archive.