From General Reward to Targeted Reward: Improving Open-ended Long-context Generation Models.
Zhihan Guo, Jiele Wu, Wenqian Cui, Yifei Zhang, Minda Hu, Yufei Wang, Irwin King
Browse the full EMNLP paper archive.
Zhihan Guo, Jiele Wu, Wenqian Cui, Yifei Zhang, Minda Hu, Yufei Wang, Irwin King
Browse the full EMNLP paper archive.