reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs.
Zhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Yoon Kim, Asli Celikyilmaz, Marjan Ghazvininejad
Browse the full EMNLP paper archive.
Zhaofeng Wu, Michihiro Yasunaga, Andrew Cohen, Yoon Kim, Asli Celikyilmaz, Marjan Ghazvininejad
Browse the full EMNLP paper archive.