Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning.
Qianxi He, Qingyu Ren, Shanzhe Lei, Xuhong Wang, Yingchun Wang
Browse the full EMNLP paper archive.
Qianxi He, Qingyu Ren, Shanzhe Lei, Xuhong Wang, Yingchun Wang
Browse the full EMNLP paper archive.