An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning.
Wei Sun, Qianlong Du, Fuwei Cui, Jiajun Zhang
Browse the full ACL paper archive.
Wei Sun, Qianlong Du, Fuwei Cui, Jiajun Zhang
Browse the full ACL paper archive.