Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning.
Haolin Liu, Dian Yu, Sidi Lu, Yujun Zhou, Rui Liu, Zhenwen Liang, Haitao Mi, Chen-Yu Wei, Dong Yu
Browse the full ACL paper archive.
Haolin Liu, Dian Yu, Sidi Lu, Yujun Zhou, Rui Liu, Zhenwen Liang, Haitao Mi, Chen-Yu Wei, Dong Yu
Browse the full ACL paper archive.