Legal Mathematical Reasoning with LLMs: Procedural Alignment through Two-Stage Reinforcement Learning.
Kepu Zhang, Guofu Xie, Weijie Yu, Mingyue Xu, Xu Tang, Yaxin Li, Jun Xu
Browse the full EMNLP paper archive.
Kepu Zhang, Guofu Xie, Weijie Yu, Mingyue Xu, Xu Tang, Yaxin Li, Jun Xu
Browse the full EMNLP paper archive.