Skip to content

Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning.

Zhaohui Yang, Yuxiao Ye, Shilei Jiang, Shihong Deng, Chen Hu, Linjing Li, Daxin Jiang

VenueA*EMNLP
Year2025
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.