Skip to content

SPPD: Self-training with Process Preference Learning Using Dynamic Value Margin.

Hao Yi, Qingyang Li, Yulan Hu, Fuzheng Zhang, Di Zhang, Yong Liu

VenueA*EMNLP
Year2025
ProceedingsEMNLP (Findings)

Browse the full EMNLP paper archive.