Skip to content

Direct Value Optimization: Improving Chain-of-Thought Reasoning in LLMs with Refined Values.

Hongbo Zhang, Han Cui, Guangsheng Bao, Linyi Yang, Jun Wang, Yue Zhang

VenueA*EMNLP
Year2025
ProceedingsEMNLP

Browse the full EMNLP paper archive.