Skip to content

Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience.

Zhenwen Liang, Ruosen Li, Yujun Zhou, Linfeng Song, Dian Yu, Xinya Du, Haitao Mi, Dong Yu

VenueA*ACL
Year2026
ProceedingsACL (1)

Browse the full ACL paper archive.