Skip to content

Low-probability Tokens Sustain Exploration in Reinforcement Learning with Verifiable Reward.

Guanhua Huang, Tingqiang Xu, Mingze Wang, Qi Yi, Xue Gong, Siheng Li, Ruibin Xiong, Kejiao Li, Yuhao Jiang, Bo Zhou

VenueA*ACL
Year2026
ProceedingsACL (Findings)

Browse the full ACL paper archive.