Skip to content

J4R: Learning to Judge with Equivalent Initial State Group Relative Policy Optimization.

Austin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong, Shafiq Joty

VenueA*ACL
Year2026
ProceedingsACL (1)

Browse the full ACL paper archive.