Efficiently Learning To Reason or Not to Reason: Root-token Policy Optimization for Adaptive Thinking.
Taehyeon Kim, Hyunsoo Lee, Youngsoo Jang, Moontae Lee
Browse the full ACL paper archive.
Taehyeon Kim, Hyunsoo Lee, Youngsoo Jang, Moontae Lee
Browse the full ACL paper archive.