TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback.
Eunseop Yoon, Hee Suk Yoon, SooHwan Eom, Gunsoo Han, Daniel Wontae Nam, Daejin Jo, Kyoung-Woon On, Mark Hasegawa-Johnson, Sungwoong Kim, Chang Dong Yoo
Browse the full ACL paper archive.