Skip to content

ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization.

Yuxuan Zhang

VenueA*ACL
Year2026
ProceedingsACL (1)

Browse the full ACL paper archive.