Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling.
Pankayaraj Pathmanathan, Furong Huang
Browse the full ACL paper archive.
Pankayaraj Pathmanathan, Furong Huang
Browse the full ACL paper archive.