Skip to content

RLHFPoison: Reward Poisoning Attack for Reinforcement Learning with Human Feedback in Large Language Models.

Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik, Chaowei Xiao

VenueA*ACL
Year2024
ProceedingsACL (1)

Browse the full ACL paper archive.