Countering Reward Over-Optimization in LLM with Demonstration-Guided Reinforcement Learning.
Mathieu Rita, Florian Strub, Rahma Chaabouni, Paul Michel, Emmanuel Dupoux, Olivier Pietquin
Browse the full ACL paper archive.
Mathieu Rita, Florian Strub, Rahma Chaabouni, Paul Michel, Emmanuel Dupoux, Olivier Pietquin
Browse the full ACL paper archive.