Skip to content

Countering Reward Over-Optimization in LLM with Demonstration-Guided Reinforcement Learning.

Mathieu Rita, Florian Strub, Rahma Chaabouni, Paul Michel, Emmanuel Dupoux, Olivier Pietquin

VenueA*ACL
Year2024
ProceedingsACL (Findings)

Browse the full ACL paper archive.