Skip to content

Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models.

Alessio Galatolo, Zhenbang Dai, Katie Winkle, Meriem Beloucif

VenueA*ACL
Year2025
ProceedingsACL (Findings)

Browse the full ACL paper archive.