Visualising Policy-Reward Interplay to Inform Zeroth-Order Preference Optimisation of Large Language Models.
Alessio Galatolo, Zhenbang Dai, Katie Winkle, Meriem Beloucif
Browse the full ACL paper archive.
Alessio Galatolo, Zhenbang Dai, Katie Winkle, Meriem Beloucif
Browse the full ACL paper archive.