MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking.
Sebastian Farquhar, Vikrant Varma, David Lindner, David Elson, Caleb Biddulph, Ian Goodfellow, Rohin Shah
Browse the full ICML paper archive.
Sebastian Farquhar, Vikrant Varma, David Lindner, David Elson, Caleb Biddulph, Ian Goodfellow, Rohin Shah
Browse the full ICML paper archive.