REvolve: Reward Evolution with Large Language Models using Human Feedback.
Rishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi, Pedro Zuidberg Dos Martires
Browse the full ICLR paper archive.
Rishi Hazra, Alkis Sygkounas, Andreas Persson, Amy Loutfi, Pedro Zuidberg Dos Martires
Browse the full ICLR paper archive.