Can You Rely on Synthetic Labellers in Preference-Based Reinforcement Learning? It's Complicated.
Katherine Metcalf, Miguel Sarabia, Masha Fedzechkina, Barry-John Theobald
Browse the full AAAI paper archive.
Katherine Metcalf, Miguel Sarabia, Masha Fedzechkina, Barry-John Theobald
Browse the full AAAI paper archive.