Skip to content

Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles.

Bhrij Patel, Wesley A. Suttle, Alec Koppel, Vaneet Aggarwal, Brian M. Sadler, Dinesh Manocha, Amrit S. Bedi

VenueA*ICML
Year2024
ProceedingsICML

Browse the full ICML paper archive.