Skip to content

Efficient Exploration in Average-Reward Constrained Reinforcement Learning: Achieving Near-Optimal Regret With Posterior Sampling.

Danil Provodin, Maurits Clemens Kaptein, Mykola Pechenizkiy

VenueA*ICML
Year2024
ProceedingsICML

Browse the full ICML paper archive.