Efficient Exploration in Average-Reward Constrained Reinforcement Learning: Achieving Near-Optimal Regret With Posterior Sampling.
Danil Provodin, Maurits Clemens Kaptein, Mykola Pechenizkiy
Browse the full ICML paper archive.
Danil Provodin, Maurits Clemens Kaptein, Mykola Pechenizkiy
Browse the full ICML paper archive.