Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods.
Xin Guo, Anran Hu, Junzi Zhang
Browse the full AAAI paper archive.
Xin Guo, Anran Hu, Junzi Zhang
Browse the full AAAI paper archive.