Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field Regime.
James-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz Szpruch
Browse the full ICML paper archive.
James-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz Szpruch
Browse the full ICML paper archive.