Skip to content

Global optimality of softmax policy gradient with single hidden layer neural networks in the mean-field regime.

Andrea Agazzi, Jianfeng Lu

VenueA*ICLR
Year2021
ProceedingsICLR

Browse the full ICLR paper archive.