Learning Near-Optimal Policies with Bellman-Residual Minimization Based Fitted Policy Iteration and a Single Sample Path.
Andrs Antos, Csaba Szepesvri, Rmi Munos
Browse the full COLT paper archive.
Andrs Antos, Csaba Szepesvri, Rmi Munos
Browse the full COLT paper archive.