Smoothed Action Value Functions for Learning Gaussian Policies.
Ofir Nachum, Mohammad Norouzi, George Tucker, Dale Schuurmans
Browse the full ICML paper archive.
Ofir Nachum, Mohammad Norouzi, George Tucker, Dale Schuurmans
Browse the full ICML paper archive.