Behaviour Policy Optimization: Provably Lower Variance Return Estimates for Off-Policy Reinforcement Learning.
Alexander W. Goodall, Edwin Hamel-De le Court, Francesco Belardinelli
Browse the full AAAI paper archive.
Alexander W. Goodall, Edwin Hamel-De le Court, Francesco Belardinelli
Browse the full AAAI paper archive.