A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes.
Chengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan Jiang
Browse the full ICML paper archive.
Chengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan Jiang
Browse the full ICML paper archive.