FAST-Q: Fast-track Exploration with Adversarially Balanced State Representations for Counterfactual Action Estimation in Offline Reinforcement Learning.
Pulkit Agrawal, Rukma Talwadker, Aditya Pareek, Tridib Mukherjee
Browse the full WWW paper archive.