A reinforcement learning framework based on regret minimization for approximating best response in fictitious self-play.
Yanran Xu, Kangxin He, Shu Hu, Hui Li
Browse the full HPCC paper archive.
Yanran Xu, Kangxin He, Shu Hu, Hui Li
Browse the full HPCC paper archive.