| 2025 | AAAI | Logarithmic Regret for Linear Markov Decision Processes with Adversarial Corruptions. | Canzhe Zhao, Xiangcheng Zhang, Baoxiang Wang, Shuai Li |
| 2025 | ICML | Learning Imperfect Information Extensive-form Games with Last-iterate Convergence under Bandit Feedback. | Canzhe Zhao, Yutian Cheng, Jing Dong, Baoxiang Wang, Shuai Li |
| 2025 | UAI | Towards Provably Efficient Learning of Imperfect Information Extensive-Form Games with Linear Function Approximation. | Canzhe Zhao, Shuze Chen, Weiming Liu, Haobo Fu, Qiang Fu, Shuai Li |
| 2023 | COLT | Best-of-three-worlds Analysis for Linear Bandits with Follow-the-regularized-leader Algorithm. | Fang Kong, Canzhe Zhao, Shuai Li |
| 2023 | ICLR | Learning Adversarial Linear Mixture Markov Decision Processes with Bandit Feedback and Unknown Transition. | Canzhe Zhao, Ruofeng Yang, Baoxiang Wang, Shuai Li |
| 2023 | IJCAI | DPMAC: Differentially Private Communication for Cooperative Multi-Agent Reinforcement Learning. | Canzhe Zhao, Yanjie Ze, Jing Dong, Baoxiang Wang, Shuai Li |
| 2023 | WSDM | Differentially Private Temporal Difference Learning with Stochastic Nonconvex-Strongly-Concave Optimization. | Canzhe Zhao, Yanjie Ze, Jing Dong, Baoxiang Wang, Shuai Li |
| 2022 | AAAI | Simultaneously Learning Stochastic and Adversarial Bandits under the Position-Based Model. | Cheng Chen, Canzhe Zhao, Shuai Li |
| 2022 | WWW | Knowledge-aware Conversational Preference Elicitation with Bandit Feedback. | Canzhe Zhao, Tong Yu, Zhihui Xie, Shuai Li |
| 2021 | CIKM | Clustering of Conversational Bandits for User Preference Learning and Elicitation. | Junda Wu, Canzhe Zhao, Tong Yu, Jingyang Li, Shuai Li |
| 2021 | SIGIR | Comparison-based Conversational Recommender System with Relative Bandit Feedback. | Zhihui Xie, Tong Yu, Canzhe Zhao, Shuai Li |