Yunhao Tang
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
37
Venues
6
Active years
2018–2025
Best venue rank
A*
Where they publish
Papers
37 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | AISTATS | A Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning. | Khimya Khetarpal, Zhaohan Daniel Guo, Bernardo vila Pires, Yunhao Tang, Clare Lyle, Mark Rowland, Nicolas Heess, Diana L. Borsa, Arthur Guez, Will Dabney |
| 2025 | ICML | Categorical Distributional Reinforcement Learning with Kullback-Leibler Divergence: Convergence and Asymptotics. | Tyler Kastner, Mark Rowland, Yunhao Tang, Murat A. Erdogdu, Amir-massoud Farahmand |
| 2025 | ICML | Optimizing Language Models for Inference Time Objectives using Reinforcement Learning. | Yunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rmi Munos |
| 2024 | AAAI | Learning Uncertainty-Aware Temporally-Extended Actions. | Joongkyu Lee, Seung Joon Park, Yunhao Tang, Min-hwan Oh |
| 2024 | ICML | Human Alignment of Large Language Models through Online Preference Optimisation. | Daniele Calandriello, Zhaohan Daniel Guo, Rmi Munos, Mark Rowland, Yunhao Tang, Bernardo vila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot |
| 2024 | ICML | Nash Learning from Human Feedback. | Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot |
| 2024 | ICML | Generalized Preference Optimization: A Unified Approach to Offline Alignment. | Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rmi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo vila Pires, Bilal Piot |
| 2024 | ICML | A Distributional Analogue to the Successor Representation. | Harley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang, Andr Barreto, Will Dabney, Marc G. Bellemare, Mark Rowland |
| 2023 | ICML | Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition. | Yash Chandak, Shantanu Thakoor, Zhaohan Daniel Guo, Yunhao Tang, Rmi Munos, Will Dabney, Diana L. Borsa |
| 2023 | ICML | Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice. | Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Mnard, Mohammad Gheshlaghi Azar, Rmi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvri, Wataru Kumagai, Yutaka Matsuo |
| 2023 | ICML | Quantile Credit Assignment. | Thomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang, Mark Rowland, Theophane Weber, Clare Lyle, Audrunas Gruslys, Michal Valko, Will Dabney, Georg Ostrovski, Eric Moulines, Rmi Munos |
| 2023 | ICML | The Edge of Orthogonality: A Simple View of What Makes BYOL Tick. | Pierre Harvey Richemond, Allison C. Tam, Yunhao Tang, Florian Strub, Bilal Piot, Felix Hill |
| 2023 | ICML | The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation. | Mark Rowland, Yunhao Tang, Clare Lyle, Rmi Munos, Marc G. Bellemare, Will Dabney |
| 2023 | ICML | Understanding Self-Predictive Learning for Reinforcement Learning. | Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko |
| 2023 | ICML | DoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm. | Yunhao Tang, Tadashi Kozuno, Mark Rowland, Anna Harutyunyan, Rmi Munos, Bernardo vila Pires, Michal Valko |
| 2023 | ICML | Towards a better understanding of representation dynamics under TD-learning. | Yunhao Tang, Rmi Munos |
| 2023 | ICML | VA-learning as a more efficient alternative to Q-learning. | Yunhao Tang, Rmi Munos, Mark Rowland, Michal Valko |
| 2023 | ICML | Fast Rates for Maximum Entropy Exploration. | Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines, Rmi Munos, Alexey Naumov, Pierre Perrault, Yunhao Tang, Michal Valko, Pierre Mnard |
| 2022 | AISTATS | Marginalized Operators for Off-policy Reinforcement Learning. | Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko |
| 2022 | ICML | Biased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement Learning. | Yunhao Tang |
| 2022 | ICML | From Dirichlet to Rubin: Optimistic Exploration in RL without Bonuses. | Daniil Tiapkin, Denis Belomestny, Eric Moulines, Alexey Naumov, Sergey Samsonov, Yunhao Tang, Michal Valko, Pierre Mnard |
| 2021 | AISTATS | Hindsight Expectation Maximization for Goal-conditioned Reinforcement Learning. | Yunhao Tang, Alp Kucukelbir |
| 2021 | ICML | Revisiting Peng's Q(λ) for Modern Reinforcement Learning. | Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rmi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel |
| 2021 | ICML | Taylor Expansion of Discount Factors. | Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko |
| 2020 | AAAI | Discretizing Continuous Action Space for On-Policy Optimization. | Yunhao Tang, Shipra Agrawal |
| 2020 | AISTATS | Practical Nonisotropic Monte Carlo Sampling in High Dimensions via Determinantal Point Processes. | Krzysztof Choromanski, Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang |
| 2020 | AISTATS | Variance Reduction for Evolution Strategies via Structured Control Variates. | Yunhao Tang, Krzysztof Choromanski, Alp Kucukelbir |
| 2020 | AISTATS | Discrete Action On-Policy Learning with Action-Value Critic. | Yuguang Yue, Yunhao Tang, Mingzhang Yin, Mingyuan Zhou |
| 2020 | ICLR | ES-MAML: Simple Hessian-Free Meta Learning. | Xingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski, Aldo Pacchiano, Yunhao Tang |
| 2020 | ICML | Monte-Carlo Tree Search as Regularized Policy Optimization. | Jean-Bastien Grill, Florent Altch, Yunhao Tang, Thomas Hubert, Michal Valko, Ioannis Antonoglou, Rmi Munos |
| 2020 | ICML | Learning to Score Behaviors for Guided Policy Optimization. | Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski, Anna Choromanska, Michael I. Jordan |
| 2020 | ICML | Reinforcement Learning for Integer Programming: Learning to Cut. | Yunhao Tang, Shipra Agrawal, Yuri Faenza |
| 2020 | ICML | Taylor Expansion Policy Optimization. | Yunhao Tang, Michal Valko, Rmi Munos |
| 2019 | AISTATS | KAMA-NNs: Low-dimensional Rotation Based Neural Networks. | Krzysztof Choromanski, Aldo Pacchiano, Jeffrey Pennington, Yunhao Tang |
| 2019 | AISTATS | Orthogonal Estimation of Wasserstein Distances. | Mark Rowland, Jiri Hron, Yunhao Tang, Krzysztof Choromanski, Tams Sarls, Adrian Weller |
| 2019 | CoRL | Provably Robust Blackbox Optimization for Reinforcement Learning. | Krzysztof Choromanski, Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Deepali Jain, Yuxiang Yang, Atil Iscen, Jasmine Hsu, Vikas Sindhwani |
| 2018 | IJCAI | Exploration by Distributional Reinforcement Learning. | Yunhao Tang, Shipra Agrawal |