Skip to content

Yunhao Tang

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

37

Venues

6

Active years

2018–2025

Best venue rank

A*

Where they publish

Papers

37 indexed papers, newest first.

YearVenueTitleAuthors
2025AISTATSA Unifying Framework for Action-Conditional Self-Predictive Reinforcement Learning.Khimya Khetarpal, Zhaohan Daniel Guo, Bernardo vila Pires, Yunhao Tang, Clare Lyle, Mark Rowland, Nicolas Heess, Diana L. Borsa, Arthur Guez, Will Dabney
2025ICMLCategorical Distributional Reinforcement Learning with Kullback-Leibler Divergence: Convergence and Asymptotics.Tyler Kastner, Mark Rowland, Yunhao Tang, Murat A. Erdogdu, Amir-massoud Farahmand
2025ICMLOptimizing Language Models for Inference Time Objectives using Reinforcement Learning.Yunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rmi Munos
2024AAAILearning Uncertainty-Aware Temporally-Extended Actions.Joongkyu Lee, Seung Joon Park, Yunhao Tang, Min-hwan Oh
2024ICMLHuman Alignment of Large Language Models through Online Preference Optimisation.Daniele Calandriello, Zhaohan Daniel Guo, Rmi Munos, Mark Rowland, Yunhao Tang, Bernardo vila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot
2024ICMLNash Learning from Human Feedback.Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot
2024ICMLGeneralized Preference Optimization: A Unified Approach to Offline Alignment.Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rmi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo vila Pires, Bilal Piot
2024ICMLA Distributional Analogue to the Successor Representation.Harley Wiltzer, Jesse Farebrother, Arthur Gretton, Yunhao Tang, Andr Barreto, Will Dabney, Marc G. Bellemare, Mark Rowland
2023ICMLRepresentations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition.Yash Chandak, Shantanu Thakoor, Zhaohan Daniel Guo, Yunhao Tang, Rmi Munos, Will Dabney, Diana L. Borsa
2023ICMLRegularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice.Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Mnard, Mohammad Gheshlaghi Azar, Rmi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvri, Wataru Kumagai, Yutaka Matsuo
2023ICMLQuantile Credit Assignment.Thomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang, Mark Rowland, Theophane Weber, Clare Lyle, Audrunas Gruslys, Michal Valko, Will Dabney, Georg Ostrovski, Eric Moulines, Rmi Munos
2023ICMLThe Edge of Orthogonality: A Simple View of What Makes BYOL Tick.Pierre Harvey Richemond, Allison C. Tam, Yunhao Tang, Florian Strub, Bilal Piot, Felix Hill
2023ICMLThe Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation.Mark Rowland, Yunhao Tang, Clare Lyle, Rmi Munos, Marc G. Bellemare, Will Dabney
2023ICMLUnderstanding Self-Predictive Learning for Reinforcement Learning.Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko
2023ICMLDoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm.Yunhao Tang, Tadashi Kozuno, Mark Rowland, Anna Harutyunyan, Rmi Munos, Bernardo vila Pires, Michal Valko
2023ICMLTowards a better understanding of representation dynamics under TD-learning.Yunhao Tang, Rmi Munos
2023ICMLVA-learning as a more efficient alternative to Q-learning.Yunhao Tang, Rmi Munos, Mark Rowland, Michal Valko
2023ICMLFast Rates for Maximum Entropy Exploration.Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines, Rmi Munos, Alexey Naumov, Pierre Perrault, Yunhao Tang, Michal Valko, Pierre Mnard
2022AISTATSMarginalized Operators for Off-policy Reinforcement Learning.Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko
2022ICMLBiased Gradient Estimate with Drastic Variance Reduction for Meta Reinforcement Learning.Yunhao Tang
2022ICMLFrom Dirichlet to Rubin: Optimistic Exploration in RL without Bonuses.Daniil Tiapkin, Denis Belomestny, Eric Moulines, Alexey Naumov, Sergey Samsonov, Yunhao Tang, Michal Valko, Pierre Mnard
2021AISTATSHindsight Expectation Maximization for Goal-conditioned Reinforcement Learning.Yunhao Tang, Alp Kucukelbir
2021ICMLRevisiting Peng's Q(λ) for Modern Reinforcement Learning.Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rmi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel
2021ICMLTaylor Expansion of Discount Factors.Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko
2020AAAIDiscretizing Continuous Action Space for On-Policy Optimization.Yunhao Tang, Shipra Agrawal
2020AISTATSPractical Nonisotropic Monte Carlo Sampling in High Dimensions via Determinantal Point Processes.Krzysztof Choromanski, Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang
2020AISTATSVariance Reduction for Evolution Strategies via Structured Control Variates.Yunhao Tang, Krzysztof Choromanski, Alp Kucukelbir
2020AISTATSDiscrete Action On-Policy Learning with Action-Value Critic.Yuguang Yue, Yunhao Tang, Mingzhang Yin, Mingyuan Zhou
2020ICLRES-MAML: Simple Hessian-Free Meta Learning.Xingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski, Aldo Pacchiano, Yunhao Tang
2020ICMLMonte-Carlo Tree Search as Regularized Policy Optimization.Jean-Bastien Grill, Florent Altch, Yunhao Tang, Thomas Hubert, Michal Valko, Ioannis Antonoglou, Rmi Munos
2020ICMLLearning to Score Behaviors for Guided Policy Optimization.Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Krzysztof Choromanski, Anna Choromanska, Michael I. Jordan
2020ICMLReinforcement Learning for Integer Programming: Learning to Cut.Yunhao Tang, Shipra Agrawal, Yuri Faenza
2020ICMLTaylor Expansion Policy Optimization.Yunhao Tang, Michal Valko, Rmi Munos
2019AISTATSKAMA-NNs: Low-dimensional Rotation Based Neural Networks.Krzysztof Choromanski, Aldo Pacchiano, Jeffrey Pennington, Yunhao Tang
2019AISTATSOrthogonal Estimation of Wasserstein Distances.Mark Rowland, Jiri Hron, Yunhao Tang, Krzysztof Choromanski, Tams Sarls, Adrian Weller
2019CoRLProvably Robust Blackbox Optimization for Reinforcement Learning.Krzysztof Choromanski, Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Deepali Jain, Yuxiang Yang, Atil Iscen, Jasmine Hsu, Vikas Sindhwani
2018IJCAIExploration by Distributional Reinforcement Learning.Yunhao Tang, Shipra Agrawal