Mohammad Gheshlaghi Azar
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
16
Venues
6
Active years
2012–2025
Best venue rank
A*
Where they publish
Papers
16 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | ICLR | Self-Improving Robust Preference Optimization. | Eugene Choi, Arash Ahmadian, Matthieu Geist, Olivier Pietquin, Mohammad Gheshlaghi Azar |
| 2024 | AISTATS | A General Theoretical Paradigm to Understand Learning from Human Preferences. | Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Rmi Munos, Mark Rowland, Michal Valko, Daniele Calandriello |
| 2024 | EMNLP | Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion. | Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi, Bill Wu, Chris Cremer, Arash Ahmadian, Yash Chandak, Mohammad Gheshlaghi Azar, Olivier Pietquin, Matthieu Geist |
| 2024 | ICML | Nash Learning from Human Feedback. | Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot |
| 2023 | ICML | Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice. | Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Mnard, Mohammad Gheshlaghi Azar, Rmi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvri, Wataru Kumagai, Yutaka Matsuo |
| 2023 | ICML | Understanding Self-Predictive Learning for Reinforcement Learning. | Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko |
| 2022 | ICLR | Large-Scale Representation Learning on Graphs via Bootstrapping. | Shantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou, Eva L. Dyer, Rmi Munos, Petar Velickovic, Michal Valko |
| 2020 | ICML | Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning. | Zhaohan Daniel Guo, Bernardo vila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altch, Rmi Munos, Mohammad Gheshlaghi Azar |
| 2020 | ICML | Fast computation of Nash Equilibria in Imperfect Information Games. | Rmi Munos, Julien Prolat, Jean-Baptiste Lespiau, Mark Rowland, Bart De Vylder, Marc Lanctot, Finbarr Timbers, Daniel Hennes, Shayegan Omidshafiei, Audrunas Gruslys, Mohammad Gheshlaghi Azar, Edward Lockhart, Karl Tuyls |
| 2018 | AAAI | Rainbow: Combining Improvements in Deep Reinforcement Learning. | Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, David Silver |
| 2018 | ICLR | Noisy Networks For Exploration. | Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rmi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg |
| 2018 | ICLR | The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning. | Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc G. Bellemare, Rmi Munos |
| 2017 | ICML | Minimax Regret Bounds for Reinforcement Learning. | Mohammad Gheshlaghi Azar, Ian Osband, Rmi Munos |
| 2016 | UAI | Convex Relaxation Regression: Black-Box Optimization of Smooth Functions by Learning Their Convex Envelopes. | Mohammad Gheshlaghi Azar, Eva L. Dyer, Konrad P. Krding |
| 2014 | ICML | Online Stochastic Optimization under Correlated Bandit Feedback. | Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill |
| 2012 | ICML | On the Sample Complexity of Reinforcement Learning with a Generative Model . | Mohammad Gheshlaghi Azar, Rmi Munos, Bert Kappen |