Skip to content

Mohammad Gheshlaghi Azar

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

16

Venues

6

Active years

2012–2025

Best venue rank

A*

Where they publish

Papers

16 indexed papers, newest first.

YearVenueTitleAuthors
2025ICLRSelf-Improving Robust Preference Optimization.Eugene Choi, Arash Ahmadian, Matthieu Geist, Olivier Pietquin, Mohammad Gheshlaghi Azar
2024AISTATSA General Theoretical Paradigm to Understand Learning from Human Preferences.Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Rmi Munos, Mark Rowland, Michal Valko, Daniele Calandriello
2024EMNLPContrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion.Yannis Flet-Berliac, Nathan Grinsztajn, Florian Strub, Eugene Choi, Bill Wu, Chris Cremer, Arash Ahmadian, Yash Chandak, Mohammad Gheshlaghi Azar, Olivier Pietquin, Matthieu Geist
2024ICMLNash Learning from Human Feedback.Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot
2023ICMLRegularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice.Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Mnard, Mohammad Gheshlaghi Azar, Rmi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvri, Wataru Kumagai, Yutaka Matsuo
2023ICMLUnderstanding Self-Predictive Learning for Reinforcement Learning.Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko
2022ICLRLarge-Scale Representation Learning on Graphs via Bootstrapping.Shantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou, Eva L. Dyer, Rmi Munos, Petar Velickovic, Michal Valko
2020ICMLBootstrap Latent-Predictive Representations for Multitask Reinforcement Learning.Zhaohan Daniel Guo, Bernardo vila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altch, Rmi Munos, Mohammad Gheshlaghi Azar
2020ICMLFast computation of Nash Equilibria in Imperfect Information Games.Rmi Munos, Julien Prolat, Jean-Baptiste Lespiau, Mark Rowland, Bart De Vylder, Marc Lanctot, Finbarr Timbers, Daniel Hennes, Shayegan Omidshafiei, Audrunas Gruslys, Mohammad Gheshlaghi Azar, Edward Lockhart, Karl Tuyls
2018AAAIRainbow: Combining Improvements in Deep Reinforcement Learning.Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, David Silver
2018ICLRNoisy Networks For Exploration.Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rmi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg
2018ICLRThe Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning.Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc G. Bellemare, Rmi Munos
2017ICMLMinimax Regret Bounds for Reinforcement Learning.Mohammad Gheshlaghi Azar, Ian Osband, Rmi Munos
2016UAIConvex Relaxation Regression: Black-Box Optimization of Smooth Functions by Learning Their Convex Envelopes.Mohammad Gheshlaghi Azar, Eva L. Dyer, Konrad P. Krding
2014ICMLOnline Stochastic Optimization under Correlated Bandit Feedback.Mohammad Gheshlaghi Azar, Alessandro Lazaric, Emma Brunskill
2012ICMLOn the Sample Complexity of Reinforcement Learning with a Generative Model .Mohammad Gheshlaghi Azar, Rmi Munos, Bert Kappen