Skip to content

Rmi Munos

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

94

Venues

13

Active years

1996–2025

Best venue rank

A*

Where they publish

Papers

94 indexed papers, newest first.

YearVenueTitleAuthors
2025ICMLTemporal Difference Flows.Jesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Rmi Munos, Alessandro Lazaric, Ahmed Touati
2025ICMLOptimizing Language Models for Inference Time Objectives using Reinforcement Learning.Yunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rmi Munos
2024AISTATSA General Theoretical Paradigm to Understand Learning from Human Preferences.Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Rmi Munos, Mark Rowland, Michal Valko, Daniele Calandriello
2024ICMLHuman Alignment of Large Language Models through Online Preference Optimisation.Daniele Calandriello, Zhaohan Daniel Guo, Rmi Munos, Mark Rowland, Yunhao Tang, Bernardo vila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot
2024ICMLNash Learning from Human Feedback.Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot
2024ICMLGeneralized Preference Optimization: A Unified Approach to Offline Alignment.Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rmi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo vila Pires, Bilal Piot
2023ICMLRepresentations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition.Yash Chandak, Shantanu Thakoor, Zhaohan Daniel Guo, Yunhao Tang, Rmi Munos, Will Dabney, Diana L. Borsa
2023ICMLAdapting to game trees in zero-sum imperfect information games.Cme Fiegel, Pierre Mnard, Tadashi Kozuno, Rmi Munos, Vianney Perchet, Michal Valko
2023ICMLCuriosity in Hindsight: Intrinsic Exploration in Stochastic Environments.Daniel Jarrett, Corentin Tallec, Florent Altch, Thomas Mesnard, Rmi Munos, Michal Valko
2023ICMLRegularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice.Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Mnard, Mohammad Gheshlaghi Azar, Rmi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvri, Wataru Kumagai, Yutaka Matsuo
2023ICMLQuantile Credit Assignment.Thomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang, Mark Rowland, Theophane Weber, Clare Lyle, Audrunas Gruslys, Michal Valko, Will Dabney, Georg Ostrovski, Eric Moulines, Rmi Munos
2023ICMLThe Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation.Mark Rowland, Yunhao Tang, Clare Lyle, Rmi Munos, Marc G. Bellemare, Will Dabney
2023ICMLUnderstanding Self-Predictive Learning for Reinforcement Learning.Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko
2023ICMLDoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm.Yunhao Tang, Tadashi Kozuno, Mark Rowland, Anna Harutyunyan, Rmi Munos, Bernardo vila Pires, Michal Valko
2023ICMLTowards a better understanding of representation dynamics under TD-learning.Yunhao Tang, Rmi Munos
2023ICMLVA-learning as a more efficient alternative to Q-learning.Yunhao Tang, Rmi Munos, Mark Rowland, Michal Valko
2023ICMLFast Rates for Maximum Entropy Exploration.Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines, Rmi Munos, Alexey Naumov, Pierre Perrault, Yunhao Tang, Michal Valko, Pierre Mnard
2022AISTATSMarginalized Operators for Off-policy Reinforcement Learning.Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko
2022ICLRLarge-Scale Representation Learning on Graphs via Bootstrapping.Shantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou, Eva L. Dyer, Rmi Munos, Petar Velickovic, Michal Valko
2022ICMLGeneralised Policy Improvement with Geometric Policy Composition.Shantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney, Rmi Munos, Andr Barreto
2021ICMLRevisiting Peng's Q(λ) for Modern Reinforcement Learning.Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rmi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel
2021ICMLCounterfactual Credit Assignment in Model-Free Reinforcement Learning.Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S. Stepleton, Nicolas Heess, Arthur Guez, Eric Moulines, Marcus Hutter, Lars Buesing, Rmi Munos
2021ICMLFrom Poincar Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization.Julien Prolat, Rmi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro A. Ortega, Neil Burch, Thomas W. Anthony, David Balduzzi, Bart De Vylder, Georgios Piliouras, Marc Lanctot, Karl Tuyls
2021ICMLTaylor Expansion of Discount Factors.Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko
2020AISTATSAdaptive Trade-Offs in Off-Policy Learning.Mark Rowland, Will Dabney, Rmi Munos
2020AISTATSConditional Importance Sampling for Off-Policy Learning.Mark Rowland, Anna Harutyunyan, Hado van Hasselt, Diana Borsa, Tom Schaul, Rmi Munos, Will Dabney
2020ICLRA Generalized Training Approach for Multiagent Learning.Paul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls, Julien Prolat, Siqi Liu, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes, Zhe Wang, Guy Lever, Nicolas Heess, Thore Graepel, Rmi Munos
2020ICMLMonte-Carlo Tree Search as Regularized Policy Optimization.Jean-Bastien Grill, Florent Altch, Yunhao Tang, Thomas Hubert, Michal Valko, Ioannis Antonoglou, Rmi Munos
2020ICMLBootstrap Latent-Predictive Representations for Multitask Reinforcement Learning.Zhaohan Daniel Guo, Bernardo vila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altch, Rmi Munos, Mohammad Gheshlaghi Azar
2020ICMLFast computation of Nash Equilibria in Imperfect Information Games.Rmi Munos, Julien Prolat, Jean-Baptiste Lespiau, Mark Rowland, Bart De Vylder, Marc Lanctot, Finbarr Timbers, Daniel Hennes, Shayegan Omidshafiei, Audrunas Gruslys, Mohammad Gheshlaghi Azar, Edward Lockhart, Karl Tuyls
2020ICMLTaylor Expansion Policy Optimization.Yunhao Tang, Michal Valko, Rmi Munos
2019AISTATSThe Termination Critic.Anna Harutyunyan, Will Dabney, Diana Borsa, Nicolas Heess, Rmi Munos, Doina Precup
2019ICLRUniversal Successor Features Approximators.Diana Borsa, Andr Barreto, John Quan, Daniel J. Mankowitz, Hado van Hasselt, Rmi Munos, David Silver, Tom Schaul
2019ICLRRecurrent Experience Replay in Distributed Reinforcement Learning.Steven Kapturowski, Georg Ostrovski, John Quan, Rmi Munos, Will Dabney
2019ICMLStatistics and Samples in Distributional Reinforcement Learning.Mark Rowland, Robert Dadashi, Saurabh Kumar, Rmi Munos, Marc G. Bellemare, Will Dabney
2018AAAIDistributional Reinforcement Learning With Quantile Regression.Will Dabney, Mark Rowland, Marc G. Bellemare, Rmi Munos
2018AISTATSAn Analysis of Categorical Distributional Reinforcement Learning.Mark Rowland, Marc G. Bellemare, Will Dabney, Rmi Munos, Yee Whye Teh
2018ICLRMaximum a Posteriori Policy Optimisation.Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rmi Munos, Nicolas Heess, Martin A. Riedmiller
2018ICLRNoisy Networks For Exploration.Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rmi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg
2018ICLRThe Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning.Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc G. Bellemare, Rmi Munos
2018ICMLTransfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement.Andr Barreto, Diana Borsa, John Quan, Tom Schaul, David Silver, Matteo Hessel, Daniel J. Mankowitz, Augustin Zdek, Rmi Munos
2018ICMLImplicit Quantile Networks for Distributional Reinforcement Learning.Will Dabney, Georg Ostrovski, David Silver, Rmi Munos
2018ICMLIMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.Lasse Espeholt, Hubert Soyer, Rmi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, Koray Kavukcuoglu
2018ICMLLearning to Search with MCTSnets.Arthur Guez, Theophane Weber, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals, Daan Wierstra, Rmi Munos, David Silver
2018ICMLThe Uncertainty Bellman Equation and Exploration.Brendan O'Donoghue, Ian Osband, Rmi Munos, Volodymyr Mnih
2018ICMLAutoregressive Quantile Networks for Generative Modeling.Georg Ostrovski, Will Dabney, Rmi Munos
2017CogSciLearning to reinforcement learn.Jane Wang, Zeb Kurth-Nelson, Hubert Soyer, Joel Z. Leibo, Dhruva Tirumala, Rmi Munos, Charles Blundell, Dharshan Kumaran, Matt M. Botvinick
2017ICLRSample Efficient Actor-Critic with Experience Replay.Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Rmi Munos, Koray Kavukcuoglu, Nando de Freitas
2017ICLRCombining policy gradient and Q-learning.Brendan O'Donoghue, Rmi Munos, Koray Kavukcuoglu, Volodymyr Mnih
2017ICMLMinimax Regret Bounds for Reinforcement Learning.Mohammad Gheshlaghi Azar, Ian Osband, Rmi Munos
2017ICMLA Distributional Perspective on Reinforcement Learning.Marc G. Bellemare, Will Dabney, Rmi Munos
2017ICMLAutomated Curriculum Learning for Neural Networks.Alex Graves, Marc G. Bellemare, Jacob Menick, Rmi Munos, Koray Kavukcuoglu
2017ICMLCount-Based Exploration with Neural Density Models.Georg Ostrovski, Marc G. Bellemare, Aron van den Oord, Rmi Munos
2016AAAIIncreasing the Action Gap: New Operators for Reinforcement Learning.Marc G. Bellemare, Georg Ostrovski, Arthur Guez, Philip S. Thomas, Rmi Munos
2016AAAIGeneralized Emphatic Temporal Difference Learning: Bias-Variance Analysis.Assaf Hallak, Aviv Tamar, Rmi Munos, Shie Mannor
2016ALTQ(λ) with Off-Policy Corrections.Anna Harutyunyan, Marc G. Bellemare, Tom Stepleton, Rmi Munos
2015AAAIFast Gradient Descent for Drifting Least Squares Regression, with Application to Bandits.Nathaniel Korda, Prashanth L. A., Rmi Munos
2015AISTATSToward Minimax Off-policy Value Estimation.Lihong Li, Rmi Munos, Csaba Szepesvri
2015ICMLCheap Bandits.Manjesh Kumar Hanawal, Venkatesh Saligrama, Michal Valko, Rmi Munos
2014AAAISpectral Thompson Sampling.Toms Kock, Michal Valko, Rmi Munos, Shipra Agrawal
2014CECBandits attack function optimization.Philippe Preux, Rmi Munos, Michal Valko
2014ICMLSpectral Bandits for Smooth Graph Functions.Michal Valko, Rmi Munos, Branislav Kveton, Toms Kock
2014ICMLRelative Upper Confidence Bound for the K-Armed Dueling Bandit Problem.Masrour Zoghi, Shimon Whiteson, Rmi Munos, Maarten de Rijke
2014WSDMRelative confidence sampling for efficient on-line ranker evaluation.Masrour Zoghi, Shimon Whiteson, Maarten de Rijke, Rmi Munos
2013ALTEditors' Introduction.Sanjay Jain, Rmi Munos, Frank Stephan, Thomas Zeugmann
2013ICMLToward Optimal Stratification for Stratified Monte-Carlo Integration.Alexandra Carpentier, Rmi Munos
2013ICMLStochastic Simultaneous Optimistic Optimization.Michal Valko, Alexandra Carpentier, Rmi Munos
2013UAIFinite-Time Analysis of Kernelised Contextual Bandits.Michal Valko, Nathaniel Korda, Rmi Munos, Ilias N. Flaounas, Nello Cristianini
2012ALTMinimax Number of Strata for Online Stratified Sampling Given Noisy Samples.Alexandra Carpentier, Rmi Munos
2012ALTThompson Sampling: An Asymptotically Optimal Finite-Time Analysis.Emilie Kaufmann, Nathaniel Korda, Rmi Munos
2012ALTRegret Bounds for Restless Markov Bandits.Ronald Ortner, Daniil Ryabko, Peter Auer, Rmi Munos
2012ICMLOn the Sample Complexity of Reinforcement Learning with a Generative Model .Mohammad Gheshlaghi Azar, Rmi Munos, Bert Kappen
2011ALTUpper-Confidence-Bound Algorithms for Active Learning in Multi-armed Bandits.Alexandra Carpentier, Alessandro Lazaric, Mohammad Ghavamzadeh, Rmi Munos, Peter Auer
2011ICMLFinite-Sample Analysis of Lasso-TD.Mohammad Ghavamzadeh, Alessandro Lazaric, Rmi Munos, Matthew W. Hoffman
2010COLTBest Arm Identification in Multi-Armed Bandits.Jean-Yves Audibert, Sbastien Bubeck, Rmi Munos
2010COLTOpen Loop Optimistic Planning.Sbastien Bubeck, Rmi Munos
2010ICMLAnalysis of a Classification-based Policy Iteration Algorithm.Alessandro Lazaric, Mohammad Ghavamzadeh, Rmi Munos
2010ICMLFinite-Sample Analysis of LSTD.Alessandro Lazaric, Mohammad Ghavamzadeh, Rmi Munos
2009ALTPure Exploration in Multi-armed Bandits Problems.Sbastien Bubeck, Rmi Munos, Gilles Stoltz
2009COLTHybrid Stochastic-Adversarial On-line Learning.Alessandro Lazaric, Rmi Munos
2009ICMLWorkshop summary: On-line learning with limited feedback.Jean-Yves Audibert, Peter Auer, Alessandro Lazaric, Rmi Munos, Daniil Ryabko, Csaba Szepesvri
2008ECAIAdaptive play in Texas Hold'em Poker.Raphal Matrepierre, Jrmie Mary, Rmi Munos
2007ALTTuning Bandit Algorithms in Stochastic Environments.Jean-Yves Audibert, Rmi Munos, Csaba Szepesvri
2007UAIBandit Algorithms for Tree Search.Pierre-Arnaud Coquelin, Rmi Munos
2006COLTLearning Near-Optimal Policies with Bellman-Residual Minimization Based Fitted Policy Iteration and a Single Sample Path.Andrs Antos, Csaba Szepesvri, Rmi Munos
2005AAAIError Bounds for Approximate Value Iteration.Rmi Munos
2005AAAIGeometric Variance Reduction in Markov Chains. Application to Value Function and Gradient Estimation.Rmi Munos
2005ICMLFinite time bounds for sampling based fitted value iteration.Csaba Szepesvri, Rmi Munos
2003ICMLError Bounds for Approximate Policy Iteration.Rmi Munos
2000ICMLRates of Convergence for Variable Resolution Schemes in Optimal Control.Rmi Munos, Andrew W. Moore
1999IJCAIVariable Resolution Discretization for High-Accuracy Solutions of Optimal Control Problems.Rmi Munos, Andrew W. Moore
1999IJCNNGradient descent approaches to neural-net-based solutions of the Hamilton-Jacobi-Bellman equation.Rmi Munos, Leemon C. Baird III, Andrew W. Moore
1997IJCAIA Convergent Reinforcement Learning Algorithm in the Continuous Case Based on a Finite Difference Method.Rmi Munos
1996ICMLA Convergent Reinforcement Learning Algorithm in the Continuous Case: The Finite-Element Reinforcement Learning.Rmi Munos