Rmi Munos
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
94
Venues
13
Active years
1996–2025
Best venue rank
A*
Where they publish
Papers
94 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2025 | ICML | Temporal Difference Flows. | Jesse Farebrother, Matteo Pirotta, Andrea Tirinzoni, Rmi Munos, Alessandro Lazaric, Ahmed Touati |
| 2025 | ICML | Optimizing Language Models for Inference Time Objectives using Reinforcement Learning. | Yunhao Tang, Kunhao Zheng, Gabriel Synnaeve, Rmi Munos |
| 2024 | AISTATS | A General Theoretical Paradigm to Understand Learning from Human Preferences. | Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bilal Piot, Rmi Munos, Mark Rowland, Michal Valko, Daniele Calandriello |
| 2024 | ICML | Human Alignment of Large Language Models through Online Preference Optimisation. | Daniele Calandriello, Zhaohan Daniel Guo, Rmi Munos, Mark Rowland, Yunhao Tang, Bernardo vila Pires, Pierre Harvey Richemond, Charline Le Lan, Michal Valko, Tianqi Liu, Rishabh Joshi, Zeyu Zheng, Bilal Piot |
| 2024 | ICML | Nash Learning from Human Feedback. | Rmi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar, Mark Rowland, Daniel Guo, Yunhao Tang, Matthieu Geist, Thomas Mesnard, Cme Fiegel, Andrea Michi, Marco Selvi, Sertan Girgin, Nikola Momchev, Olivier Bachem, Daniel J. Mankowitz, Doina Precup, Bilal Piot |
| 2024 | ICML | Generalized Preference Optimization: A Unified Approach to Offline Alignment. | Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Rmi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo vila Pires, Bilal Piot |
| 2023 | ICML | Representations and Exploration for Deep Reinforcement Learning using Singular Value Decomposition. | Yash Chandak, Shantanu Thakoor, Zhaohan Daniel Guo, Yunhao Tang, Rmi Munos, Will Dabney, Diana L. Borsa |
| 2023 | ICML | Adapting to game trees in zero-sum imperfect information games. | Cme Fiegel, Pierre Mnard, Tadashi Kozuno, Rmi Munos, Vianney Perchet, Michal Valko |
| 2023 | ICML | Curiosity in Hindsight: Intrinsic Exploration in Stochastic Environments. | Daniel Jarrett, Corentin Tallec, Florent Altch, Thomas Mesnard, Rmi Munos, Michal Valko |
| 2023 | ICML | Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice. | Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Wenhao Yang, Jincheng Mei, Pierre Mnard, Mohammad Gheshlaghi Azar, Rmi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvri, Wataru Kumagai, Yutaka Matsuo |
| 2023 | ICML | Quantile Credit Assignment. | Thomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang, Mark Rowland, Theophane Weber, Clare Lyle, Audrunas Gruslys, Michal Valko, Will Dabney, Georg Ostrovski, Eric Moulines, Rmi Munos |
| 2023 | ICML | The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation. | Mark Rowland, Yunhao Tang, Clare Lyle, Rmi Munos, Marc G. Bellemare, Will Dabney |
| 2023 | ICML | Understanding Self-Predictive Learning for Reinforcement Learning. | Yunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo vila Pires, Yash Chandak, Rmi Munos, Mark Rowland, Mohammad Gheshlaghi Azar, Charline Le Lan, Clare Lyle, Andrs Gyrgy, Shantanu Thakoor, Will Dabney, Bilal Piot, Daniele Calandriello, Michal Valko |
| 2023 | ICML | DoMo-AC: Doubly Multi-step Off-policy Actor-Critic Algorithm. | Yunhao Tang, Tadashi Kozuno, Mark Rowland, Anna Harutyunyan, Rmi Munos, Bernardo vila Pires, Michal Valko |
| 2023 | ICML | Towards a better understanding of representation dynamics under TD-learning. | Yunhao Tang, Rmi Munos |
| 2023 | ICML | VA-learning as a more efficient alternative to Q-learning. | Yunhao Tang, Rmi Munos, Mark Rowland, Michal Valko |
| 2023 | ICML | Fast Rates for Maximum Entropy Exploration. | Daniil Tiapkin, Denis Belomestny, Daniele Calandriello, Eric Moulines, Rmi Munos, Alexey Naumov, Pierre Perrault, Yunhao Tang, Michal Valko, Pierre Mnard |
| 2022 | AISTATS | Marginalized Operators for Off-policy Reinforcement Learning. | Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko |
| 2022 | ICLR | Large-Scale Representation Learning on Graphs via Bootstrapping. | Shantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou, Eva L. Dyer, Rmi Munos, Petar Velickovic, Michal Valko |
| 2022 | ICML | Generalised Policy Improvement with Geometric Policy Composition. | Shantanu Thakoor, Mark Rowland, Diana Borsa, Will Dabney, Rmi Munos, Andr Barreto |
| 2021 | ICML | Revisiting Peng's Q(λ) for Modern Reinforcement Learning. | Tadashi Kozuno, Yunhao Tang, Mark Rowland, Rmi Munos, Steven Kapturowski, Will Dabney, Michal Valko, David Abel |
| 2021 | ICML | Counterfactual Credit Assignment in Model-Free Reinforcement Learning. | Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S. Stepleton, Nicolas Heess, Arthur Guez, Eric Moulines, Marcus Hutter, Lars Buesing, Rmi Munos |
| 2021 | ICML | From Poincar Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization. | Julien Prolat, Rmi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro A. Ortega, Neil Burch, Thomas W. Anthony, David Balduzzi, Bart De Vylder, Georgios Piliouras, Marc Lanctot, Karl Tuyls |
| 2021 | ICML | Taylor Expansion of Discount Factors. | Yunhao Tang, Mark Rowland, Rmi Munos, Michal Valko |
| 2020 | AISTATS | Adaptive Trade-Offs in Off-Policy Learning. | Mark Rowland, Will Dabney, Rmi Munos |
| 2020 | AISTATS | Conditional Importance Sampling for Off-Policy Learning. | Mark Rowland, Anna Harutyunyan, Hado van Hasselt, Diana Borsa, Tom Schaul, Rmi Munos, Will Dabney |
| 2020 | ICLR | A Generalized Training Approach for Multiagent Learning. | Paul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls, Julien Prolat, Siqi Liu, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes, Zhe Wang, Guy Lever, Nicolas Heess, Thore Graepel, Rmi Munos |
| 2020 | ICML | Monte-Carlo Tree Search as Regularized Policy Optimization. | Jean-Bastien Grill, Florent Altch, Yunhao Tang, Thomas Hubert, Michal Valko, Ioannis Antonoglou, Rmi Munos |
| 2020 | ICML | Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning. | Zhaohan Daniel Guo, Bernardo vila Pires, Bilal Piot, Jean-Bastien Grill, Florent Altch, Rmi Munos, Mohammad Gheshlaghi Azar |
| 2020 | ICML | Fast computation of Nash Equilibria in Imperfect Information Games. | Rmi Munos, Julien Prolat, Jean-Baptiste Lespiau, Mark Rowland, Bart De Vylder, Marc Lanctot, Finbarr Timbers, Daniel Hennes, Shayegan Omidshafiei, Audrunas Gruslys, Mohammad Gheshlaghi Azar, Edward Lockhart, Karl Tuyls |
| 2020 | ICML | Taylor Expansion Policy Optimization. | Yunhao Tang, Michal Valko, Rmi Munos |
| 2019 | AISTATS | The Termination Critic. | Anna Harutyunyan, Will Dabney, Diana Borsa, Nicolas Heess, Rmi Munos, Doina Precup |
| 2019 | ICLR | Universal Successor Features Approximators. | Diana Borsa, Andr Barreto, John Quan, Daniel J. Mankowitz, Hado van Hasselt, Rmi Munos, David Silver, Tom Schaul |
| 2019 | ICLR | Recurrent Experience Replay in Distributed Reinforcement Learning. | Steven Kapturowski, Georg Ostrovski, John Quan, Rmi Munos, Will Dabney |
| 2019 | ICML | Statistics and Samples in Distributional Reinforcement Learning. | Mark Rowland, Robert Dadashi, Saurabh Kumar, Rmi Munos, Marc G. Bellemare, Will Dabney |
| 2018 | AAAI | Distributional Reinforcement Learning With Quantile Regression. | Will Dabney, Mark Rowland, Marc G. Bellemare, Rmi Munos |
| 2018 | AISTATS | An Analysis of Categorical Distributional Reinforcement Learning. | Mark Rowland, Marc G. Bellemare, Will Dabney, Rmi Munos, Yee Whye Teh |
| 2018 | ICLR | Maximum a Posteriori Policy Optimisation. | Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rmi Munos, Nicolas Heess, Martin A. Riedmiller |
| 2018 | ICLR | Noisy Networks For Exploration. | Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Matteo Hessel, Ian Osband, Alex Graves, Volodymyr Mnih, Rmi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg |
| 2018 | ICLR | The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning. | Audrunas Gruslys, Will Dabney, Mohammad Gheshlaghi Azar, Bilal Piot, Marc G. Bellemare, Rmi Munos |
| 2018 | ICML | Transfer in Deep Reinforcement Learning Using Successor Features and Generalised Policy Improvement. | Andr Barreto, Diana Borsa, John Quan, Tom Schaul, David Silver, Matteo Hessel, Daniel J. Mankowitz, Augustin Zdek, Rmi Munos |
| 2018 | ICML | Implicit Quantile Networks for Distributional Reinforcement Learning. | Will Dabney, Georg Ostrovski, David Silver, Rmi Munos |
| 2018 | ICML | IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures. | Lasse Espeholt, Hubert Soyer, Rmi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, Koray Kavukcuoglu |
| 2018 | ICML | Learning to Search with MCTSnets. | Arthur Guez, Theophane Weber, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals, Daan Wierstra, Rmi Munos, David Silver |
| 2018 | ICML | The Uncertainty Bellman Equation and Exploration. | Brendan O'Donoghue, Ian Osband, Rmi Munos, Volodymyr Mnih |
| 2018 | ICML | Autoregressive Quantile Networks for Generative Modeling. | Georg Ostrovski, Will Dabney, Rmi Munos |
| 2017 | CogSci | Learning to reinforcement learn. | Jane Wang, Zeb Kurth-Nelson, Hubert Soyer, Joel Z. Leibo, Dhruva Tirumala, Rmi Munos, Charles Blundell, Dharshan Kumaran, Matt M. Botvinick |
| 2017 | ICLR | Sample Efficient Actor-Critic with Experience Replay. | Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Rmi Munos, Koray Kavukcuoglu, Nando de Freitas |
| 2017 | ICLR | Combining policy gradient and Q-learning. | Brendan O'Donoghue, Rmi Munos, Koray Kavukcuoglu, Volodymyr Mnih |
| 2017 | ICML | Minimax Regret Bounds for Reinforcement Learning. | Mohammad Gheshlaghi Azar, Ian Osband, Rmi Munos |
| 2017 | ICML | A Distributional Perspective on Reinforcement Learning. | Marc G. Bellemare, Will Dabney, Rmi Munos |
| 2017 | ICML | Automated Curriculum Learning for Neural Networks. | Alex Graves, Marc G. Bellemare, Jacob Menick, Rmi Munos, Koray Kavukcuoglu |
| 2017 | ICML | Count-Based Exploration with Neural Density Models. | Georg Ostrovski, Marc G. Bellemare, Aron van den Oord, Rmi Munos |
| 2016 | AAAI | Increasing the Action Gap: New Operators for Reinforcement Learning. | Marc G. Bellemare, Georg Ostrovski, Arthur Guez, Philip S. Thomas, Rmi Munos |
| 2016 | AAAI | Generalized Emphatic Temporal Difference Learning: Bias-Variance Analysis. | Assaf Hallak, Aviv Tamar, Rmi Munos, Shie Mannor |
| 2016 | ALT | Q(λ) with Off-Policy Corrections. | Anna Harutyunyan, Marc G. Bellemare, Tom Stepleton, Rmi Munos |
| 2015 | AAAI | Fast Gradient Descent for Drifting Least Squares Regression, with Application to Bandits. | Nathaniel Korda, Prashanth L. A., Rmi Munos |
| 2015 | AISTATS | Toward Minimax Off-policy Value Estimation. | Lihong Li, Rmi Munos, Csaba Szepesvri |
| 2015 | ICML | Cheap Bandits. | Manjesh Kumar Hanawal, Venkatesh Saligrama, Michal Valko, Rmi Munos |
| 2014 | AAAI | Spectral Thompson Sampling. | Toms Kock, Michal Valko, Rmi Munos, Shipra Agrawal |
| 2014 | CEC | Bandits attack function optimization. | Philippe Preux, Rmi Munos, Michal Valko |
| 2014 | ICML | Spectral Bandits for Smooth Graph Functions. | Michal Valko, Rmi Munos, Branislav Kveton, Toms Kock |
| 2014 | ICML | Relative Upper Confidence Bound for the K-Armed Dueling Bandit Problem. | Masrour Zoghi, Shimon Whiteson, Rmi Munos, Maarten de Rijke |
| 2014 | WSDM | Relative confidence sampling for efficient on-line ranker evaluation. | Masrour Zoghi, Shimon Whiteson, Maarten de Rijke, Rmi Munos |
| 2013 | ALT | Editors' Introduction. | Sanjay Jain, Rmi Munos, Frank Stephan, Thomas Zeugmann |
| 2013 | ICML | Toward Optimal Stratification for Stratified Monte-Carlo Integration. | Alexandra Carpentier, Rmi Munos |
| 2013 | ICML | Stochastic Simultaneous Optimistic Optimization. | Michal Valko, Alexandra Carpentier, Rmi Munos |
| 2013 | UAI | Finite-Time Analysis of Kernelised Contextual Bandits. | Michal Valko, Nathaniel Korda, Rmi Munos, Ilias N. Flaounas, Nello Cristianini |
| 2012 | ALT | Minimax Number of Strata for Online Stratified Sampling Given Noisy Samples. | Alexandra Carpentier, Rmi Munos |
| 2012 | ALT | Thompson Sampling: An Asymptotically Optimal Finite-Time Analysis. | Emilie Kaufmann, Nathaniel Korda, Rmi Munos |
| 2012 | ALT | Regret Bounds for Restless Markov Bandits. | Ronald Ortner, Daniil Ryabko, Peter Auer, Rmi Munos |
| 2012 | ICML | On the Sample Complexity of Reinforcement Learning with a Generative Model . | Mohammad Gheshlaghi Azar, Rmi Munos, Bert Kappen |
| 2011 | ALT | Upper-Confidence-Bound Algorithms for Active Learning in Multi-armed Bandits. | Alexandra Carpentier, Alessandro Lazaric, Mohammad Ghavamzadeh, Rmi Munos, Peter Auer |
| 2011 | ICML | Finite-Sample Analysis of Lasso-TD. | Mohammad Ghavamzadeh, Alessandro Lazaric, Rmi Munos, Matthew W. Hoffman |
| 2010 | COLT | Best Arm Identification in Multi-Armed Bandits. | Jean-Yves Audibert, Sbastien Bubeck, Rmi Munos |
| 2010 | COLT | Open Loop Optimistic Planning. | Sbastien Bubeck, Rmi Munos |
| 2010 | ICML | Analysis of a Classification-based Policy Iteration Algorithm. | Alessandro Lazaric, Mohammad Ghavamzadeh, Rmi Munos |
| 2010 | ICML | Finite-Sample Analysis of LSTD. | Alessandro Lazaric, Mohammad Ghavamzadeh, Rmi Munos |
| 2009 | ALT | Pure Exploration in Multi-armed Bandits Problems. | Sbastien Bubeck, Rmi Munos, Gilles Stoltz |
| 2009 | COLT | Hybrid Stochastic-Adversarial On-line Learning. | Alessandro Lazaric, Rmi Munos |
| 2009 | ICML | Workshop summary: On-line learning with limited feedback. | Jean-Yves Audibert, Peter Auer, Alessandro Lazaric, Rmi Munos, Daniil Ryabko, Csaba Szepesvri |
| 2008 | ECAI | Adaptive play in Texas Hold'em Poker. | Raphal Matrepierre, Jrmie Mary, Rmi Munos |
| 2007 | ALT | Tuning Bandit Algorithms in Stochastic Environments. | Jean-Yves Audibert, Rmi Munos, Csaba Szepesvri |
| 2007 | UAI | Bandit Algorithms for Tree Search. | Pierre-Arnaud Coquelin, Rmi Munos |
| 2006 | COLT | Learning Near-Optimal Policies with Bellman-Residual Minimization Based Fitted Policy Iteration and a Single Sample Path. | Andrs Antos, Csaba Szepesvri, Rmi Munos |
| 2005 | AAAI | Error Bounds for Approximate Value Iteration. | Rmi Munos |
| 2005 | AAAI | Geometric Variance Reduction in Markov Chains. Application to Value Function and Gradient Estimation. | Rmi Munos |
| 2005 | ICML | Finite time bounds for sampling based fitted value iteration. | Csaba Szepesvri, Rmi Munos |
| 2003 | ICML | Error Bounds for Approximate Policy Iteration. | Rmi Munos |
| 2000 | ICML | Rates of Convergence for Variable Resolution Schemes in Optimal Control. | Rmi Munos, Andrew W. Moore |
| 1999 | IJCAI | Variable Resolution Discretization for High-Accuracy Solutions of Optimal Control Problems. | Rmi Munos, Andrew W. Moore |
| 1999 | IJCNN | Gradient descent approaches to neural-net-based solutions of the Hamilton-Jacobi-Bellman equation. | Rmi Munos, Leemon C. Baird III, Andrew W. Moore |
| 1997 | IJCAI | A Convergent Reinforcement Learning Algorithm in the Continuous Case Based on a Finite Difference Method. | Rmi Munos |
| 1996 | ICML | A Convergent Reinforcement Learning Algorithm in the Continuous Case: The Finite-Element Reinforcement Learning. | Rmi Munos |