| 2020 | AISTATS | Momentum in Reinforcement Learning. | Nino Vieillard, Bruno Scherrer, Olivier Pietquin, Matthieu Geist |
| 2019 | AAAI | How to Combine Tree-Search Methods in Reinforcement Learning. | Yonathan Efroni, Gal Dalal, Bruno Scherrer, Shie Mannor |
| 2019 | ICML | A Theory of Regularized Markov Decision Processes. | Matthieu Geist, Bruno Scherrer, Olivier Pietquin |
| 2018 | ICML | Beyond the One-Step Greedy Approach in Reinforcement Learning. | Yonathan Efroni, Gal Dalal, Bruno Scherrer, Shie Mannor |
| 2016 | AISTATS | On the Use of Non-Stationary Strategies for Solving Two-Player Zero-Sum Markov Games. | Julien Prolat, Bilal Piot, Bruno Scherrer, Olivier Pietquin |
| 2016 | ICML | Softened Approximate Policy Iteration for Markov Games. | Julien Prolat, Bilal Piot, Matthieu Geist, Bruno Scherrer, Olivier Pietquin |
| 2015 | ICML | Non-Stationary Approximate Modified Policy Iteration. | Boris Lesner, Bruno Scherrer |
| 2015 | ICML | Approximate Dynamic Programming for Two-Player Zero-Sum Markov Games. | Julien Prolat, Bruno Scherrer, Bilal Piot, Olivier Pietquin |
| 2015 | ICML | On the Rate of Convergence and Error Bounds for LSTD(\(\lambda\)). | Manel Tagorti, Bruno Scherrer |
| 2014 | ICML | Approximate Policy Iteration Schemes: A Comparison. | Bruno Scherrer |
| 2012 | ICML | A Dantzig Selector Approach to Temporal Difference Learning. | Matthieu Geist, Bruno Scherrer, Alessandro Lazaric, Mohammad Ghavamzadeh |
| 2012 | ICML | Approximate Modified Policy Iteration. | Bruno Scherrer, Victor Gabillon, Mohammad Ghavamzadeh, Matthieu Geist |
| 2011 | ICML | Classification-based Policy Iteration with a Critic. | Victor Gabillon, Alessandro Lazaric, Mohammad Ghavamzadeh, Bruno Scherrer |
| 2010 | ICML | Should one compute the Temporal Difference fix point or minimize the Bellman Residual? The unified oblique projection view. | Bruno Scherrer |
| 2010 | ICML | Least-Squares Policy Iteration: Bias-Variance Trade-off in Control Problems. | Christophe Thiery, Bruno Scherrer |
| 2007 | CEC | Convergence and rate of convergence of a foraging ant model. | Amine M. Boumaza, Bruno Scherrer |
| 2007 | ICRA | Optimal control subsumes harmonic control. | Amine M. Boumaza, Bruno Scherrer |
| 2003 | ESANN | Parallel asynchronous distributed computations of optimal control in large state space Markov Decision processes. | Bruno Scherrer |
| 2003 | IJCAI | Modular self-organization for a long-living autonomous agent. | Bruno Scherrer |
| 2002 | ICTAI | Cooperative Co-Learning: A Model-Based Approach for Solving Multi Agent Reinforcement Problems. | Bruno Scherrer, Franois Charpillet |
| 2002 | SAC | A heuristic approach for solving decentralized-POMDP: assessment on the pursuit problem. | Iadine Chades, Bruno Scherrer, Franois Charpillet |