| 2025 | ICML | MetaOptimize: A Framework for Optimizing Step Sizes and Other Meta-parameters. | Arsalan Sharifnassab, Saber Salehkaleybar, Richard S. Sutton |
| 2024 | AAAI | Reward-Respecting Subtasks for Model-Based Reinforcement Learning (Abstract Reprint). | Richard S. Sutton, Marlos C. Machado, G. Zacharias Holland, David Szepesvari, Finbarr Timbers, Brian Tanner, Adam White |
| 2023 | ICML | Toward Efficient Gradient-Based Value Estimation. | Arsalan Sharifnassab, Richard S. Sutton |
| 2021 | ICML | Learning and Planning in Average-Reward Markov Decision Processes. | Yi Wan, Abhishek Naik, Richard S. Sutton |
| 2021 | ICML | Average-Reward Off-Policy Policy Evaluation with Function Approximation. | Shangtong Zhang, Yi Wan, Richard S. Sutton, Shimon Whiteson |
| 2020 | AAAI | Fixed-Horizon Temporal Difference Methods for Stable Reinforcement Learning. | Kristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton, Daniel Graves |
| 2020 | ICLR | Behaviour Suite for Reinforcement Learning. | Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvri, Satinder Singh, Benjamin Van Roy, Richard S. Sutton, David Silver, Hado van Hasselt |
| 2019 | IJCAI | Extending Sliding-Step Importance Weighting from Supervised Learning to Reinforcement Learning. | Tian Tian, Richard S. Sutton |
| 2019 | IJCAI | Planning with Expectation Models. | Yi Wan, Muhammad Zaheer, Adam White, Martha White, Richard S. Sutton |
| 2018 | AAAI | Multi-Step Reinforcement Learning: A Unifying Algorithm. | Kristopher De Asis, J. Fernando Hernandez-Garcia, G. Zacharias Holland, Richard S. Sutton |
| 2018 | UAI | Per-decision Multi-step Temporal Difference Learning with Control Variates. | Kristopher De Asis, Richard S. Sutton |
| 2018 | UAI | Comparing Direct and Indirect Temporal-Difference Methods for Estimating the Variance of the Return. | Craig Sherstan, Dylan R. Ashley, Brendan Bennett, Kenny Young, Adam White, Martha White, Richard S. Sutton |
| 2017 | AI | On Generalized Bellman Equations and Temporal-Difference Learning. | Huizhen Yu, Ashique Rupam Mahmood, Richard S. Sutton |
| 2015 | ICML | A Deeper Look at Planning as Learning from Replay. | Harm Vanseijen, Richard S. Sutton |
| 2015 | UAI | Off-policy learning based on weighted importance sampling with linear computational complexity. | Ashique Rupam Mahmood, Richard S. Sutton |
| 2014 | ICML | True Online TD(lambda). | Harm van Seijen, Richard S. Sutton |
| 2014 | ICML | A new Q(lambda) with interim forward view and Monte Carlo equivalence. | Richard S. Sutton, Ashique Rupam Mahmood, Doina Precup, Hado van Hasselt |
| 2014 | UAI | Off-policy TD( l) with a true online equivalence. | Hado van Hasselt, Ashique Rupam Mahmood, Richard S. Sutton |
| 2013 | AAAI | Representation Search through Generate and Test. | Ashique Rupam Mahmood, Richard S. Sutton |
| 2013 | ICML | Planning by Prioritized Sweeping with Small Backups. | Harm van Seijen, Richard S. Sutton |
| 2012 | ICASSP | Tuning-free step-size adaptation. | Ashique Rupam Mahmood, Richard S. Sutton, Thomas Degris, Patrick M. Pilarski |
| 2012 | ICML | Linear Off-Policy Actor-Critic. | Thomas Degris, Martha White, Richard S. Sutton |
| 2012 | SMC | Acquiring a broad range of empirical knowledge in real time by temporal-difference learning. | Joseph Modayil, Adam White, Patrick M. Pilarski, Richard S. Sutton |
| 2011 | ILP | Beyond Reward: The Problem of Knowledge and Data. | Richard S. Sutton |
| 2010 | ICML | Toward Off-Policy Learning Control with Function Approximation. | Hamid Reza Maei, Csaba Szepesvri, Shalabh Bhatnagar, Richard S. Sutton |
| 2009 | ICML | Fast gradient-descent methods for temporal-difference learning with linear function approximation. | Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvri, Eric Wiewiora |
| 2008 | ICML | Sample-based learning and search with permanent and transient memories. | David Silver, Richard S. Sutton, Martin Mller |
| 2008 | UAI | Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping. | Richard S. Sutton, Csaba Szepesvri, Alborz Geramifard, Michael H. Bowling |
| 2007 | ICML | On the role of tracking in stationary environments. | Richard S. Sutton, Anna Koop, David Silver |
| 2007 | IJCAI | Reinforcement Learning of Local Shape in the Game of Go. | David Silver, Richard S. Sutton, Martin Mller |
| 2006 | AAAI | Incremental Least-Squares Temporal Difference Learning. | Alborz Geramifard, Michael H. Bowling, Richard S. Sutton |
| 2005 | ICML | TD(lambda) networks: temporal-difference networks with eligibility traces. | Brian Tanner, Richard S. Sutton |
| 2005 | IJCAI | Using Predictive Representations to Improve Generalization in Reinforcement Learning. | Eddie J. Rafols, Mark B. Ring, Richard S. Sutton, Brian Tanner |
| 2005 | IJCAI | Temporal-Difference Networks with History. | Brian Tanner, Richard S. Sutton |
| 2001 | ICML | Off-Policy Temporal Difference Learning with Function Approximation. | Doina Precup, Richard S. Sutton, Sanjoy Dasgupta |
| 2001 | ICML | Scaling Reinforcement Learning toward RoboCup Soccer. | Peter Stone, Richard S. Sutton |
| 2001 | RoboCup | Keepaway Soccer: A Machine Learning Testbed. | Peter Stone, Richard S. Sutton |
| 2000 | ICML | Eligibility Traces for Off-Policy Policy Evaluation. | Doina Precup, Richard S. Sutton, Satinder Singh |
| 2000 | RoboCup | Reinforcement Learning for 3 vs. 2 Keepaway | Peter Stone, Richard S. Sutton, Satinder Singh |
| 1998 | ICML | Intra-Option Learning about Temporally Abstract Actions. | Richard S. Sutton, Doina Precup, Satinder Singh |
| 1997 | ICANN | On the Significance of Markov Decision Processes. | Richard S. Sutton |
| 1997 | ICML | Exponentiated Gradient Methods for Reinforcement Learning. | Doina Precup, Richard S. Sutton |
| 1995 | ICML | TD Models: Modeling the World at a Mixture of Time Scales. | Richard S. Sutton |
| 1993 | ICML | Online Learning with Random Representations. | Richard S. Sutton, Steven D. Whitehead |
| 1992 | AAAI | Adapting Bias by Gradient Descent: An Incremental Version of Delta-Bar-Delta. | Richard S. Sutton |
| 1991 | ICML | Planning by Incremental Dynamic Programming. | Richard S. Sutton |
| 1991 | ICML | Learning Polynomial Functions by Feature Construction. | Richard S. Sutton, Christopher J. Matheus |
| 1990 | ICML | Integrated Architectures for Learning, Planning, and Reacting Based on Approximating Dynamic Programming. | Richard S. Sutton |
| 1985 | IJCAI | Training and Tracking in Robotics. | Oliver G. Selfridge, Richard S. Sutton, Andrew G. Barto |