Philip S. Thomas
Publication record assembled from the DBLP archive of ranked conferences.
Papers indexed
36
Venues
11
Active years
2009–2024
Best venue rank
A*
Where they publish
Papers
36 indexed papers, newest first.
| Year | Venue | Title | Authors |
|---|---|---|---|
| 2024 | AAAI | From Past to Future: Rethinking Eligibility Traces. | Dhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu, Philip S. Thomas, Bruno Castro da Silva |
| 2024 | ICML | Position: Benchmarking is Limited in Reinforcement Learning Research. | Scott M. Jordan, Adam White, Bruno Castro da Silva, Martha White, Philip S. Thomas |
| 2023 | AISTATS | Asymptotically Unbiased Off-Policy Policy Evaluation when Reusing Old Data in Nonstationary Environments. | Vincent Liu, Yash Chandak, Philip S. Thomas, Martha White |
| 2023 | ICSE | Seldonian Toolkit: Building Software with Safe and Fair Machine Learning. | Austin Hoag, James E. Kostas, Bruno Castro da Silva, Philip S. Thomas, Yuriy Brun |
| 2022 | ICLR | Fairness Guarantees under Demographic Shift. | Stephen Giguere, Blossom Metevier, Bruno Castro da Silva, Yuriy Brun, Philip S. Thomas, Scott Niekum |
| 2022 | ITP | Mechanizing Soundness of Off-Policy Evaluation. | Jared Yeager, J. Eliot B. Moss, Michael Norrish, Philip S. Thomas |
| 2021 | AAAI | High-Confidence Off-Policy (or Counterfactual) Variance Estimation. | Yash Chandak, Shiv Shankar, Philip S. Thomas |
| 2021 | ICML | High Confidence Generalization for Reinforcement Learning. | James E. Kostas, Yash Chandak, Scott M. Jordan, Georgios Theocharous, Philip S. Thomas |
| 2021 | ICML | Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods. | Chris Nota, Philip S. Thomas, Bruno C. da Silva |
| 2021 | ICML | Towards Practical Mean Bounds for Small Samples. | My Phan, Philip S. Thomas, Erik G. Learned-Miller |
| 2021 | RecSys | Large-scale Interactive Conversational Recommendation System using Actor-Critic Framework. | Ali Montazeralghaem, James Allan, Philip S. Thomas |
| 2020 | AAAI | Reinforcement Learning When All Actions Are Not Always Available. | Yash Chandak, Georgios Theocharous, Blossom Metevier, Philip S. Thomas |
| 2020 | AAAI | Lifelong Learning with a Changing Action Set. | Yash Chandak, Georgios Theocharous, Chris Nota, Philip S. Thomas |
| 2020 | ICML | Optimizing for the Future in Non-Stationary MDPs. | Yash Chandak, Georgios Theocharous, Shiv Shankar, Martha White, Sridhar Mahadevan, Philip S. Thomas |
| 2020 | ICML | Evaluating the Performance of Reinforcement Learning Algorithms. | Scott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, Philip S. Thomas |
| 2020 | ICML | Asynchronous Coagent Networks. | James E. Kostas, Chris Nota, Philip S. Thomas |
| 2019 | AAAI | Natural Option Critic. | Saket Tiwari, Philip S. Thomas |
| 2019 | ICML | Learning Action Representations for Reinforcement Learning. | Yash Chandak, Georgios Theocharous, James E. Kostas, Scott M. Jordan, Philip S. Thomas |
| 2019 | ICML | Concentration Inequalities for Conditional Value at Risk. | Philip S. Thomas, Erik G. Learned-Miller |
| 2018 | ICML | Decoupling Gradient-Like Learning Rules from Representations. | Philip S. Thomas, Christoph Dann, Emma Brunskill |
| 2018 | IJCAI | Importance Sampling for Fair Policy Selection. | Shayan Doroudi, Philip S. Thomas, Emma Brunskill |
| 2017 | AAAI | Importance Sampling with Unequal Support. | Philip S. Thomas, Emma Brunskill |
| 2017 | AAAI | Predictive Off-Policy Policy Evaluation for Nonstationary Decision Problems, with Applications to Digital Marketing. | Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh, Ishan Durugkar, Emma Brunskill |
| 2017 | ICML | Data-Efficient Policy Evaluation Through Behavior Policy Search. | Josiah P. Hanna, Philip S. Thomas, Peter Stone, Scott Niekum |
| 2017 | UAI | Importance Sampling for Fair Policy Selection. | Shayan Doroudi, Philip S. Thomas, Emma Brunskill |
| 2016 | AAAI | Increasing the Action Gap: New Operators for Reinforcement Learning. | Marc G. Bellemare, Georg Ostrovski, Arthur Guez, Philip S. Thomas, Rmi Munos |
| 2016 | ICML | Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. | Philip S. Thomas, Emma Brunskill |
| 2016 | ICML | Energetic Natural Gradient Descent. | Philip S. Thomas, Bruno Castro da Silva, Christoph Dann, Emma Brunskill |
| 2015 | AAAI | High-Confidence Off-Policy Evaluation. | Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh |
| 2015 | ICML | High Confidence Policy Improvement. | Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh |
| 2015 | IJCAI | Personalized Ad Recommendation Systems for Life-Time Value Optimization with Guarantees. | Georgios Theocharous, Philip S. Thomas, Mohammad Ghavamzadeh |
| 2015 | WWW | Ad Recommendation Systems for Life-Time Value Optimization. | Georgios Theocharous, Philip S. Thomas, Mohammad Ghavamzadeh |
| 2014 | AAAI | Natural Temporal Difference Learning. | William Dabney, Philip S. Thomas |
| 2011 | AAAI | Value Function Approximation in Reinforcement Learning Using the Fourier Basis. | George Dimitri Konidaris, Sarah Osentoski, Philip S. Thomas |
| 2011 | ICML | Conjugate Markov Decision Processes. | Philip S. Thomas, Andrew G. Barto |
| 2009 | IAAI | Application of the Actor-Critic Architecture to Functional Electrical Stimulation Control of a Human Arm. | Philip S. Thomas, Antonie J. van den Bogert, Kathleen M. Jagodnik, Michael S. Branicky |