Skip to content

Philip S. Thomas

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

36

Venues

11

Active years

2009–2024

Best venue rank

A*

Where they publish

Papers

36 indexed papers, newest first.

YearVenueTitleAuthors
2024AAAIFrom Past to Future: Rethinking Eligibility Traces.Dhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu, Philip S. Thomas, Bruno Castro da Silva
2024ICMLPosition: Benchmarking is Limited in Reinforcement Learning Research.Scott M. Jordan, Adam White, Bruno Castro da Silva, Martha White, Philip S. Thomas
2023AISTATSAsymptotically Unbiased Off-Policy Policy Evaluation when Reusing Old Data in Nonstationary Environments.Vincent Liu, Yash Chandak, Philip S. Thomas, Martha White
2023ICSESeldonian Toolkit: Building Software with Safe and Fair Machine Learning.Austin Hoag, James E. Kostas, Bruno Castro da Silva, Philip S. Thomas, Yuriy Brun
2022ICLRFairness Guarantees under Demographic Shift.Stephen Giguere, Blossom Metevier, Bruno Castro da Silva, Yuriy Brun, Philip S. Thomas, Scott Niekum
2022ITPMechanizing Soundness of Off-Policy Evaluation.Jared Yeager, J. Eliot B. Moss, Michael Norrish, Philip S. Thomas
2021AAAIHigh-Confidence Off-Policy (or Counterfactual) Variance Estimation.Yash Chandak, Shiv Shankar, Philip S. Thomas
2021ICMLHigh Confidence Generalization for Reinforcement Learning.James E. Kostas, Yash Chandak, Scott M. Jordan, Georgios Theocharous, Philip S. Thomas
2021ICMLPosterior Value Functions: Hindsight Baselines for Policy Gradient Methods.Chris Nota, Philip S. Thomas, Bruno C. da Silva
2021ICMLTowards Practical Mean Bounds for Small Samples.My Phan, Philip S. Thomas, Erik G. Learned-Miller
2021RecSysLarge-scale Interactive Conversational Recommendation System using Actor-Critic Framework.Ali Montazeralghaem, James Allan, Philip S. Thomas
2020AAAIReinforcement Learning When All Actions Are Not Always Available.Yash Chandak, Georgios Theocharous, Blossom Metevier, Philip S. Thomas
2020AAAILifelong Learning with a Changing Action Set.Yash Chandak, Georgios Theocharous, Chris Nota, Philip S. Thomas
2020ICMLOptimizing for the Future in Non-Stationary MDPs.Yash Chandak, Georgios Theocharous, Shiv Shankar, Martha White, Sridhar Mahadevan, Philip S. Thomas
2020ICMLEvaluating the Performance of Reinforcement Learning Algorithms.Scott M. Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, Philip S. Thomas
2020ICMLAsynchronous Coagent Networks.James E. Kostas, Chris Nota, Philip S. Thomas
2019AAAINatural Option Critic.Saket Tiwari, Philip S. Thomas
2019ICMLLearning Action Representations for Reinforcement Learning.Yash Chandak, Georgios Theocharous, James E. Kostas, Scott M. Jordan, Philip S. Thomas
2019ICMLConcentration Inequalities for Conditional Value at Risk.Philip S. Thomas, Erik G. Learned-Miller
2018ICMLDecoupling Gradient-Like Learning Rules from Representations.Philip S. Thomas, Christoph Dann, Emma Brunskill
2018IJCAIImportance Sampling for Fair Policy Selection.Shayan Doroudi, Philip S. Thomas, Emma Brunskill
2017AAAIImportance Sampling with Unequal Support.Philip S. Thomas, Emma Brunskill
2017AAAIPredictive Off-Policy Policy Evaluation for Nonstationary Decision Problems, with Applications to Digital Marketing.Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh, Ishan Durugkar, Emma Brunskill
2017ICMLData-Efficient Policy Evaluation Through Behavior Policy Search.Josiah P. Hanna, Philip S. Thomas, Peter Stone, Scott Niekum
2017UAIImportance Sampling for Fair Policy Selection.Shayan Doroudi, Philip S. Thomas, Emma Brunskill
2016AAAIIncreasing the Action Gap: New Operators for Reinforcement Learning.Marc G. Bellemare, Georg Ostrovski, Arthur Guez, Philip S. Thomas, Rmi Munos
2016ICMLData-Efficient Off-Policy Policy Evaluation for Reinforcement Learning.Philip S. Thomas, Emma Brunskill
2016ICMLEnergetic Natural Gradient Descent.Philip S. Thomas, Bruno Castro da Silva, Christoph Dann, Emma Brunskill
2015AAAIHigh-Confidence Off-Policy Evaluation.Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh
2015ICMLHigh Confidence Policy Improvement.Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh
2015IJCAIPersonalized Ad Recommendation Systems for Life-Time Value Optimization with Guarantees.Georgios Theocharous, Philip S. Thomas, Mohammad Ghavamzadeh
2015WWWAd Recommendation Systems for Life-Time Value Optimization.Georgios Theocharous, Philip S. Thomas, Mohammad Ghavamzadeh
2014AAAINatural Temporal Difference Learning.William Dabney, Philip S. Thomas
2011AAAIValue Function Approximation in Reinforcement Learning Using the Fourier Basis.George Dimitri Konidaris, Sarah Osentoski, Philip S. Thomas
2011ICMLConjugate Markov Decision Processes.Philip S. Thomas, Andrew G. Barto
2009IAAIApplication of the Actor-Critic Architecture to Functional Electrical Stimulation Control of a Human Arm.Philip S. Thomas, Antonie J. van den Bogert, Kathleen M. Jagodnik, Michael S. Branicky