| 2025 | ICLR | Adding Conditional Control to Diffusion Models with Reinforcement Learning. | Yulai Zhao, Masatoshi Uehara, Gabriele Scalia, Sun-Yuan Kung, Tommaso Biancalani, Sergey Levine, Ehsan Hajiramezanali |
| 2025 | ICLR | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design. | Chenyu Wang, Masatoshi Uehara, Yichun He, Amy Wang, Avantika Lal, Tommi S. Jaakkola, Sergey Levine, Aviv Regev, Hanchen Wang, Tommaso Biancalani |
| 2025 | ICML | Reward-Guided Iterative Refinement in Diffusion Models at Test-Time with Applications to Protein and DNA Design. | Masatoshi Uehara, Xingyu Su, Yulai Zhao, Xiner Li, Aviv Regev, Shuiwang Ji, Sergey Levine, Tommaso Biancalani |
| 2024 | AISTATS | Functional Graphical Models: Structure Enables Offline Data-Driven Optimization. | Kuba Grudzien Kuba, Masatoshi Uehara, Sergey Levine, Pieter Abbeel |
| 2024 | ICLR | Provable Reward-Agnostic Preference-Based Reinforcement Learning. | Wenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. Lee |
| 2024 | ICLR | Provable Offline Preference-Based Reinforcement Learning. | Wenhao Zhan, Masatoshi Uehara, Nathan Kallus, Jason D. Lee, Wen Sun |
| 2024 | ICML | Feedback Efficient Online Fine-Tuning of Diffusion Models. | Masatoshi Uehara, Yulai Zhao, Kevin Black, Ehsan Hajiramezanali, Gabriele Scalia, Nathaniel Lee Diamant, Alex M. Tseng, Sergey Levine, Tommaso Biancalani |
| 2023 | COLT | Inference on Strongly Identified Functionals of Weakly Identified Functions. | Andrew Bennett, Nathan Kallus, Xiaojie Mao, Whitney Newey, Vasilis Syrgkanis, Masatoshi Uehara |
| 2023 | COLT | Minimax Instrumental Variable Regression and L | Andrew Bennett, Nathan Kallus, Xiaojie Mao, Whitney Newey, Vasilis Syrgkanis, Masatoshi Uehara |
| 2023 | ICLR | PAC Reinforcement Learning for Predictive State Representations. | Wenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. Lee |
| 2023 | ICML | Computationally Efficient PAC RL in POMDPs with Latent Determinism and Conditional Embeddings. | Masatoshi Uehara, Ayush Sekhari, Jason D. Lee, Nathan Kallus, Wen Sun |
| 2023 | ICML | Distributional Offline Policy Evaluation with Predictive Error Guarantees. | Runzhe Wu, Masatoshi Uehara, Wen Sun |
| 2023 | KDD | Off-Policy Evaluation of Ranking Policies under Diverse User Behavior. | Haruka Kiyohara, Masatoshi Uehara, Yusuke Narita, Nobuyuki Shimizu, Yasuo Yamamoto, Yuta Saito |
| 2022 | ICLR | Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage. | Masatoshi Uehara, Wen Sun |
| 2022 | ICLR | Representation Learning for Online and Offline RL in Low-rank MDPs. | Masatoshi Uehara, Xuezhou Zhang, Wen Sun |
| 2022 | ICML | A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes. | Chengchun Shi, Masatoshi Uehara, Jiawei Huang, Nan Jiang |
| 2022 | ICML | Efficient Reinforcement Learning in Block MDPs: A Model-free Representation Learning approach. | Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang, Alekh Agarwal, Wen Sun |
| 2021 | COLT | Fast Rates for the Regret of Offline Reinforcement Learning. | Yichun Hu, Nathan Kallus, Masatoshi Uehara |
| 2021 | ICML | Optimal Off-Policy Evaluation from Multiple Logging Policies. | Nathan Kallus, Yuta Saito, Masatoshi Uehara |
| 2020 | AISTATS | A Unified Statistically Efficient Estimation Framework for Unnormalized Models. | Masatoshi Uehara, Takafumi Kanamori, Takashi Takenouchi, Takeru Matsuda |
| 2020 | AISTATS | Imputation estimators for unnormalized models with missing data. | Masatoshi Uehara, Takeru Matsuda, Jae Kwang Kim |
| 2020 | ICML | Double Reinforcement Learning for Efficient and Robust Off-Policy Evaluation. | Nathan Kallus, Masatoshi Uehara |
| 2020 | ICML | Statistically Efficient Off-Policy Policy Gradients. | Nathan Kallus, Masatoshi Uehara |
| 2020 | ICML | Minimax Weight and Q-Function Learning for Off-Policy Evaluation. | Masatoshi Uehara, Jiawei Huang, Nan Jiang |