Skip to content

Alekh Agarwal

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

59

Venues

11

Active years

2006–2025

Best venue rank

A*

Where they publish

Papers

59 indexed papers, newest first.

YearVenueTitleAuthors
2025ACLOptimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models.Lior Belenki, Alekh Agarwal, Tianze Shi, Kristina Toutanova
2025ICLRRewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, Aviral Kumar
2025ICMLDesign Considerations in Offline Preference-based RL.Alekh Agarwal, Christoph Dann, Teodor Vanislavov Marinov
2025ICMLTheoretical guarantees on the best-of-n alignment policy.Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander Nicholas D'Amour, Jacob Eisenstein, Chirag Nagpal, Ananda Theertha Suresh
2025ICMLCatoni Contextual Bandits are Robust to Heavy-tailed Rewards.Chenlu Ye, Yujia Jin, Alekh Agarwal, Tong Zhang
2024ALTA Mechanism for Sample-Efficient In-Context Learning for Sparse Retrieval Tasks.Jacob D. Abernethy, Alekh Agarwal, Teodor Vanislavov Marinov, Manfred K. Warmuth
2024EMNLPConditional Language Policy: A General Framework For Steerable Multi-Objective Finetuning.Kaiwen Wang, Rahul Kidambi, Ryan Sullivan, Alekh Agarwal, Christoph Dann, Andrea Michi, Marco Gelmi, Yunxuan Li, Raghav Gupta, Avinava Dubey, Alexandre Ram, Johan Ferret, Geoffrey Cideron, Le Hou, Hongkun Yu, Amr Ahmed, Aranyak Mehta, Lonard Hussenot, Olivier Bachem, Edouard Leurent
2024ICMLThe Non-linear F-Design and Applications to Interactive Learning.Alekh Agarwal, Jian Qian, Alexander Rakhlin, Tong Zhang
2024ICMLA Minimaximalist Approach to Reinforcement Learning from Human Feedback.Gokul Swamy, Christoph Dann, Rahul Kidambi, Steven Wu, Alekh Agarwal
2024ICMLMore Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning.Kaiwen Wang, Owen Oertell, Alekh Agarwal, Nathan Kallus, Wen Sun
2024NAACLEfficient End-to-End Visual Document Understanding with Rationale Distillation.Wang Zhu, Alekh Agarwal, Mandar Joshi, Robin Jia, Jesse Thomason, Kristina Toutanova
2023COLTProvable Benefits of Representational Transfer in Reinforcement Learning.Alekh Agarwal, Yuda Song, Wen Sun, Kaiwen Wang, Mengdi Wang, Xuezhou Zhang
2023COLTVOQL: Towards Optimal Regret in Model-free RL with Nonlinear Function Approximation.Alekh Agarwal, Yujia Jin, Tong Zhang
2023ICMLLearning in POMDPs is Sample-Efficient with Hindsight Observability.Jonathan Lee, Alekh Agarwal, Christoph Dann, Tong Zhang
2023ICMLStochastic Gradient Succeeds for Bandits.Jincheng Mei, Zixin Zhong, Bo Dai, Alekh Agarwal, Csaba Szepesvri, Dale Schuurmans
2022COLTMinimax Regret Optimization for Robust Machine Learning under Distribution Shift.Alekh Agarwal, Tong Zhang
2022COLTNon-Linear Reinforcement Learning in Large Action Spaces: Structural Conditions and Sample-efficiency of Posterior Sampling.Alekh Agarwal, Tong Zhang
2022ICLRProvably Filtering Exogenous Distractors using Multistep Inverse Dynamics.Yonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal, John Langford
2022ICMLAdversarially Trained Actor Critic for Offline Reinforcement Learning.Ching-An Cheng, Tengyang Xie, Nan Jiang, Alekh Agarwal
2022ICMLEfficient Reinforcement Learning in Block MDPs: A Model-free Representation Learning approach.Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang, Alekh Agarwal, Wen Sun
2021COLTTowards a Dimension-Free Understanding of Adaptive Linear Control.Juan C. Perdomo, Max Simchowitz, Alekh Agarwal, Peter L. Bartlett
2021COLTCautiously Optimistic Policy Optimization and Exploration with Linear Function Approximation.Andrea Zanette, Ching-An Cheng, Alekh Agarwal
2021ICMLProvably Correct Optimization and Exploration with Non-linear Policies.Fei Feng, Wotao Yin, Alekh Agarwal, Lin Yang
2020AAAIMetareasoning in Modular Software Systems: On-the-Fly Configuration Using Reinforcement Learning with Rich Contextual Representations.Aditya Modi, Debadeepta Dey, Alekh Agarwal, Adith Swaminathan, Besmira Nushi, Sean Andrist, Eric Horvitz
2020COLTOptimality and Approximation with Policy Gradient Methods in Markov Decision Processes.Alekh Agarwal, Sham M. Kakade, Jason D. Lee, Gaurav Mahajan
2020COLTModel-Based Reinforcement Learning with a Generative Model is Minimax Optimal.Alekh Agarwal, Sham M. Kakade, Lin F. Yang
2020COLTTaking a hint: How to leverage loss predictors in contextual bandits?Chen-Yu Wei, Haipeng Luo, Alekh Agarwal
2020ICLRDeep Batch Active Learning by Diverse, Uncertain Gradient Lower Bounds.Jordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, Alekh Agarwal
2019COLTModel-based RL in Contextual Decision Processes: PAC bounds and Exponential Improvements over Model-free Approaches.Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford
2019ICLRBias Correction of Learned Generative Models via Likelihood-free Importance Weighting.Aditya Grover, Jiaming Song, Ashish Kapoor, Kenneth Tran, Alekh Agarwal, Eric Horvitz, Stefano Ermon
2019ICMLFair Regression: Quantitative Definitions and Reduction-Based Algorithms.Alekh Agarwal, Miroslav Dudk, Zhiwei Steven Wu
2019ICMLProvably efficient RL with Rich Observations via Latent State Decoding.Simon S. Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudk, John Langford
2019ICMLWarm-starting Contextual Bandits: Robustly Combining Supervised and Bandit Feedback.Chicheng Zhang, Alekh Agarwal, Hal Daum III, John Langford, Sahand Negahban
2019UAIOff-Policy Policy Gradient with Stationary Distribution Correction.Yao Liu, Adith Swaminathan, Alekh Agarwal, Emma Brunskill
2018COLTOpen Problem: The Dependence of Sample Complexity Lower Bounds on Planning Horizon.Nan Jiang, Alekh Agarwal
2018COLTEfficient Contextual Bandits in Non-stationary Worlds.Haipeng Luo, Chen-Yu Wei, Alekh Agarwal, John Langford
2018ICMLHierarchical Imitation and Reinforcement Learning.Hoang Minh Le, Nan Jiang, Alekh Agarwal, Miroslav Dudk, Yisong Yue, Hal Daum III
2018ICMLA Reductions Approach to Fair Classification.Alekh Agarwal, Alina Beygelzimer, Miroslav Dudk, John Langford, Hanna M. Wallach
2018ICMLPractical Contextual Bandits with Regression Oracles.Dylan J. Foster, Alekh Agarwal, Miroslav Dudk, Haipeng Luo, Robert E. Schapire
2017COLTOpen Problem: First-Order Regret Bounds for Contextual Bandits.Alekh Agarwal, Akshay Krishnamurthy, John Langford, Haipeng Luo, Robert E. Schapire
2017COLTCorralling a Band of Bandit Algorithms.Alekh Agarwal, Haipeng Luo, Behnam Neyshabur, Robert E. Schapire
2017ICMLContextual Decision Processes with low Bellman rank are PAC-Learnable.Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, Robert E. Schapire
2017ICMLActive Learning for Cost-Sensitive Classification.Akshay Krishnamurthy, Alekh Agarwal, Tzu-Kuo Huang, Hal Daum III, John Langford
2017ICMLOptimal and Adaptive Off-policy Evaluation in Contextual Bandits.Yu-Xiang Wang, Alekh Agarwal, Miroslav Dudk
2015ICMLA Lower Bound for the Optimization of Finite Sums.Alekh Agarwal, Lon Bottou
2015ICMLLearning to Search Better than Your Teacher.Kai-Wei Chang, Akshay Krishnamurthy, Alekh Agarwal, Hal Daum III, John Langford
2014CISSStochastic optimization and sparse statistical recovery: An optimal algorithm for high dimensions.Alekh Agarwal, Sahand N. Negahban, Martin J. Wainwright
2014COLTLearning Sparsely Used Overcomplete Dictionaries.Alekh Agarwal, Animashree Anandkumar, Prateek Jain, Praneeth Netrapalli, Rashish Tandon
2014COLTRobust Multi-objective Learning with Mentor Feedback.Alekh Agarwal, Ashwinkumar Badanidiyuru, Miroslav Dudk, Robert E. Schapire, Aleksandrs Slivkins
2014ICMLTaming the Monster: A Fast and Simple Algorithm for Contextual Bandits.Alekh Agarwal, Daniel J. Hsu, Satyen Kale, John Langford, Lihong Li, Robert E. Schapire
2014ICMLLeast Squares Revisited: Scalable Approaches for Multi-class Prediction.Alekh Agarwal, Sham M. Kakade, Nikos Karampatziakis, Le Song, Gregory Valiant
2013ICMLSelective sampling algorithms for cost-sensitive multiclass prediction.Alekh Agarwal
2011ICMLNoisy matrix decomposition via convex relaxation: Optimal rates in high dimensions.Alekh Agarwal, Sahand N. Negahban, Martin J. Wainwright
2011UAILearning with Missing Features.Afshin Rostamizadeh, Alekh Agarwal, Peter L. Bartlett
2010COLTOptimal Algorithms for Online Convex Optimization with Multi-Point Bandit Feedback.Alekh Agarwal, Ofer Dekel, Lin Xiao
2009COLTA Stochastic View of Optimal Regret through Minimax Duality.Jacob D. Abernethy, Alekh Agarwal, Peter L. Bartlett, Alexander Rakhlin
2008ICMLMessage-passing for graph-structured linear programs: proximal projections, convergence and rounding schemes.Pradeep Ravikumar, Alekh Agarwal, Martin J. Wainwright
2007ICMLLearning random walks to rank nodes in graphs.Alekh Agarwal, Soumen Chakrabarti
2006KDDLearning to rank networked entities.Alekh Agarwal, Soumen Chakrabarti, Sunny Aggarwal