Skip to content

Rishabh Agarwal

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

32

Venues

7

Active years

2017–2025

Best venue rank

A*

Where they publish

Papers

32 indexed papers, newest first.

YearVenueTitleAuthors
2025ICLRSmaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.Hritik Bansal, Arian Hosseini, Rishabh Agarwal, Vinh Q. Tran, Mehran Kazemi
2025ICLRInference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models.Yinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang, Bo Dai, Aviral Kumar, Rishabh Agarwal, Sridhar Thiagarajan, Craig Boutilier, Aleksandra Faust
2025ICLRTraining Language Models to Self-Correct via Reinforcement Learning.Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal M. P. Behbahani, Aleksandra Faust
2025ICLRAsynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models.Michael Noukhovitch, Shengyi Huang, Sophie Xhonneux, Arian Hosseini, Rishabh Agarwal, Aaron C. Courville
2025ICLRRewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.Amrith Setlur, Chirag Nagpal, Adam Fisch, Xinyang Geng, Jacob Eisenstein, Rishabh Agarwal, Alekh Agarwal, Jonathan Berant, Aviral Kumar
2025ICLRSpeculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling.Wenda Xu, Rujun Han, Zifeng Wang, Long T. Le, Dhruv Madeka, Lei Li, William Yang Wang, Rishabh Agarwal, Chen-Yu Lee, Tomas Pfister
2025ICLRGenerative Verifiers: Reward Modeling as Next-Token Prediction.Lunjun Zhang, Arian Hosseini, Hritik Bansal, Mehran Kazemi, Aviral Kumar, Rishabh Agarwal
2025ICMLReward-Guided Prompt Evolving in Reinforcement Learning for LLMs.Ziyu Ye, Rishabh Agarwal, Tianqi Liu, Rishabh Joshi, Sarmishta Velury, Quoc V. Le, Qijun Tan, Yuan Liu
2024ICLROn-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, Olivier Bachem
2024ICLRDistillSpec: Improving Speculative Decoding via Knowledge Distillation.Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-Franois Kagy, Rishabh Agarwal
2024ICMLStop Regressing: Training Value Functions via Classification for Scalable Deep RL.Jesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taga, Yevgen Chebotar, Ted Xiao, Alex Irpan, Sergey Levine, Pablo Samuel Castro, Aleksandra Faust, Aviral Kumar, Rishabh Agarwal
2024ICMLSiT: Symmetry-invariant Transformers for Generalisation in Reinforcement Learning.Matthias Weissenbacher, Rishabh Agarwal, Yoshinobu Kawahara
2023AISTATSA Novel Stochastic Gradient Descent Algorithm for Learning Principal Subspaces.Charline Le Lan, Joshua Greaves, Jesse Farebrother, Mark Rowland, Fabian Pedregosa, Rishabh Agarwal, Marc G. Bellemare
2023ICLRProto-Value Networks: Scaling Representation Learning with Auxiliary Tasks.Jesse Farebrother, Joshua Greaves, Rishabh Agarwal, Charline Le Lan, Ross Goroshin, Pablo Samuel Castro, Marc G. Bellemare
2023ICLROffline Q-learning on Diverse Multi-Task Data Both Scales And Generalizes.Aviral Kumar, Rishabh Agarwal, Xinyang Geng, George Tucker, Sergey Levine
2023ICLRRevisiting Bisimulation: A Sampling-Based State Similarity Pseudo-metric.Charline Le Lan, Rishabh Agarwal
2023ICLRInvestigating Multi-task Pretraining and Generalization in Reinforcement Learning.Adrien Ali Taga, Rishabh Agarwal, Jesse Farebrother, Aaron C. Courville, Marc G. Bellemare
2023ICMLBootstrapped Representations in Reinforcement Learning.Charline Le Lan, Stephen Tu, Mark Rowland, Anna Harutyunyan, Rishabh Agarwal, Marc G. Bellemare, Will Dabney
2023ICMLBigger, Better, Faster: Human-level Atari with human-level efficiency.Max Schwarzer, Johan S. Obando-Ceron, Aaron C. Courville, Marc G. Bellemare, Rishabh Agarwal, Pablo Samuel Castro
2023ICMLThe Dormant Neuron Phenomenon in Deep Reinforcement Learning.Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku Evci
2023ICMLRevisiting Bellman Errors for Offline Model Selection.Joshua P. Zitovsky, Daniel de Marchi, Rishabh Agarwal, Michael Rene Kosorok
2022AAAIControl-Oriented Model-Based Reinforcement Learning with Implicit Differentiation.Evgenii Nikishin, Romina Abachi, Rishabh Agarwal, Pierre-Luc Bacon
2022AISTATSOn the Generalization of Representations in Reinforcement Learning.Charline Le Lan, Stephen Tu, Adam Oberman, Rishabh Agarwal, Marc G. Bellemare
2022ICLRDR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization.Aviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville, George Tucker, Sergey Levine
2022IGARSSDetection Of Crop Water Stress In Maize Using Drone Based Hyperspectral Imaging.Jayantrao Mohite, Suryakant A. Sawant, Rishabh Agarwal, Ankur Pandit, Srinivasu Pappula
2021ICLRContrastive Behavioral Similarity Embeddings for Generalization in Reinforcement Learning.Rishabh Agarwal, Marlos C. Machado, Pablo Samuel Castro, Marc G. Bellemare
2021ICLRImplicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement Learning.Aviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey Levine
2020ICMLAn Optimistic Perspective on Offline Reinforcement Learning.Rishabh Agarwal, Dale Schuurmans, Mohammad Norouzi
2020ICMLRevisiting Fundamentals of Experience Replay.William Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio, Hugo Larochelle, Mark Rowland, Will Dabney
2019ICMLLearning to Generalize from Sparse and Underspecified Rewards.Rishabh Agarwal, Chen Liang, Dale Schuurmans, Mohammad Norouzi
2017GLOBECOMS-Pencil: A Smart Pencil Grip Monitoring System for Kids Using Sensors.Prakhar Gupta, Rishabh Agarwal, Surbhi Saraswat, Hari Prabhat Gupta, Tanima Dutta
2017ISDAComputing Theory Prime Implicates in Modal Logic.Manoj K. Raut, Tushar V. Kokane, Rishabh Agarwal