| 2025 | AAAI | Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation. | Prashansa Panda, Shalabh Bhatnagar |
| 2025 | ICCV | One Encoder to Rule Them All: Representation Learning for Model-Free Visual Reinforcement Learning Using Fourier Neural Operators. | Parag Dutta, Mohd Ayyoob, Shalabh Bhatnagar, Ambedkar Dukkipati |
| 2025 | SiggraphA | Gradient-Weighted Feature Back-Projection: A Fast Alternative to Feature Distillation in 3D Gaussian Splatting. | Joji Joseph, Bharadwaj Amrutur, Shalabh Bhatnagar |
| 2024 | AISTATS | A Cubic-regularized Policy Newton Algorithm for Reinforcement Learning. | Mizhaan Prajit Maniyar, Prashanth L. A., Akash Mondal, Shalabh Bhatnagar |
| 2024 | ICPR | Learning Dynamic Representations in Large Language Models for Evolving Data Streams. | Ashish Srivastava, Shalabh Bhatnagar, M. Narasimha Murty, Jagannathan Ramanujam |
| 2024 | UAI | Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms. | Prashansa Panda, Shalabh Bhatnagar |
| 2024 | SiggraphA | Segmentation of 3D Gaussians using Masked Gradients. | Joji Joseph, Bharadwaj Amrutur, Shalabh Bhatnagar |
| 2023 | CISS | Generalized Simultaneous Perturbation Stochastic Approximation with Reduced Estimator Bias. | Shalabh Bhatnagar, Prashanth L. A. |
| 2023 | ICML | Off-Policy Average Reward Actor-Critic with Deterministic Policy Search. | Naman Saxena, Subhojyoti Khastagir, Shishir Kolathaya, Shalabh Bhatnagar |
| 2023 | RO-MAN | Autonomous UAV Navigation in Complex Environments using Human Feedback. | Sambhu H. Karumanchi, Raghuram Bharadwaj Diddigi, Prabuchandran K. J., Shalabh Bhatnagar |
| 2022 | AAAI | Gradient Temporal Difference with Momentum: Stability and Convergence. | Rohan Deb, Shalabh Bhatnagar |
| 2022 | ICAART | Robust Traffic Signal Timing Control using Multiagent Twin Delayed Deep Deterministic Policy Gradients. | Priya Shanmugasundaram, Shalabh Bhatnagar |
| 2022 | ICAART | Co-operative Multi-agent Twin Delayed DDPG for Robust Phase Duration Optimization of Large Road Networks. | Priya Shanmugasundaram, Shalabh Bhatnagar |
| 2022 | IJCNN | Neural Network Compatible Off-Policy Natural Actor-Critic Algorithm. | Raghuram Bharadwaj Diddigi, Prateek Jain, Prabuchandran K. J., Shalabh Bhatnagar |
| 2022 | ICRA | Dynamic Mirror Descent based Model Predictive Control for Accelerating Robot Learning. | Utkarsh A. Mishra, Soumya R. Samineni, Prakhar Goel, Chandravaran Kunjeti, Himanshu Lodha, Aman Singh, Aditya Sagi, Shalabh Bhatnagar, Shishir Kolathaya |
| 2022 | SMC | Data Efficient Safe Reinforcement Learning. | Sindhu Padakandla, Prabuchandran K. J., Sourav Ganguly, Shalabh Bhatnagar |
| 2020 | AAAI | Hierarchical Average Reward Policy Gradient Algorithms (Student Abstract). | Akshay Dharmavaram, Matthew Riemer, Shalabh Bhatnagar |
| 2020 | CoRL | Robust Quadrupedal Locomotion on Sloped Terrains: A Linear Policy Approach. | Kartik Paigwar, Lokesh Krishna, Sashank Tirumala, Naman Khetan, Aditya Varma, Ashish Joglekar, Shalabh Bhatnagar, Ashitava Ghosal, Bharadwaj Amrutur, Shishir Kolathaya |
| 2020 | ECAI | A Convergent Off-Policy Temporal Difference Algorithm. | Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Shalabh Bhatnagar |
| 2020 | IJCNN | Deep Reinforcement Learning with Successive Over-Relaxation and its Application in Autoscaling Cloud Resources. | Indu John, Shalabh Bhatnagar |
| 2020 | PIMRC | Learning-Based Resource Allocation in Industrial IoT Systems. | Sindhu Padakandla, Shilpa Rao, Shalabh Bhatnagar |
| 2020 | RO-MAN | Learning Stable Manoeuvres in Quadruped Robots from Expert Demonstrations. | Sashank Tirumala, Sagar Venkatesh Gubbi, Kartik Paigwar, Aditya Sagi, Ashish Joglekar, Shalabh Bhatnagar, Ashitava Ghosal, Bharadwaj Amrutur, Shishir Kolathaya |
| 2019 | COMAD | Efficient Budget Allocation and Task Assignment in Crowdsourcing. | Indu John, Shalabh Bhatnagar |
| 2019 | ICMLA | Predictive and Prescriptive Analytics for Performance Optimization: Framework and a Case Study on a Large-Scale Enterprise System. | Indu John, Ravikumar Karumanchi, Shalabh Bhatnagar |
| 2019 | ICRA | Realizing Learned Quadruped Locomotion Behaviors through Kinematic Motion Primitives. | Abhik Singla, Shounak Bhattacharya, Dhaivat Dholakiya, Shalabh Bhatnagar, Ashitava Ghosal, Bharadwaj Amrutur, Shishir Kolathaya |
| 2019 | RO-MAN | Learning Active Spine Behaviors for Dynamic and Efficient Locomotion in Quadruped Robots. | Shounak Bhattacharya, Abhik Singla, Abhimanyu, Dhaivat Dholakiya, Shalabh Bhatnagar, Bharadwaj Amrutur, Ashitava Ghosal, Shishir Kolathaya |
| 2019 | RO-MAN | Trajectory based Deep Policy Search for Quadrupedal Walking. | Shishir Kolathaya, Ashitava Ghosal, Bharadwaj Amrutur, Ashish Joglekar, Suhan Shetty, Dhaivat Dholakiya, Abhimanyu, Aditya Sagi, Shounak Bhattacharya, Abhik Singla, Shalabh Bhatnagar |
| 2017 | IJCNN | A model based search method for prediction in model-free Markov decision process. | Ajin George Joseph, Shalabh Bhatnagar |
| 2017 | IJCNN | Bounds for off-policy prediction in reinforcement learning. | Ajin George Joseph, Shalabh Bhatnagar |
| 2016 | ECAI | Revisiting the Cross Entropy Method with Applications in Stochastic Global Optimization and Reinforcement Learning. | Ajin George Joseph, Shalabh Bhatnagar |
| 2016 | ECAI | Shaping Proto-Value Functions Using Rewards. | Raj Kumar Maity, Chandrashekar Lakshminarayanan, Sindhu Padakandla, Shalabh Bhatnagar |
| 2016 | IJCNN | Scalable focussed entity resolution. | Ranganath B. N., Shalabh Bhatnagar |
| 2016 | WSC | A randomized algorithm for continuous optimization. | Ajin George Joseph, Shalabh Bhatnagar |
| 2015 | AAAI | A Generalized Reduced Linear Program for Markov Decision Processes. | Chandrashekar Lakshminarayanan, Shalabh Bhatnagar |
| 2015 | COMSNETS | Decentralized learning for traffic signal control. | Prabuchandran K. J., Hemanth Kumar A. N, Shalabh Bhatnagar |
| 2015 | ICONIP | A Stochastic Approximation Algorithm for Quantile Estimation. | Ajin George Joseph, Shalabh Bhatnagar |
| 2014 | COMSNETS | Adaptive sleep-wake control using reinforcement learning in sensor networks. | Prashanth L. A., Abhranil Chatterjee, Shalabh Bhatnagar |
| 2014 | HCOMP | A Markov Decision Process Framework for Predictable Job Completion Times on Crowdsourcing Platforms. | Chandrashekar Lakshminarayanan, Ayush Dubey, Shalabh Bhatnagar, Chithralekha Balamurugan |
| 2014 | WSC | Simulation optimization via gradient-based stochastic search. | Enlu Zhou, Shalabh Bhatnagar, Xi Chen |
| 2012 | ISIT | q-Gaussian based Smoothed Functional algorithms for stochastic optimization. | Debarghya Ghoshdastidar, Ambedkar Dukkipati, Shalabh Bhatnagar |
| 2011 | ICDCIT | Smoothed Functional and Quasi-Newton Algorithms for Routing in Multi-stage Queueing Network with Constraints. | K. Lakshmanan, Shalabh Bhatnagar |
| 2011 | ICSOC | Stochastic Optimization for Adaptive Labor Staffing in Service Systems. | Prashanth L. A., H. L. Prasad, Nirmit Desai, Shalabh Bhatnagar, Gargi Banerjee Dasgupta |
| 2010 | ICML | Toward Off-Policy Learning Control with Function Approximation. | Hamid Reza Maei, Csaba Szepesvri, Shalabh Bhatnagar, Richard S. Sutton |
| 2009 | ICML | Fast gradient-descent methods for temporal-difference learning with linear function approximation. | Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvri, Eric Wiewiora |
| 2008 | MMSP | SPSA based feature relevance estimation for video retrieval. | Sudha Velusamy, Shalabh Bhatnagar, S. V. Basavaraja, V. Sridhar |
| 2007 | ACC | Parametrized Actor-Critic Algorithms for Finite-Horizon MDPs. | Mohammed Shahid Abdulla, Shalabh Bhatnagar |
| 2007 | ACC | Solving MDPs using Two-timescale Simulated Annealing with Multiplicative Weights. | Mohammed Shahid Abdulla, Shalabh Bhatnagar |
| 2007 | ICDCIT | An Efficient and Optimized Bluetooth Scheduling Algorithm for Piconets. | Vijay Prakash Chaturvedi, V. Rakesh, Shalabh Bhatnagar |
| 2007 | ICDCIT | An Optimal Weighted-Average Congestion Based Pricing Scheme for Enhanced QoS. | Koteswara Rao Vemu, Shalabh Bhatnagar, N. Hemachandra |
| 2006 | WSC | SPSA algorithms with measurement reuse. | Mohammed Shahid Abdulla, Shalabh Bhatnagar |
| 2005 | CEC | Information theoretic justification of Boltzmann selection and its generalization to Tsallis case. | Ambedkar Dukkipati, M. Narasimha Murty, Shalabh Bhatnagar |
| 2005 | ISIT | Properties of Kullback-Leibler cross-entropy minimization in nonextensive framework. | Ambedkar Dukkipati, Narasimha Murty Musti, Shalabh Bhatnagar |
| 2004 | CEC | Cauchy annealing schedule: an annealing schedule for Boltzmann selection scheme in evolutionary algorithms. | Ambedkar Dukkipati, M. Narasimha Murty, Shalabh Bhatnagar |
| 2004 | ICPR | A Pattern Synthesis Technique with an Efficient Nearest Neighbor Classifier for Binary Pattern Recognition. | P. Viswanath, M. Narasimha Murty, Shalabh Bhatnagar |
| 2003 | CEC | Quotient evolutionary space: abstraction of evolutionary process w.r.t macroscopic properties. | Ambedkar Dukkipati, M. Narasimha Murty, Shalabh Bhatnagar |