| 2025 | ICLR | To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions. | Noah Marshall, Ke Liang Xiao, Atish Agarwala, Elliot Paquette |
| 2025 | ICML | Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks. | Shikai Qiu, Lechao Xiao, Andrew Gordon Wilson, Jeffrey Pennington, Atish Agarwala |
| 2025 | ICML | Avoiding spurious sharpness minimization broadens applicability of SAM. | Sidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. Dauphin |
| 2025 | ICML | Exact risk curves of signSGD in High-Dimensions: quantifying preconditioning and noise-compression effects. | Ke Liang Xiao, Noah Marshall, Atish Agarwala, Elliot Paquette |
| 2023 | ICML | SAM operates far from home: eigenvalue regularization as a dynamical phenomenon. | Atish Agarwala, Yann N. Dauphin |
| 2023 | ICML | Second-order regression models exhibit progressive sharpening to the edge of stability. | Atish Agarwala, Fabian Pedregosa, Jeffrey Pennington |
| 2022 | ICML | Deep equilibrium networks are sensitive to initialization statistics. | Atish Agarwala, Samuel S. Schoenholz |
| 2021 | ICLR | One Network Fits All? Modular versus Monolithic Task Formulations in Neural Networks. | Atish Agarwala, Abhimanyu Das, Brendan Juba, Rina Panigrahy, Vatsal Sharan, Xin Wang, Qiuyi Zhang |