| 2025 | ICLR | The Belief State Transformer. | Edward S. Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Dinesh Jayaraman, Alex Lamb, John Langford |
| 2025 | ICLR | Does SGD really happen in tiny subspaces? | Minhak Song, Kwangjun Ahn, Chulhee Yun |
| 2025 | ICML | General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization. | Kwangjun Ahn, Gagik Magakyan, Ashok Cutkosky |
| 2024 | ICLR | Linear attention is (maybe) all you need (to understand Transformer optimization). | Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie, Suvrit Sra |
| 2024 | ICML | How to Escape Sharp Minima with Random Perturbations. | Kwangjun Ahn, Ali Jadbabaie, Suvrit Sra |
| 2024 | ICML | Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise. | Kwangjun Ahn, Zhiyu Zhang, Yunbum Kook, Yan Dai |
| 2022 | ICML | Understanding the unstable convergence of gradient descent. | Kwangjun Ahn, Jingzhao Zhang, Suvrit Sra |
| 2022 | ICML | Agnostic Learnability of Halfspaces via Logistic Loss. | Ziwei Ji, Kwangjun Ahn, Pranjal Awasthi, Satyen Kale, Stefani Karp |
| 2021 | COLT | Optimal dimension dependence of the Metropolis-Adjusted Langevin Algorithm. | Sinho Chewi, Chen Lu, Kwangjun Ahn, Xiang Cheng, Thibaut Le Gouic, Philippe Rigollet |
| 2020 | COLT | From Nesterov's Estimate Sequence to Riemannian Acceleration. | Kwangjun Ahn, Suvrit Sra |
| 2017 | ISIT | Information-theoretic limits of subspace clustering. | Kwangjun Ahn, Kangwook Lee, Changho Suh |