| 2025 | ICLR | Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count. | Hanseul Cho, Jaeyoung Cha, Srinadh Bhojanapalli, Chulhee Yun |
| 2025 | ICLR | Convergence and Implicit Bias of Gradient Descent on Continual Linear Classification. | Hyunji Jung, Hanseul Cho, Chulhee Yun |
| 2025 | ICLR | Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo. | Hyunsu Kim, Giung Nam, Chulhee Yun, Hongseok Yang, Juho Lee |
| 2025 | ICLR | Does SGD really happen in tiny subspaces? | Minhak Song, Kwangjun Ahn, Chulhee Yun |
| 2025 | ICML | Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty. | Yeseul Cho, Baekrok Shin, Changmin Kang, Chulhee Yun |
| 2025 | ICML | Provable Benefit of Random Permutations over Uniform Sampling in Stochastic Coordinate Descent. | Donghwa Kim, Jaewook Lee, Chulhee Yun |
| 2025 | ICML | Incremental Gradient Descent with Small Epoch Counts is Surprisingly Slow on Ill-Conditioned Problems. | Yujun Kim, Jaeyoung Cha, Chulhee Yun |
| 2025 | ICML | Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and More. | Geonhui Yoo, Minhak Song, Chulhee Yun |
| 2024 | ICLR | Linear attention is (maybe) all you need (to understand Transformer optimization). | Kwangjun Ahn, Xiang Cheng, Minhak Song, Chulhee Yun, Ali Jadbabaie, Suvrit Sra |
| 2024 | ICML | Fundamental Benefit of Alternating Updates in Minimax Optimization. | Jaewook Lee, Hanseul Cho, Chulhee Yun |
| 2023 | ICLR | SGDA with shuffling: faster convergence for nonconvex-PŁ minimax optimization. | Hanseul Cho, Chulhee Yun |
| 2023 | ICML | Tighter Lower Bounds for Shuffling SGD: Random Permutations and Beyond. | Jaeyoung Cha, Jaewook Lee, Chulhee Yun |
| 2023 | ICML | Provable Benefit of Mixup for Finding Optimal Decision Boundaries. | Junsoo Oh, Chulhee Yun |
| 2023 | ICML | On the Training Instability of Shuffling SGD with Batch Normalization. | David Xing Wu, Chulhee Yun, Suvrit Sra |
| 2022 | ICLR | Minibatch vs Local SGD with Shuffling: Tight Convergence Bounds and Beyond. | Chulhee Yun, Shashank Rajput, Suvrit Sra |
| 2021 | COLT | Provable Memorization via Deep Neural Networks using Sub-linear Parameters. | Sejun Park, Jaeho Lee, Chulhee Yun, Jinwoo Shin |
| 2021 | COLT | Open Problem: Can Single-Shuffle SGD be Better than Reshuffling SGD and GD? | Chulhee Yun, Suvrit Sra, Ali Jadbabaie |
| 2021 | ICLR | Minimum Width for Universal Approximation. | Sejun Park, Chulhee Yun, Jaeho Lee, Jinwoo Shin |
| 2021 | ICLR | A unifying view on implicit bias in training linear neural networks. | Chulhee Yun, Shankar Krishnan, Hossein Mobahi |
| 2020 | ICLR | Are Transformers universal approximators of sequence-to-sequence functions? | Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi, Sanjiv Kumar |
| 2020 | ICML | Low-Rank Bottleneck in Multi-head Attention Models. | Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi, Sanjiv Kumar |
| 2019 | ICLR | Efficiently testing local optimality and escaping saddles for ReLU networks. | Chulhee Yun, Suvrit Sra, Ali Jadbabaie |
| 2019 | ICLR | Small nonlinearities in activation functions create bad local minima in neural networks. | Chulhee Yun, Suvrit Sra, Ali Jadbabaie |
| 2018 | COLT | Minimax Bounds on Stochastic Batched Convex Optimization. | John C. Duchi, Feng Ruan, Chulhee Yun |
| 2018 | ICLR | Global Optimality Conditions for Deep Neural Networks. | Chulhee Yun, Suvrit Sra, Ali Jadbabaie |
| 2015 | ICASSP | Face detection using Local Hybrid Patterns. | Chulhee Yun, Donghoon Lee, Chang Dong Yoo |