| 2026 | COLT | On the Statistical Query Complexity of Learning Semiautomata: a Random Walk Approach. | George Giapitzakis, Kimon Fountoulakis, Eshaan Nichani, Jason D. Lee |
| 2026 | COLT | Provable Learning of Random Hierarchy Models and Hierarchical Shallow-to-Deep Chaining. | Yunwei Ren, Yatin Dandi, Florent Krzakala, Jason D. Lee |
| 2026 | COLT | Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization (Extended Abstract). | Jingfeng Wu, Peter L. Bartlett, Sham M. Kakade, Jason D. Lee, Bin Yu |
| 2025 | AISTATS | How Well Can Transformers Emulate In-Context Newton's Method? | Angeliki Giannou, Liu Yang, Tianhao Wang, Dimitris Papailiopoulos, Jason D. Lee |
| 2025 | COLT | Learning Compositional Functions with Transformers from Easy-to-Hard Data. | Zixuan Wang, Eshaan Nichani, Alberto Bietti, Alex Damian, Daniel Hsu, Jason D. Lee, Denny Wu |
| 2025 | COLT | Anytime Acceleration of Gradient Descent. | Zihan Zhang, Jason D. Lee, Simon S. Du, Yuxin Chen |
| 2025 | ICLR | Understanding Optimization in Deep Learning with Central Flows. | Jeremy Cohen, Alex Damian, Ameet Talwalkar, J. Zico Kolter, Jason D. Lee |
| 2025 | ICLR | Learning Hierarchical Polynomials of Multiple Nonlinear Features. | Hengyu Fu, Zihao Wang, Eshaan Nichani, Jason D. Lee |
| 2025 | ICLR | Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF. | Zhaolin Gao, Wenhao Zhan, Jonathan Daniel Chang, Gokul Swamy, Kiant Brantley, Jason D. Lee, Wen Sun |
| 2025 | ICLR | Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought. | Jianhao Huang, Zixuan Wang, Jason D. Lee |
| 2025 | ICLR | Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization. | Audrey Huang, Wenhao Zhan, Tengyang Xie, Jason D. Lee, Wen Sun, Akshay Krishnamurthy, Dylan J. Foster |
| 2025 | ICLR | Understanding Factual Recall in Transformers via Associative Memories. | Eshaan Nichani, Jason D. Lee, Alberto Bietti |
| 2025 | ICLR | Transformers Provably Learn Two-Mixture of Linear Classification via Gradient Flow. | Hongru Yang, Zhangyang Wang, Jason D. Lee, Yingbin Liang |
| 2025 | ICLR | Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank. | Wenhao Zhan, Scott Fujimoto, Zheqing Zhu, Jason D. Lee, Daniel Jiang, Yonathan Efroni |
| 2025 | ICML | Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension. | Yijun Dong, Yicheng Li, Yunai Li, Jason D. Lee, Qi Lei |
| 2025 | ICML | Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation. | Juno Kim, Denny Wu, Jason D. Lee, Taiji Suzuki |
| 2025 | ICML | Minimax Optimal Regret Bound for Reinforcement Learning with Trajectory Feedback. | Zihan Zhang, Yuxin Chen, Jason D. Lee, Simon Shaolei Du, Ruosong Wang |
| 2025 | ICML | Rethinking Addressing in Language Models via Contextualized Equivariant Positional Encoding. | Jiajun Zhu, Peihao Wang, Ruisi Cai, Jason D. Lee, Pan Li, Zhangyang Wang |
| 2024 | COLT | Computational-Statistical Gaps in Gaussian Single-Index Models (Extended Abstract). | Alex Damian, Loucas Pillaud-Vivien, Jason D. Lee, Joan Bruna |
| 2024 | COLT | Settling the sample complexity of online reinforcement learning. | Zihan Zhang, Yuxin Chen, Jason D. Lee, Simon S. Du |
| 2024 | COLT | Optimal Multi-Distribution Learning. | Zihan Zhang, Wenhao Zhan, Yuxin Chen, Simon S. Du, Jason D. Lee |
| 2024 | ICLR | Provably Efficient CVaR RL in Low-rank MDPs. | Yulai Zhao, Wenhao Zhan, Xiaoyan Hu, Ho-fung Leung, Farzan Farnia, Wen Sun, Jason D. Lee |
| 2024 | ICLR | Teaching Arithmetic to Small Transformers. | Nayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee, Dimitris Papailiopoulos |
| 2024 | ICLR | Dichotomy of Early and Late Phase Implicit Biases Can Provably Induce Grokking. | Kaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon Shaolei Du, Jason D. Lee, Wei Hu |
| 2024 | ICLR | Learning Hierarchical Polynomials with Three-Layer Neural Networks. | Zihao Wang, Eshaan Nichani, Jason D. Lee |
| 2024 | ICLR | Horizon-Free Regret for Linear Markov Decision Processes. | Zihan Zhang, Jason D. Lee, Yuxin Chen, Simon Shaolei Du |
| 2024 | ICLR | Provable Reward-Agnostic Preference-Based Reinforcement Learning. | Wenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. Lee |
| 2024 | ICLR | Provable Offline Preference-Based Reinforcement Learning. | Wenhao Zhan, Masatoshi Uehara, Nathan Kallus, Jason D. Lee, Wen Sun |
| 2024 | ICML | Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads. | Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao |
| 2024 | ICML | LoRA Training in the NTK Regime has No Spurious Local Minima. | Uijeong Jang, Jason D. Lee, Ernest K. Ryu |
| 2024 | ICML | An Information-Theoretic Analysis of In-Context Learning. | Hong Jun Jeon, Jason D. Lee, Qi Lei, Benjamin Van Roy |
| 2024 | ICML | How Transformers Learn Causal Structure with Gradient Descent. | Eshaan Nichani, Alex Damian, Jason D. Lee |
| 2024 | ICML | Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot. | Zixuan Wang, Stanley Wei, Daniel Hsu, Jason D. Lee |
| 2024 | ICML | Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark. | Yihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li, Yimeng Zhang, Wenqing Zheng, Pin-Yu Chen, Jason D. Lee, Wotao Yin, Mingyi Hong, Zhangyang Wang, Sijia Liu, Tianlong Chen |
| 2024 | NAACL | REST: Retrieval-Based Speculative Decoding. | Zhenyu He, Zexuan Zhong, Tianle Cai, Jason D. Lee, Di He |
| 2023 | AISTATS | Provable Hierarchy-Based Meta-Reinforcement Learning. | Kurtland Chua, Qi Lei, Jason D. Lee |
| 2023 | AISTATS | Optimal Sample Complexity Bounds for Non-convex Optimization under Kurdyka-Lojasiewicz Condition. | Qian Yu, Yining Wang, Baihe Huang, Qi Lei, Jason D. Lee |
| 2023 | AISTATS | Provably Efficient Reinforcement Learning via Surprise Bound. | Hanlin Zhu, Ruosong Wang, Jason D. Lee |
| 2023 | ICLR | Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability. | Alex Damian, Eshaan Nichani, Jason D. Lee |
| 2023 | ICLR | Can We Find Nash Equilibria at a Linear Rate in Markov Games? | Zhuoqing Song, Jason D. Lee, Zhuoran Yang |
| 2023 | ICLR | Decentralized Optimistic Hyperpolicy Mirror Descent: Provably No-Regret Learning in Markov Games. | Wenhao Zhan, Jason D. Lee, Zhuoran Yang |
| 2023 | ICLR | PAC Reinforcement Learning for Predictive State Representations. | Wenhao Zhan, Masatoshi Uehara, Wen Sun, Jason D. Lee |
| 2023 | ICML | Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning. | Yulai Zhao, Zhuoran Yang, Zhaoran Wang, Jason D. Lee |
| 2023 | ICML | Efficient displacement convex optimization with particle gradient descent. | Hadi Daneshmand, Jason D. Lee, Chi Jin |
| 2023 | ICML | Looped Transformers as Programmable Computers. | Angeliki Giannou, Shashank Rajput, Jy-yong Sohn, Kangwook Lee, Jason D. Lee, Dimitris Papailiopoulos |
| 2023 | ICML | Understanding Incremental Learning of Gradient Descent: A Fine-grained Analysis of Matrix Sensing. | Jikai Jin, Zhiyuan Li, Kaifeng Lyu, Simon Shaolei Du, Jason D. Lee |
| 2023 | ICML | Computationally Efficient PAC RL in POMDPs with Latent Determinism and Conditional Embeddings. | Masatoshi Uehara, Ayush Sekhari, Jason D. Lee, Nathan Kallus, Wen Sun |
| 2022 | AISTATS | Provably Efficient Policy Optimization for Two-Player Zero-Sum Markov Games. | Yulai Zhao, Yuandong Tian, Jason D. Lee, Simon S. Du |
| 2022 | COLT | Neural Networks can Learn Representations with Gradient Descent. | Alexandru Damian, Jason D. Lee, Mahdi Soltanolkotabi |
| 2022 | COLT | Optimization-Based Separations for Neural Networks. | Itay Safran, Jason D. Lee |
| 2022 | COLT | Offline Reinforcement Learning with Realizability and Single-policy Concentrability. | Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, Jason D. Lee |
| 2022 | ICASSP | Competitive Multi-Agent Reinforcement Learning with Self-Supervised Representation. | DiJia Su, Jason D. Lee, John M. Mulvey, H. Vincent Poor |
| 2022 | ICLR | Towards General Function Approximation in Zero-Sum Markov Games. | Baihe Huang, Jason D. Lee, Zhaoran Wang, Zhuoran Yang |
| 2021 | COLT | Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks. | Cong Fang, Jason D. Lee, Pengkun Yang, Tong Zhang |
| 2021 | COLT | Shape Matters: Understanding the Implicit Bias of the Noise Covariance. | Jeff Z. HaoChen, Colin Wei, Jason D. Lee, Tengyu Ma |
| 2021 | ICLR | Few-Shot Learning via Learning the Representation, Provably. | Simon Shaolei Du, Wei Hu, Sham M. Kakade, Jason D. Lee, Qi Lei |
| 2021 | ICLR | Impact of Representation Learning in Linear Bandits. | Jiaqi Yang, Wei Hu, Jason D. Lee, Simon Shaolei Du |
| 2021 | ICML | How Important is the Train-Validation Split in Meta-Learning? | Yu Bai, Minshuo Chen, Pan Zhou, Tuo Zhao, Jason D. Lee, Sham M. Kakade, Huan Wang, Caiming Xiong |
| 2021 | ICML | A Theory of Label Propagation for Subpopulation Shift. | Tianle Cai, Ruiqi Gao, Jason D. Lee, Qi Lei |
| 2021 | ICML | Bilinear Classes: A Structural Framework for Provable Generalization in RL. | Simon S. Du, Sham M. Kakade, Jason D. Lee, Shachar Lovett, Gaurav Mahajan, Wen Sun, Ruosong Wang |
| 2021 | ICML | Near-Optimal Linear Regression under Distribution Shift. | Qi Lei, Wei Hu, Jason D. Lee |
| 2020 | COLT | Optimality and Approximation with Policy Gradient Methods in Markov Decision Processes. | Alekh Agarwal, Sham M. Kakade, Jason D. Lee, Gaurav Mahajan |
| 2020 | COLT | Kernel and Rich Regimes in Overparametrized Models. | Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, Nathan Srebro |
| 2020 | ICLR | Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks. | Yu Bai, Jason D. Lee |
| 2020 | ICML | SGD Learns One-Layer Networks in WGANs. | Qi Lei, Jason D. Lee, Alex Dimakis, Constantinos Daskalakis |
| 2020 | ICML | Optimal transport mapping via input convex neural networks. | Ashok Vardhan Makkuva, Amirhossein Taghvaei, Sewoong Oh, Jason D. Lee |
| 2019 | AISTATS | Convergence of Gradient Descent on Separable Data. | Mor Shpigel Nacson, Jason D. Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, Daniel Soudry |
| 2019 | ICML | Gradient Descent Finds Global Minima of Deep Neural Networks. | Simon S. Du, Jason D. Lee, Haochuan Li, Liwei Wang, Xiyu Zhai |
| 2019 | ICML | Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models. | Mor Shpigel Nacson, Suriya Gunasekar, Jason D. Lee, Nathan Srebro, Daniel Soudry |
| 2018 | ICLR | Learning One-hidden-layer Neural Networks with Landscape Design. | Rong Ge, Jason D. Lee, Tengyu Ma |
| 2018 | ICLR | When is a Convolutional Filter Easy to Learn? | Simon S. Du, Jason D. Lee, Yuandong Tian |
| 2018 | ICLR | No Spurious Local Minima in a Two Hidden Unit ReLU Network. | Chenwei Wu, Jiajun Luo, Jason D. Lee |
| 2018 | ICML | On the Power of Over-parametrization in Neural Networks with Quadratic Activation. | Simon S. Du, Jason D. Lee |
| 2018 | ICML | Gradient Descent Learns One-hidden-layer CNN: Don't be Afraid of Spurious Local Minima. | Simon S. Du, Jason D. Lee, Yuandong Tian, Aarti Singh, Barnabs Pczos |
| 2018 | ICML | Characterizing Implicit Bias in Terms of Optimization Geometry. | Suriya Gunasekar, Jason D. Lee, Daniel Soudry, Nathan Srebro |
| 2018 | ICML | Gradient Primal-Dual Algorithm Converges to Second-Order Stationary Solution for Nonconvex Distributed Optimization Over Networks. | Mingyi Hong, Meisam Razaviyayn, Jason D. Lee |
| 2017 | AISTATS | Black-box Importance Sampling. | Qiang Liu, Jason D. Lee |
| 2017 | AISTATS | Sketching Meets Random Projection in the Dual: A Provable Recovery Algorithm for Big and High-dimensional Data. | Jialei Wang, Jason D. Lee, Mehrdad Mahdavi, Mladen Kolar, Nati Srebro |
| 2017 | AISTATS | On the Learnability of Fully-Connected Neural Networks. | Yuchen Zhang, Jason D. Lee, Martin J. Wainwright, Michael I. Jordan |
| 2016 | COLT | Gradient Descent Only Converges to Minimizers. | Jason D. Lee, Max Simchowitz, Michael I. Jordan, Benjamin Recht |
| 2016 | ICML | A Kernelized Stein Discrepancy for Goodness-of-fit Tests. | Qiang Liu, Jason D. Lee, Michael I. Jordan |
| 2016 | ICML | L1-regularized Neural Networks are Improperly Learnable in Polynomial Time. | Yuchen Zhang, Jason D. Lee, Michael I. Jordan |
| 2013 | AISTATS | Structure Learning of Mixed Graphical Models. | Jason D. Lee, Trevor Hastie |
| 2009 | ICCD | A distributed concurrent on-line test scheduling protocol for many-core NoC-based systems. | Jason D. Lee, Rabi N. Mahapatra, Praveen Bhojwani |
| 2008 | ICCD | In-field NoC-based SoC testing with distributed test vector storage. | Jason D. Lee, Rabi N. Mahapatra |
| 2007 | ISLPED | SAPP: scalable and adaptable peak power management in nocs. | Praveen Bhojwani, Jason D. Lee, Rabi N. Mahapatra |