| 2025 | COLT | A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis. | Akash Kumar, Rahul Parhi, Mikhail Belkin |
| 2025 | ICML | Task Generalization with Autoregressive Compositional Structure: Can Learning from D Tasks Generalize to DT Tasks? | Amirhesam Abedsoltan, Huaqing Zhang, Kaiyue Wen, Hongzhou Lin, Jingzhao Zhang, Mikhail Belkin |
| 2025 | ICML | Emergence in non-neural models: grokking modular arithmetic via average gradient outer product. | Neil Mallinar, Daniel Beaglehole, Libin Zhu, Adityanarayanan Radhakrishnan, Parthe Pandit, Mikhail Belkin |
| 2025 | NAACL | UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models. | Yijiang River Dong, Hongzhou Lin, Mikhail Belkin, Ramn Huerta, Ivan Vulic |
| 2024 | AISTATS | On the Nystrm Approximation for Preconditioning in Kernel Machines. | Amirhesam Abedsoltan, Parthe Pandit, Luis Rademacher, Mikhail Belkin |
| 2024 | ICLR | More is Better: when Infinite Overparameterization is Optimal and Overfitting is Obligatory. | James B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail Belkin |
| 2024 | ICLR | Quadratic models for understanding catapult dynamics of neural networks. | Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, Mikhail Belkin |
| 2024 | ICML | Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning. | Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, Mikhail Belkin |
| 2024 | UAI | Uncertainty Estimation with Recursive Feature Machines. | Daniel Gedon, Amirhesam Abedsoltan, Thomas B. Schn, Mikhail Belkin |
| 2023 | ICLR | Restricted Strong Convexity of Deep Learning Models with Smooth Activations. | Arindam Banerjee, Pedro Cisneros-Velarde, Libin Zhu, Mikhail Belkin |
| 2023 | ICML | Toward Large Kernel Models. | Amirhesam Abedsoltan, Mikhail Belkin, Parthe Pandit |
| 2023 | ICML | Cut your Losses with Squentropy. | Like Hui, Mikhail Belkin, Stephen Wright |
| 2023 | UAI | Neural tangent kernel at initialization: linear width suffices. | Arindam Banerjee, Pedro Cisneros-Velarde, Libin Zhu, Mikhail Belkin |
| 2022 | ICLR | Transition to Linearity of Wide Neural Networks is an Emerging Property of Assembling Weak Models. | Chaoyue Liu, Libin Zhu, Mikhail Belkin |
| 2021 | ICLR | Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification Tasks. | Like Hui, Mikhail Belkin |
| 2020 | ICLR | Accelerating SGD with momentum for over-parameterized learning. | Chaoyue Liu, Mikhail Belkin |
| 2019 | AISTATS | Does data interpolation contradict statistical optimality? | Mikhail Belkin, Alexander Rakhlin, Alexandre B. Tsybakov |
| 2019 | Interspeech | Kernel Machines Beat Deep Neural Networks on Mask-Based Single-Channel Speech Enhancement. | Like Hui, Siyuan Ma, Mikhail Belkin |
| 2018 | ALT | Unperturbed: spectral analysis beyond Davis-Kahan. | Justin Eldridge, Mikhail Belkin, Yusu Wang |
| 2018 | COLT | Approximation beats concentration? An approximation view on inference with smooth radial kernels. | Mikhail Belkin |
| 2018 | ICML | To Understand Deep Learning We Need to Understand Kernel Learning. | Mikhail Belkin, Siyuan Ma, Soumik Mandal |
| 2018 | ICML | The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning. | Siyuan Ma, Raef Bassily, Mikhail Belkin |
| 2016 | AAAI | The Hidden Convexity of Spectral Clustering. | James R. Voss, Mikhail Belkin, Luis Rademacher |
| 2016 | AISTATS | Back to the Future: Radial Basis Function Networks Revisited. | Qichao Que, Mikhail Belkin |
| 2016 | COLT | Basis Learning as an Algorithmic Primitive. | Mikhail Belkin, Luis Rademacher, James R. Voss |
| 2016 | ICML | Learning privately from multiparty data. | Jihun Hamm, Yingjun Cao, Mikhail Belkin |
| 2015 | COLT | Beyond Hartigan Consistency: Merge Distortion Metric for Hierarchical Clustering. | Justin Eldridge, Mikhail Belkin, Yusu Wang |
| 2015 | ICDCS | Crowd-ML: A Privacy-Preserving Learning Framework for a Crowd of Smart Devices. | Jihun Hamm, Adam C. Champion, Guoxing Chen, Mikhail Belkin, Dong Xuan |
| 2014 | COLT | The More, the Merrier: the Blessing of Dimensionality for Learning Large Gaussian Mixtures. | Joseph Anderson, Mikhail Belkin, Navin Goyal, Luis Rademacher, James R. Voss |
| 2013 | COLT | Blind Signal Separation in the Presence of Gaussian Noise. | Mikhail Belkin, Luis Rademacher, James R. Voss |
| 2011 | KDD | An iterated graph laplacian approach for ranking on manifolds. | Xueyuan Zhou, Mikhail Belkin, Nathan Srebro |
| 2010 | COLT | Toward Learning Gaussian Mixtures with Arbitrary Separation. | Mikhail Belkin, Kaushik Sinha |
| 2010 | FOCS | Polynomial Learning of Distribution Families. | Mikhail Belkin, Kaushik Sinha |
| 2010 | Interspeech | Learning speaker normalization using semisupervised manifold alignment. | Andrew R. Plummer, Mary E. Beckman, Mikhail Belkin, Eric Fosler-Lussier, Benjamin Munson |
| 2009 | COLT | A Note on Learning with Integral Operators. | Lorenzo Rosasco, Mikhail Belkin, Ernesto De Vito |
| 2009 | SODA | Constructing Laplace operator from point clouds in | Mikhail Belkin, Jian Sun, Yusu Wang |
| 2008 | ICML | Data spectroscopy: learning mixture models using eigenspaces of convolution operators. | Tao Shi, Mikhail Belkin, Bin Yu |
| 2008 | ICPR | Probabilistic mixtures of differential profiles for shape recognition. | Lei Ding, Mikhail Belkin |
| 2006 | FOCS | Heat Flow and a Faster Algorithm to Compute the Surface Area of a Convex Body. | Mikhail Belkin, Hariharan Narayanan, Partha Niyogi |
| 2005 | COLT | Towards a Theoretical Foundation for Laplacian-Based Manifold Methods. | Mikhail Belkin, Partha Niyogi |
| 2005 | ICML | Beyond the point cloud: from transductive to semi-supervised learning. | Vikas Sindhwani, Partha Niyogi, Mikhail Belkin |
| 2004 | COLT | Regularization and Semi-supervised Learning on Large Graphs. | Mikhail Belkin, Irina Matveeva, Partha Niyogi |
| 2004 | COLT | On the Convergence of Spectral Clustering on Random Samples: The Normalized Case. | Ulrike von Luxburg, Olivier Bousquet, Mikhail Belkin |
| 2004 | ICASSP | Tikhonov regularization and semi-supervised learning on large graphs. | Mikhail Belkin, Irina Matveeva, Partha Niyogi |