| 2025 | ICML | Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression. | Jingfeng Wu, Peter L. Bartlett, Matus Telgarsky, Bin Yu |
| 2024 | AISTATS | Spectrum Extraction and Clipping for Implicitly Linear Layers. | Ali Ebrahimpour Boroojeny, Matus Telgarsky, Hari Sundaram |
| 2024 | COLT | Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency. | Jingfeng Wu, Peter L. Bartlett, Matus Telgarsky, Bin Yu |
| 2024 | ICML | Transformers, parallel computation, and logarithmic depth. | Clayton Sanford, Daniel Hsu, Matus Telgarsky |
| 2023 | ICLR | On Achieving Optimal Adversarial Test Error. | Justin D. Li, Matus Telgarsky |
| 2023 | ICLR | Feature selection and low test error in shallow low-rotation ReLU networks. | Matus Telgarsky |
| 2022 | COLT | Stochastic linear optimization never overfits with quadratically-bounded losses on general data. | Matus Telgarsky |
| 2022 | ICLR | Actor-critic is implicitly biased towards high entropy optimal policies. | Yuzheng Hu, Ziwei Ji, Matus Telgarsky |
| 2021 | ALT | Characterizing the implicit bias via a primal-dual analysis. | Ziwei Ji, Matus Telgarsky |
| 2021 | ICLR | Generalization bounds via distillation. | Daniel Hsu, Ziwei Ji, Matus Telgarsky, Lan Wang |
| 2021 | ICML | Fast margin maximization via dual acceleration. | Ziwei Ji, Nathan Srebro, Matus Telgarsky |
| 2020 | COLT | Gradient descent follows the regularization path for general losses. | Ziwei Ji, Miroslav Dudk, Robert E. Schapire, Matus Telgarsky |
| 2020 | ICLR | Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks. | Ziwei Ji, Matus Telgarsky |
| 2020 | ICLR | Neural tangent kernels, transportation mappings, and universal approximation. | Ziwei Ji, Matus Telgarsky, Ruicheng Xian |
| 2019 | COLT | The implicit bias of gradient descent on nonseparable data. | Ziwei Ji, Matus Telgarsky |
| 2019 | ICLR | Gradient descent aligns the layers of deep linear networks. | Ziwei Ji, Matus Telgarsky |
| 2019 | ICML | A Gradual, Semi-Discrete Approach to Generative Network Training via Explicit Wasserstein Minimization. | Yucheng Chen, Matus Telgarsky, Chao Zhang, Bolton Bailey, Daniel Hsu, Jian Peng |
| 2017 | COLT | Non-convex learning via Stochastic Gradient Langevin Dynamics: a nonasymptotic analysis. | Maxim Raginsky, Alexander Rakhlin, Matus Telgarsky |
| 2017 | ICML | Neural Networks and Rational Functions. | Matus Telgarsky |
| 2016 | COLT | benefits of depth in neural networks. | Matus Telgarsky |
| 2015 | ALT | Tensor Decompositions for Learning Latent Variable Models (A Survey for ALT). | Anima Anandkumar, Rong Ge, Daniel J. Hsu, Sham M. Kakade, Matus Telgarsky |
| 2015 | COLT | Convex Risk Minimization and Conditional Probability Estimation. | Matus Telgarsky, Miroslav Dudk |
| 2013 | COLT | Boosting with the Logistic Loss is Consistent. | Matus Telgarsky |
| 2013 | ICML | Margins, Shrinkage, and Boosting. | Matus Telgarsky |
| 2012 | ICML | Agglomerative Bregman Clustering. | Matus Telgarsky, Sanjoy Dasgupta |
| 2007 | ICASSP | Signal Decomposition using Multiscale Admixture Models. | Matus Telgarsky, John D. Lafferty |