| 2025 | EMNLP | Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference. | Hao Mark Chen, Wayne Luk, Yiu Ka Fai Cedric, Rui Li, Konstantin Mishchenko, Stylianos I. Venieris, Hongxiang Fan |
| 2024 | ICML | Prodigy: An Expeditiously Adaptive Parameter-Free Learner. | Konstantin Mishchenko, Aaron Defazio |
| 2023 | ICML | Learning-Rate-Free Learning by D-Adaptation. | Aaron Defazio, Konstantin Mishchenko |
| 2023 | ICML | Two Losses Are Better Than One: Faster Optimization Using a Cheaper Proxy. | Blake E. Woodworth, Konstantin Mishchenko, Francis R. Bach |
| 2022 | ICLR | IntSGD: Adaptive Floatless Compression of Stochastic Gradients. | Konstantin Mishchenko, Bokun Wang, Dmitry Kovalev, Peter Richtrik |
| 2022 | ICML | Proximal and Federated Random Reshuffling. | Konstantin Mishchenko, Ahmed Khaled, Peter Richtrik |
| 2022 | ICML | ProxSkip: Yes! Local Gradient Steps Provably Lead to Communication Acceleration! Finally! | Konstantin Mishchenko, Grigory Malinovsky, Sebastian U. Stich, Peter Richtrik |
| 2020 | AISTATS | Tighter Theory for Local SGD on Identical and Heterogeneous Data. | Ahmed Khaled, Konstantin Mishchenko, Peter Richtrik |
| 2020 | AISTATS | Revisiting Stochastic Extragradient. | Konstantin Mishchenko, Dmitry Kovalev, Egor Shulgin, Peter Richtrik, Yura Malitsky |
| 2020 | AISTATS | DAve-QN: A Distributed Averaged Quasi-Newton Method with Local Superlinear Convergence Rate. | Saeed Soori, Konstantin Mishchenko, Aryan Mokhtari, Maryam Mehri Dehnavi, Mert Grbzbalaban |
| 2020 | ICML | Adaptive Gradient Descent without Descent. | Yura Malitsky, Konstantin Mishchenko |
| 2020 | UAI | 99% of Worker-Master Communication in Distributed Optimization Is Not Needed. | Konstantin Mishchenko, Filip Hanzely, Peter Richtrik |
| 2018 | ICML | A Delay-tolerant Proximal-Gradient Algorithm for Distributed Learning. | Konstantin Mishchenko, Franck Iutzeler, Jrme Malick, Massih-Reza Amini |