| 2026 | COLT | Convergence of Continual Learning in Homogeneous Deep Networks. | Matan Schliserman, Gon Buzaglo, Itay Evron, Daniel Soudry |
| 2025 | ICLR | Scaling FP8 training to trillion-token LLMs. | Maxim Fishman, Brian Chmiel, Ron Banner, Daniel Soudry |
| 2025 | ICML | When Diffusion Models Memorize: Inductive Biases in Probability Flow of Minimum-Norm Shallow Neural Nets. | Chen Zeno, Hila Manor, Greg Ongie, Nir Weinberger, Tomer Michaeli, Daniel Soudry |
| 2024 | ICLR | Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators. | Yaniv Blumenfeld, Itay Hubara, Daniel Soudry |
| 2024 | ICLR | The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting - An Analytical Model. | Daniel Goldfarb, Itay Evron, Nir Weinberger, Daniel Soudry, Paul Hand |
| 2024 | ICML | How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers. | Gon Buzaglo, Itamar Harel, Mor Shpigel Nacson, Alon Brutzkus, Nathan Srebro, Daniel Soudry |
| 2023 | AISTATS | The Role of Codeword-to-Class Assignments in Error-Correcting Codes: An Empirical Study. | Itay Evron, Ophir Onn, Tamar Weiss Orzech, Hai Azeroual, Daniel Soudry |
| 2023 | CVPR | Alias-Free Convnets: Fractional Shift Invariance via Polynomial Activations. | Hagay Michaeli, Tomer Michaeli, Daniel Soudry |
| 2023 | ICLR | Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats. | Brian Chmiel, Ron Banner, Elad Hoffer, Hilla Ben-Yaacov, Daniel Soudry |
| 2023 | ICLR | Minimum Variance Unbiased N: M Sparsity for the Neural Gradients. | Brian Chmiel, Itay Hubara, Ron Banner, Daniel Soudry |
| 2023 | ICLR | The Implicit Bias of Minima Stability in Multivariate Shallow ReLU Networks. | Mor Shpigel Nacson, Rotem Mulayoff, Greg Ongie, Tomer Michaeli, Daniel Soudry |
| 2023 | ICML | Continual Learning in Linear Classification on Separable Data. | Itay Evron, Edward Moroshko, Gon Buzaglo, Maroun Khriesh, Badea Marjieh, Nathan Srebro, Daniel Soudry |
| 2023 | ICML | Gradient Descent Monotonically Decreases the Sharpness of Gradient Flow Solutions in Scalar Networks and Beyond. | Itai Kreisler, Mor Shpigel Nacson, Daniel Soudry, Yair Carmon |
| 2022 | AAAI | Regularization Guarantees Generalization in Bayesian Reinforcement Learning through Algorithmic Stability. | Aviv Tamar, Daniel Soudry, Ev Zisselman |
| 2022 | COLT | How catastrophic can catastrophic forgetting be in linear regression? | Itay Evron, Edward Moroshko, Rachel A. Ward, Nathan Srebro, Daniel Soudry |
| 2022 | ICLR | A Statistical Framework for Efficient Out of Distribution Detection in Deep Neural Networks. | Matan Haroush, Tzviel Frostig, Ruth Heller, Daniel Soudry |
| 2022 | ICML | Implicit Bias of the Step Size in Linear Diagonal Neural Networks. | Mor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, Daniel Soudry |
| 2021 | ICLR | Neural gradients are near-lognormal: improved quantized and sparse training. | Brian Chmiel, Liad Ben-Uri, Moran Shkolnik, Elad Hoffer, Ron Banner, Daniel Soudry |
| 2021 | ICML | On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror Descent. | Shahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth, Nathan Srebro, Amir Globerson, Daniel Soudry |
| 2021 | ICML | Accurate Post Training Quantization With Small Calibration Sets. | Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, Daniel Soudry |
| 2020 | COLT | Kernel and Rich Regimes in Overparametrized Models. | Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, Nathan Srebro |
| 2020 | CVPR | The Knowledge Within: Methods for Data-Free Model Compression. | Matan Haroush, Itay Hubara, Elad Hoffer, Daniel Soudry |
| 2020 | CVPR | Augment Your Batch: Improving Generalization Through Instance Repetition. | Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, Daniel Soudry |
| 2020 | ICLR | At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks? | Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel Soudry |
| 2020 | ICLR | A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case. | Greg Ongie, Rebecca Willett, Daniel Soudry, Nathan Srebro |
| 2020 | ICML | Beyond Signal Propagation: Is Feature Diversity Necessary in Deep Neural Network Initialization? | Yaniv Blumenfeld, Dar Gilboa, Daniel Soudry |
| 2019 | AISTATS | Convergence of Gradient Descent on Separable Data. | Mor Shpigel Nacson, Jason D. Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, Daniel Soudry |
| 2019 | AISTATS | Stochastic Gradient Descent on Separable Data: Exact Convergence with a Fixed Learning Rate. | Mor Shpigel Nacson, Nathan Srebro, Daniel Soudry |
| 2019 | COLT | How do infinite width bounded norm networks look in function space? | Pedro Savarese, Itay Evron, Daniel Soudry, Nathan Srebro |
| 2019 | ICML | Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models. | Mor Shpigel Nacson, Suriya Gunasekar, Jason D. Lee, Nathan Srebro, Daniel Soudry |
| 2018 | ICLR | Fix your classifier: the marginal value of training the last weight layer. | Elad Hoffer, Itay Hubara, Daniel Soudry |
| 2018 | ICLR | Exponentially vanishing sub-optimal local minima in multilayer neural networks. | Daniel Soudry, Elad Hoffer |
| 2018 | ICLR | The Implicit Bias of Gradient Descent on Separable Data. | Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Nathan Srebro |
| 2018 | ICML | Characterizing Implicit Bias in Terms of Optimization Geometry. | Suriya Gunasekar, Jason D. Lee, Daniel Soudry, Nathan Srebro |
| 2016 | ISCAS | A fully analog memristor-based neural network with online gradient training. | Eyal Rosenthal, Sergey Greshnikov, Daniel Soudry, Shahar Kvatinsky |