| 2025 | COLING | LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks. | Akshara Prabhakar, Yuanzhi Li, Karthik Narasimhan, Sham M. Kakade, Eran Malach, Samy Jelassi |
| 2025 | ICLR | Mixture of Parrots: Experts improve memorization more than reasoning. | Samy Jelassi, Clara Mohri, David Brandfonbrener, Alex Gu, Nikhil Vyas, Nikhil Anand, David Alvarez-Melis, Yuanzhi Li, Sham M. Kakade, Eran Malach |
| 2025 | ICLR | A New Perspective on Shampoo's Preconditioner. | Depen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach, Sham M. Kakade, Lucas Janson |
| 2025 | ICLR | Don't stop me Now: Embedding based Scheduling for LLMS. | Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, Michael Mitzenmacher |
| 2025 | ICML | The Role of Sparsity for Length Generalization in LLMs. | Noah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran Malach |
| 2025 | ICML | Universal Length Generalization with Turing Programs. | Kaiying Hou, David Brandfonbrener, Sham M. Kakade, Samy Jelassi, Eran Malach |
| 2025 | ICML | The Power of Random Features and the Limits of Distribution-Free Gradient Descent. | Ari Karchmer, Eran Malach |
| 2024 | CRYPTO | Is ML-Based Cryptanalysis Inherently Limited? Simulating Cryptographic Adversaries via Gradient-Based Methods. | Avital Shafran, Eran Malach, Thomas Ristenpart, Gil Segev, Stefano Tessaro |
| 2024 | ICML | Repeat After Me: Transformers are Better than State Space Models at Copying. | Samy Jelassi, David Brandfonbrener, Sham M. Kakade, Eran Malach |
| 2024 | ICML | Auto-Regressive Next-Token Predictors are Universal Learners. | Eran Malach |
| 2022 | ICML | Efficient Learning of CNNs using Patch Based Features. | Alon Brutzkus, Amir Globerson, Eran Malach, Alon Regev Netser, Shai Shalev-Shwartz |
| 2021 | COLT | The Connection Between Approximation, Depth Separation and Learnability in Neural Networks. | Eran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad Shamir |
| 2021 | ICLR | Computational Separation Between Convolutional and Fully-Connected Networks. | Eran Malach, Shai Shalev-Shwartz |
| 2021 | ICML | Quantifying the Benefit of Using Differentiable Learning over Tangent Kernels. | Eran Malach, Pritish Kamath, Emmanuel Abbe, Nathan Srebro |
| 2020 | COLT | ID3 Learns Juntas for Smoothed Product Distributions. | Alon Brutzkus, Amit Daniely, Eran Malach |
| 2020 | ICML | Proving the Lottery Ticket Hypothesis: Pruning is All You Need. | Eran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad Shamir |
| 2018 | ICLR | SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data. | Alon Brutzkus, Amir Globerson, Eran Malach, Shai Shalev-Shwartz |