| 2024 | ICML | Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues. | Antonio Orvieto, Soham De, Caglar Gulcehre, Razvan Pascanu, Samuel L. Smith |
| 2023 | ICLR | Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation. | Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel L. Smith, Yee Whye Teh |
| 2023 | ICML | Resurrecting Recurrent Neural Networks for Long Sequences. | Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, aglar Glehre, Razvan Pascanu, Soham De |
| 2021 | ICLR | Characterizing signal propagation to close the performance gap in unnormalized ResNets. | Andrew Brock, Soham De, Samuel L. Smith |
| 2021 | ICLR | On the Origin of Implicit Regularization in Stochastic Gradient Descent. | Samuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham De |
| 2021 | ICML | High-Performance Large-Scale Image Recognition Without Normalization. | Andy Brock, Soham De, Samuel L. Smith, Karen Simonyan |
| 2020 | ICML | On the Generalization Benefit of Noise in Stochastic Gradient Descent. | Samuel L. Smith, Erich Elsen, Soham De |
| 2019 | ICML | The Effect of Network Width on Stochastic Gradient Descent and Generalization: an Empirical Study. | Daniel S. Park, Jascha Sohl-Dickstein, Quoc V. Le, Samuel L. Smith |
| 2018 | ICLR | Don't Decay the Learning Rate, Increase the Batch Size. | Samuel L. Smith, Pieter-Jan Kindermans, Chris Ying, Quoc V. Le |
| 2018 | ICLR | A Bayesian Perspective on Generalization and Stochastic Gradient Descent. | Samuel L. Smith, Quoc V. Le |
| 2018 | ICLR | Decoding Decoders: Finding Optimal Representation Spaces for Unsupervised Similarity Tasks. | Vitalii Zhelezniak, Dan Busbridge, April Shen, Samuel L. Smith, Nils Y. Hammerla |
| 2017 | ICLR | Offline bilingual word vectors, orthogonal transformations and the inverted softmax. | Samuel L. Smith, David H. P. Turban, Steven Hamblin, Nils Y. Hammerla |