| 2023 | ICLR | Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation. | Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel L. Smith, Yee Whye Teh |
| 2022 | ICLR | Deep Learning without Shortcuts: Shaping the Kernel with Tailored Rectifiers. | Guodong Zhang, Aleksandar Botev, James Martens |
| 2020 | ICLR | Hamiltonian Generative Networks. | Peter Toth, Danilo J. Rezende, Andrew Jaegle, Sbastien Racanire, Aleksandar Botev, Irina Higgins |
| 2018 | ICLR | A Scalable Laplace Approximation for Neural Networks. | Hippolyt Ritter, Aleksandar Botev, David Barber |
| 2017 | AISTATS | Complementary Sum Sampling for Likelihood Approximation in Large Scale Classification. | Aleksandar Botev, Bowen Zheng, David Barber |
| 2017 | ICML | Practical Gauss-Newton Optimisation for Deep Learning. | Aleksandar Botev, Hippolyt Ritter, David Barber |
| 2017 | IJCNN | Nesterov's accelerated gradient and momentum as approximations to regularised update descent. | Aleksandar Botev, Guy Lever, David Barber |
| 2017 | IJCNN | Overdispersed variational autoencoders. | Harshil Shah, David Barber, Aleksandar Botev |