| 2025 | ACL | Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? A Petroglyph Revisited. | Kazuki Irie |
| 2024 | ICANN | Self-organising Neural Discrete Representation Learning la Kohonen. | Kazuki Irie, Rbert Csords, Jrgen Schmidhuber |
| 2024 | ICLR | Exploring the Promise and Limits of Real-Time Recurrent Learning. | Kazuki Irie, Anand Gopalakrishnan, Jrgen Schmidhuber |
| 2023 | EMNLP | Approximating Two-Layer Feedforward Networks for Efficient Transformers. | Rbert Csords, Kazuki Irie, Jrgen Schmidhuber |
| 2023 | EMNLP | Practical Computational Power of Linear Transformers and Their Recurrent and Self-Referential Extensions. | Kazuki Irie, Rbert Csords, Jrgen Schmidhuber |
| 2023 | ICLR | Images as Weight Matrices: Sequential Image Generation Through Synaptic Learning Rules. | Kazuki Irie, Jrgen Schmidhuber |
| 2022 | EMNLP | CTL++: Evaluating Generalization on Never-Seen Compositional Patterns of Known Functions, and Compatibility of Neural Representations. | Rbert Csords, Kazuki Irie, Jrgen Schmidhuber |
| 2022 | ICLR | The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization. | Rbert Csords, Kazuki Irie, Jrgen Schmidhuber |
| 2022 | ICML | The Dual Form of Neural Networks Revisited: Connecting Test Time Predictions to Training Patterns via Spotlights of Attention. | Kazuki Irie, Rbert Csords, Jrgen Schmidhuber |
| 2022 | ICML | A Modern Self-Referential Weight Matrix That Learns to Modify Itself. | Kazuki Irie, Imanol Schlag, Rbert Csords, Jrgen Schmidhuber |
| 2021 | EMNLP | The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers. | Rbert Csords, Kazuki Irie, Jrgen Schmidhuber |
| 2021 | ICML | Linear Transformers Are Secretly Fast Weight Programmers. | Imanol Schlag, Kazuki Irie, Jrgen Schmidhuber |
| 2020 | ICASSP | Domain Robust, Fast, and Compact Neural Language Models. | Alexander Gerstenberger, Kazuki Irie, Pavel Golik, Eugen Beck, Hermann Ney |
| 2020 | ICASSP | How Much Self-Attention Do We Need? Trading Attention for Feed-Forward Layers. | Kazuki Irie, Alexander Gerstenberger, Ralf Schlter, Hermann Ney |
| 2020 | ICASSP | The Rwth Asr System for Ted-Lium Release 2: Improving Hybrid Hmm With Specaugment. | Wei Zhou, Wilfried Michel, Kazuki Irie, Markus Kitza, Ralf Schlter, Hermann Ney |
| 2019 | ASRU | Training Language Models for Long-Span Cross-Sentence Evaluation. | Kazuki Irie, Albert Zeyer, Ralf Schlter, Hermann Ney |
| 2019 | ASRU | A Comparison of Transformer and LSTM Encoder Decoder Models for ASR. | Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlter, Hermann Ney |
| 2019 | Interspeech | On the Choice of Modeling Unit for Sequence-to-Sequence Speech Recognition. | Kazuki Irie, Rohit Prabhavalkar, Anjuli Kannan, Antoine Bruguier, David Rybach, Patrick Nguyen |
| 2019 | Interspeech | Language Modeling with Deep Transformers. | Kazuki Irie, Albert Zeyer, Ralf Schlter, Hermann Ney |
| 2019 | Interspeech | RWTH ASR Systems for LibriSpeech: Hybrid vs Attention. | Christoph Lscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlter, Hermann Ney |
| 2018 | ICASSP | RADMM: Recurrent Adaptive Mixture Model with Applications to Domain Robust Language Modeling. | Kazuki Irie, Shankar Kumar, Michael Nirschl, Hank Liao |
| 2018 | ICASSP | Prediction of LSTM-RNN Full Context States as a Subtask for N-Gram Feedforward Language Models. | Kazuki Irie, Zhihong Lei, Ralf Schlter, Hermann Ney |
| 2018 | Interspeech | Investigation on Estimation of Sentence Probability by Combining Forward, Backward and Bi-directional LSTM-RNNs. | Kazuki Irie, Zhihong Lei, Liuhui Deng, Ralf Schlter, Hermann Ney |
| 2018 | Interspeech | Improved Training of End-to-end Attention Models for Speech Recognition. | Albert Zeyer, Kazuki Irie, Ralf Schlter, Hermann Ney |
| 2017 | ICASSP | Investigations on byte-level convolutional neural networks for language modeling in low resource speech recognition. | Kazuki Irie, Pavel Golik, Ralf Schlter, Hermann Ney |
| 2016 | ICASSP | Investigation on log-linear interpolation of multi-domain neural network language model. | Zoltn Tske, Kazuki Irie, Ralf Schlter, Hermann Ney |
| 2016 | Interspeech | LSTM, GRU, Highway and a Bit of Attention: An Empirical Overview for Language Modeling in Speech Recognition. | Kazuki Irie, Zoltn Tske, Tamer Alkhouli, Ralf Schlter, Hermann Ney |
| 2015 | Interspeech | On efficient training of word classes and their application to recurrent neural network language models. | Rami Botros, Kazuki Irie, Martin Sundermeyer, Hermann Ney |
| 2015 | Interspeech | Bag-of-words input for long history representation in neural network-based language models for speech recognition. | Kazuki Irie, Ralf Schlter, Hermann Ney |
| 2014 | ICASSP | The RWTH English lecture recognition system. | Simon Wiesler, Kazuki Irie, Zoltn Tske, Ralf Schlter, Hermann Ney |