| 2025 | Interspeech | ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition. | Thai-Binh Nguyen, Thi Van Nguyen, Quoc Truong Do, Chi Mai Luong |
| 2023 | ICASSP | AdapITN: A Fast, Reliable, and Dynamic Adaptive Inverse Text Normalization. | Thai Binh Nguyen, Le Duc Minh Nhat, Quang Minh Nguyen, Quoc Truong Do, Chi Mai Luong, Alexander Waibel |
| 2020 | Interspeech | Improving Vietnamese Named Entity Recognition from Speech Using Word Capitalization and Punctuation Recovery Models. | Thai Binh Nguyen, Quang Minh Nguyen, Thi Thu Hien Nguyen, Quoc Truong Do, Luong Chi Mai |
| 2018 | LREC | Construction of English-French Multimodal Affective Conversational Corpus from TV Dramas. | Sashi Novitasari, Quoc Truong Do, Sakriani Sakti, Dessi Puji Lestari, Satoshi Nakamura |
| 2017 | Interspeech | Toward Expressive Speech Translation: A Unified Sequence-to-Sequence LSTMs Approach for Translating Words and Emphasis. | Quoc Truong Do, Sakriani Sakti, Satoshi Nakamura |
| 2016 | EMNLP | Learning a Lexicon and Translation Model from Phoneme Lattices. | Oliver Adams, Graham Neubig, Trevor Cohn, Steven Bird, Quoc Truong Do, Satoshi Nakamura |
| 2016 | Interspeech | Transferring Emphasis in Speech Translation Using Hard-Attentional Neural Network Models. | Quoc Truong Do, Sakriani Sakti, Graham Neubig, Satoshi Nakamura |
| 2016 | Interspeech | A Hybrid System for Continuous Word-Level Emphasis Modeling Based on HMM State Clustering and Adaptive Training. | Quoc Truong Do, Tomoki Toda, Graham Neubig, Sakriani Sakti, Satoshi Nakamura |
| 2015 | ASRU | The NAIST ASR system for the 2015 Multi-Genre Broadcast challenge: On combination of deep learning systems using a rank-score function. | Quoc Truong Do, Michael Heck, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |
| 2015 | ICASSP | WFST-based structural classification integrating dnn acoustic features and RNN language features for speech recognition. | Quoc Truong Do, Satoshi Nakamura, Marc Delcroix, Takaaki Hori |
| 2015 | Interspeech | Preserving word-level emphasis in speech-to-speech translation using linear regression HSMMs. | Quoc Truong Do, Shinnosuke Takamichi, Sakriani Sakti, Graham Neubig, Tomoki Toda, Satoshi Nakamura |