| 2024 | ICASSP | Large Scale Self-Supervised Pretraining for Active Speaker Detection. | Otavio Braga, Wei Xia, Keith Johnson, Alice Chuang, Yunfan Ye, Olivier Siohan, Tuan Anh Nguyen |
| 2024 | ICASSP | Conformer is All You Need for Visual Speech Recognition. | Oscar Chang, Hank Liao, Dmitriy Serdyuk, Ankit Shahy, Olivier Siohan |
| 2023 | ICLR | Revisiting the Entropy Semiring for Neural Speech Recognition. | Oscar Chang, Dongseong Hwang, Olivier Siohan |
| 2023 | Interspeech | Cascaded encoders for fine-tuning ASR models on overlapped speech. | Richard Rose, Oscar Chang, Olivier Siohan |
| 2022 | ICASSP | Best of Both Worlds: Multi-Task Audio-Visual Automatic Speech Recognition and Active Speaker Detection. | Otavio Braga, Olivier Siohan |
| 2022 | Interspeech | End-to-End multi-talker audio-visual ASR using an active speaker attention module. | Richard Rose, Olivier Siohan |
| 2022 | Interspeech | Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Muti-Person Video. | Dmitriy Serdyuk, Otavio Braga, Olivier Siohan |
| 2021 | ASRU | Action Item Detection in Meetings Using Pretrained Transformers. | Kishan Sachdeva, Joshua Maynez, Olivier Siohan |
| 2021 | ASRU | Audio-Visual Speech Recognition is Worth $32\times 32\times 8$ Voxels. | Dmitriy Serdyuk, Otavio Braga, Olivier Siohan |
| 2021 | ICASSP | A Closer Look at Audio-Visual Multi-Person Speech Recognition and Active Speaker Selection. | Otavio Braga, Olivier Siohan |
| 2021 | Interspeech | Bridging the Gap Between Streaming and Non-Streaming ASR Systems by Distilling Ensembles of CTC and RNN-T Models. | Thibault Doutre, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Olivier Siohan, Liangliang Cao |
| 2021 | Interspeech | End-to-End Audio-Visual Speech Recognition for Overlapping Speech. | Richard Rose, Olivier Siohan, Anshuman Tripathi, Otavio Braga |
| 2020 | ICASSP | End-to-End Multi-Person Audio/Visual Automatic Speech Recognition. | Otavio Braga, Takaki Makino, Olivier Siohan, Hank Liao |
| 2019 | ASRU | Recurrent Neural Network Transducer for Audio-Visual Speech Recognition. | Takaki Makino, Hank Liao, Yannis M. Assael, Brendan Shillingford, Basilio Garcia, Otavio Braga, Olivier Siohan |
| 2017 | Interspeech | Acoustic Modeling for Google Home. | Bo Li, Tara N. Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean K. Chin, Khe Chai Sim, Ron J. Weiss, Kevin W. Wilson, Ehsan Variani, Chanwoo Kim, Olivier Siohan, Mitchel Weintraub, Erik McDermott, Richard Rose, Matt Shannon |
| 2017 | Interspeech | Annealed f-Smoothing as a Mechanism to Speed up Neural Network Training. | Tara N. Sainath, Vijayaditya Peddinti, Olivier Siohan, Arun Narayanan |
| 2017 | Interspeech | CTC Training of Multi-Phone Acoustic Models for Speech Recognition. | Olivier Siohan |
| 2016 | ICASSP | Sequence training of multi-task acoustic models using meta-state labels. | Olivier Siohan |
| 2016 | ICASSP | Selection and combination of hypotheses for dialectal speech recognition. | Victor Soto, Olivier Siohan, Mohamed Elfeky, Pedro J. Moreno |
| 2015 | ASRU | Multitask learning and system combination for automatic speech recognition. | Olivier Siohan, David Rybach |
| 2015 | ICASSP | Exemplar-based large vocabulary speech recognition using k-nearest neighbors. | Yanbo Xu, Olivier Siohan, David Simcha, Sanjiv Kumar, Hank Liao |
| 2015 | Interspeech | Large vocabulary automatic speech recognition for children. | Hank Liao, Golan Pundak, Olivier Siohan, Melissa K. Carroll, Noah Coccaro, Qi-Ming Jiang, Tara N. Sainath, Andrew W. Senior, Franoise Beaufays, Michiel Bacchiani |
| 2014 | ICASSP | Training data selection based on context-dependent state matching. | Olivier Siohan |
| 2014 | Interspeech | A big data approach to acoustic model training corpus selection. | Olga Kapralova, John Alex, Eugene Weinstein, Pedro J. Moreno, Olivier Siohan |
| 2013 | Interspeech | ivector-based acoustic data selection. | Olivier Siohan, Michiel Bacchiani |
| 2010 | Interspeech | Decision tree state clustering with word and syllable features. | Hank Liao, Christopher Alberti, Michiel Bacchiani, Olivier Siohan |
| 2009 | ICASSP | An audio indexing system for election video material. | Christopher Alberti, Michiel Bacchiani, Ari Bezman, Ciprian Chelba, Anastassia Drofa, Hank Liao, Pedro J. Moreno, Ted Power, Arnaud Sahuguet, Maria Shugrina, Olivier Siohan |
| 2007 | ASRU | The IBM 2007 speech transcription system for European parliamentary speeches. | Bhuvana Ramabhadran, Olivier Siohan, Abhinav Sethy |
| 2007 | ICASSP | Gaussian Mixture Language Models for Speech Recognition. | Mohamed Afify, Olivier Siohan, Ruhi Sarikaya |
| 2007 | SIGIR | Vocabulary independent spoken term detection. | Jonathan Mamou, Bhuvana Ramabhadran, Olivier Siohan |
| 2006 | ICASSP | Automated Quality Monitoring in the Call Center with ASR and Maximum Entropy. | Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
| 2006 | Interspeech | The IBM 2006 speech transcription system for european parliamentary speeches. | Bhuvana Ramabhadran, Olivier Siohan, Lidia Mangu, Geoffrey Zweig, Martin Westphal, Henrik Schulz, Alvaro Soneiro |
| 2006 | NAACL | Automated Quality Monitoring for Call Centers using Speech and NLP Technologies. | Geoffrey Zweig, Olivier Siohan, George Saon, Bhuvana Ramabhadran, Daniel Povey, Lidia Mangu, Brian Kingsbury |
| 2005 | ICASSP | Contructing Ensembles of ASR Systems Using Randomized Decision Trees. | Olivier Siohan, Bhuvana Ramabhadran, Brian Kingsbury |
| 2005 | Interspeech | Fast vocabulary-independent audio search using path-based graph indexing. | Olivier Siohan, Michiel Bacchiani |
| 2004 | Interspeech | Use of metadata to improve recognition of spontaneous speech and named entities. | Bhuvana Ramabhadran, Olivier Siohan, Geoffrey Zweig |
| 2004 | Interspeech | Speech recognition error analysis on the English MALACH corpus. | Olivier Siohan, Bhuvana Ramabhadran, Geoffrey Zweig |
| 2003 | ICASSP | Combining neighboring filter channels to improve quantile based histogram equalization. | Florian Hilger, Hermann Ney, Olivier Siohan, Frank K. Soong |
| 2003 | Interspeech | Hierarchical class n-gram language models: towards better estimation of unseen events in speech recognition. | Imed Zitouni, Olivier Siohan, Chin-Hui Lee |
| 2002 | ICASSP | A discriminative training criterion and an associated EM learning algorithm. | Mohamed Afify, Olivier Siohan |
| 2002 | ICASSP | A dynamic in-search discriminative training approach for large vocabulary speech recognition. | Hui Jiang, Olivier Siohan, Frank K. Soong, Chin-Hui Lee |
| 2002 | ICASSP | Towards knowledge-based features for HMM based large vocabulary automatic speech recognition. | Benoit Launay, Olivier Siohan, Arun C. Surendran, Chin-Hui Lee |
| 2002 | Interspeech | Bell labs approach to Aurora evaluation on connected digit recognition. | Jingdong Chen, Dimitris Dimitriadis, Hui Jiang, Qi Li, Tor Andr Myrvoll, Olivier Siohan, Frank K. Soong |
| 2002 | Interspeech | Backoff hierarchical class n-gram language modelling for automatic speech recognition systems. | Imed Zitouni, Olivier Siohan, Hong-Kwang Jeff Kuo, Chin-Hui Lee |
| 2001 | ICASSP | Sequential noise estimation with optimal forgetting for robust speech recognition. | Mohamed Afify, Olivier Siohan |
| 2001 | Interspeech | Evaluating the Aurora connected digit recognition task - a bell labs approach. | Mohamed Afify, Hui Jiang, Filipp Korkmazskiy, Chin-Hui Lee, Qi Li, Olivier Siohan, Frank K. Soong, Arun C. Surendran |
| 2001 | Interspeech | Minimax classification with parametric neighborhoods for noisy speech recognition. | Mohamed Afify, Olivier Siohan, Chin-Hui Lee |
| 2001 | Interspeech | An auditory system-based feature for robust speech recognition. | Qi Li, Frank K. Soong, Olivier Siohan |
| 2001 | Interspeech | A new verification-based fast match approach to large vocabulary speech recognition. | Feng Liu, Mohamed Afify, Hui Jiang, Olivier Siohan |
| 2001 | Interspeech | A real-time Japanese broadcast news closed-captioning system. | Olivier Siohan, Akio Ando, Mohamed Afify, Hui Jiang, Chin-Hui Lee, Qi Li, Feng Liu, Kazuo Onoe, Frank K. Soong, Qiru Zhou |
| 2000 | ICASSP | Multiple classifiers by constrained minimization. | Partha Niyogi, Jean-Benot Pierrot, Olivier Siohan |
| 2000 | ICASSP | Joint maximum a posteriori estimation of transformation and hidden Markov model parameters. | Olivier Siohan, Cristina Chesta, Chin-Hui Lee |
| 2000 | Interspeech | Constrained maximum likelihood linear regression for speaker adaptation. | Mohamed Afify, Olivier Siohan |
| 2000 | Interspeech | Extended maximum a posterior linear regression (EMAPLR) model adaptation for speech recognition. | Wu Chou, Olivier Siohan, Tor Andr Myrvoll, Chin-Hui Lee |
| 2000 | Interspeech | A high-performance auditory feature for robust speech recognition. | Qi Li, Frank K. Soong, Olivier Siohan |
| 2000 | Interspeech | Structural maximum a-posteriori linear regression for unsupervised speaker adaptation. | Tor Andr Myrvoll, Olivier Siohan, Chin-Hui Lee, Wu Chou |
| 1999 | ICASSP | Background model design for flexible and portable speaker verification systems. | Olivier Siohan, Chin-Hui Lee, Arun C. Surendran, Qi Li |
| 1999 | Interspeech | Maximum a posteriori linear regression for hidden Markov model adaptation. | Cristina Chesta, Olivier Siohan, Chin-Hui Lee |
| 1998 | ICASSP | Speaker verification using minimum verification error training. | Aaron E. Rosenberg, Olivier Siohan, Sarangarajan Parthasarathy |
| 1998 | ICASSP | Speaker identification using minimum classification error training. | Olivier Siohan, Aaron E. Rosenberg, Sarangarajan Parthasarathy |
| 1996 | ICASSP | A semi-continuous stochastic trajectory model for phoneme-based continuous speech recognition. | Olivier Siohan, Yifan Gong |
| 1995 | ICASSP | On the robustness of linear discriminant analysis as a preprocessing step for noisy speech recognition. | Olivier Siohan |
| 1995 | Interspeech | Noise adaptation using linear regression for continuous noisy speech recognition. | Olivier Siohan, Yifan Gong, Jean Paul Haton |
| 1994 | Interspeech | A comparison of three noisy speech recognition approaches. | Olivier Siohan, Yifan Gong, Jean Paul Haton |
| 1993 | Interspeech | A Bayesian approach to phone duration adaptation for lombard speech recognition. | Olivier Siohan, Yifan Gong, Jean Paul Haton |
| 1992 | Interspeech | Minimization of speech alignment error by iterative transformation for speaker adaptation. | Yifan Gong, Olivier Siohan, Jean Paul Haton |