| 2011 | Randomized maximum entropy language models. | Puyang Xu, Sanjeev Khudanpur, Asela Gunawardana |
| 2011 | Robust understanding of spoken Chinese through character-based tagging and prior knowledge exploitation. | Weiqun Xu, Changchun Bao, Yali Li, Jielin Pan, Yonghong Yan |
| 2011 | Accent level adjustment in bilingual Thai-English text-to-speech synthesis. | Chai Wutiwiwatchai, Ausdang Thangthai, Ananlada Chotimongkol, Chatchawarn Hansakunbuntheung, Nattanun Thatphithakkul |
| 2011 | A novel bottleneck-BLSTM front-end for feature-level context modeling in conversational speech recognition. | Martin Wllmer, Bjrn W. Schuller, Gerhard Rigoll |
| 2011 | Crowd-sourcing for difficult transcription of speech. | Jason D. Williams, I. Dan Melamed, Tirso Alonso, Barbara Hollister, Jay G. Wilpon |
| 2011 | A convergence analysis of log-linear training and its application to speech recognition. | Simon Wiesler, Ralf Schlter, Hermann Ney |
| 2011 | Automatic detection of unnatural word-level segments in unit-selection speech synthesis. | William Yang Wang, Kallirroi Georgila |
| 2011 | Improving reverberant VTS for hands-free robust speech recognition. | Yongqiang Wang, Mark J. F. Gales |
| 2011 | Convolutive Bottleneck Network features for LVCSR. | Karel Vesel, Martin Karafit, Frantisek Grzl |
| 2011 | Improved spoken term detection using support vector machines with acoustic and context features from pseudo-relevance feedback. | Tsung-wei Tu, Hung-yi Lee, Lin-Shan Lee |
| 2011 | Utterance verification using garbage words for a hospital appointment system with speech interface. | Mitsuru Takaoka, Hiromitsu Nishizaki, Yoshihiro Sekiguchi |
| 2011 | Discriminative splitting of Gaussian/log-linear mixture HMMs for speech recognition. | Muhammad Ali Tahir, Ralf Schlter, Hermann Ney |
| 2011 | Frame-level AnyBoost for LVCSR with the MMI Criterion. | Ryuki Tachibana, Takashi Fukuda, Upendra V. Chaudhari, Bhuvana Ramabhadran, Puming Zhan |
| 2011 | Supervised and unsupervised feature selection for inferring social nature of telephone conversations from their content. | Anthony P. Stark, Izhak Shafran, Jeffrey A. Kaye |
| 2011 | From Modern Standard Arabic to Levantine ASR: Leveraging GALE for dialects. | Hagen Soltau, Lidia Mangu, Fadi Biadsy |
| 2011 | A Trajectory-based Parallel Model Combination with a unified static and dynamic parameter compensation for noisy speech recognition. | Khe Chai Sim, Minh-Thang Luong |
| 2011 | Socio-situational setting classification based on language use. | Yangyang Shi, Pascal Wiggers, Catholijn M. Jonker |
| 2011 | Efficient determinization of tagged word lattices using categorial and lexicographic semirings. | Izhak Shafran, Richard Sproat, Mahsa Yarmohammadi, Brian Roark |
| 2011 | Factored adaptation for separable compensation of speaker and environmental variability. | Michael L. Seltzer, Alex Acero |
| 2011 | Evolutionary discriminative speaker adaptation. | Sid-Ahmed Selouani |
| 2011 | Feature engineering in Context-Dependent Deep Neural Networks for conversational speech transcription. | Frank Seide, Gang Li, Xie Chen, Dong Yu |
| 2011 | Some properties of Bayesian sensing hidden Markov models. | George Saon, Jen-Tzung Chien |
| 2011 | Discriminative reranking of ASR hypotheses with morpholexical and N-best-list features. | Hasim Sak, Murat Saraclar, Tunga Gungor |
| 2011 | A convex hull approach to sparse representations for exemplar-based speech recognition. | Tara N. Sainath, David Nahamoo, Dimitri Kanevsky, Bhuvana Ramabhadran, Parikshit M. Shah |
| 2011 | Making Deep Belief Networks effective for large vocabulary continuous speech recognition. | Tara N. Sainath, Brian Kingsbury, Bhuvana Ramabhadran, Petr Fousek, Petr Novk, Abdel-rahman Mohamed |