| 2026 | ACL | Beyond Single-Shot: Multi-step Tool Retrieval via Query Planning. | Wei Fang, James R. Glass |
| 2025 | ACL | Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution. | Kun Li, Tianhua Zhang, Yunxiang Li, Hongyin Luo, Abdalla Mohamed Salama Sayed Moustafa, Xixin Wu, James R. Glass, Helen M. Meng |
| 2025 | ACL | Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains. | Kun Li, Tianhua Zhang, Xixin Wu, Hongyin Luo, James R. Glass, Helen M. Meng |
| 2025 | ACL | PLAY2PROMPT: Zero-shot Tool Instruction Optimization for LLM Agents via Tool Play. | Wei Fang, Yang Zhang, Kaizhi Qian, James R. Glass, Yada Zhu |
| 2025 | ASRU | USAD: Universal Speech and Audio Representation via Distillation. | Heng-Jui Chang, Saurabhchand Bhati, James R. Glass, Alexander H. Liu |
| 2025 | ASRU | Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM? | Andrew Rouditchenko, Saurabhchand Bhati, Edson Araujo, Samuel Thomas, Hilde Kuehne, Rogrio Feris, James R. Glass |
| 2025 | ASRU | Recognizing Dementia from Neuropsychological Tests with State Space Models. | Liming Wang, Saurabhchand Bhati, Cody Karjadi, Rhoda Au, James R. Glass |
| 2025 | CVPR | CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment. | Edson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati, Samuel Thomas, Brian Kingsbury, Leonid Karlinsky, Rogrio Feris, James R. Glass, Hilde Kuehne |
| 2025 | EMNLP | RAG-Zeval: Enhancing RAG Responses Evaluator through End-to-End Reasoning and Ranking-Based Reinforcement Learning. | Kun Li, Yunxiang Li, Tianhua Zhang, Hongyin Luo, Xixin Wu, James R. Glass, Helen M. Meng |
| 2025 | ICCV | Teaching VLMs to Localize Specific Objects from In-Context Examples. | Sivan Doveh, Nimrod Shabtay, Eli Schwartz, Hilde Kuehne, Raja Giryes, Rogrio Feris, Leonid Karlinsky, James R. Glass, Assaf Arbelle, Shimon Ullman, Muhammad Jehanzeb Mirza |
| 2025 | ICLR | Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts. | Junmo Kang, Leonid Karlinsky, Hongyin Luo, Zhen Wang, Jacob A. Hansen, James R. Glass, David Daniel Cox, Rameswar Panda, Rogrio Feris, Alan Ritter |
| 2025 | ICLR | UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation. | Alexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong, Yu-Chiang Frank Wang, James R. Glass, Rafael Valle, Bryan Catanzaro |
| 2025 | ICLR | Quantifying Generalization Complexity for Large Language Models. | Zhenting Qi, Hongyin Luo, Xuliang Huang, Zhuokai Zhao, Yibo Jiang, Xiangjun Fan, Himabindu Lakkaraju, James R. Glass |
| 2025 | ICML | SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models. | Yung-Sung Chuang, Benjamin Cohen-Wang, Zejiang Shen, Zhaofeng Wu, Hu Xu, Xi Victoria Lin, James R. Glass, Shang-Wen Li, Wen-tau Yih |
| 2025 | Interspeech | DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models. | Heng-Jui Chang, Hongyu Gong, Changhan Wang, James R. Glass, Yu-An Chung |
| 2024 | ACL | Found in the middle: Calibrating Positional Attention Bias Improves Long Context Utilization. | Cheng-Yu Hsieh, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long T. Le, Abhishek Kumar, James R. Glass, Alexander Ratner, Chen-Yu Lee, Ranjay Krishna, Tomas Pfister |
| 2024 | ACL | Self-Specialization: Uncovering Latent Expertise within Large Language Models. | Junmo Kang, Hongyin Luo, Yada Zhu, Jacob A. Hansen, James R. Glass, David D. Cox, Alan Ritter, Rogrio Feris, Leonid Karlinsky |
| 2024 | CVPR | What, When, and Where? Self-Supervised Spatio- Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions. | Brian Chen, Nina Shvetsova, Andrew Rouditchenko, Daniel Kondermann, Samuel Thomas, Shih-Fu Chang, Rogrio Feris, James R. Glass, Hilde Kuehne |
| 2024 | EACL | Joint Inference of Retrieval and Generation for Passage Re-ranking. | Wei Fang, Yung-Sung Chuang, James R. Glass |
| 2024 | EMNLP | Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps. | Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, Ranjay Krishna, Yoon Kim, James R. Glass |
| 2024 | EMNLP | Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers. | Tianhua Zhang, Kun Li, Hongyin Luo, Xixin Wu, James R. Glass, Helen Meng |
| 2024 | ICASSP | Cross-Lingual Transfer Learning for Low-Resource Speech Translation. | Sameer Khurana, Nauman Dawalatabad, Antoine Laurent, Luis Vicente, Pablo Gimeno, Victoria Mingote, James R. Glass |
| 2024 | ICASSP | Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective. | Alexander H. Liu, Sung-Lin Yeh, James R. Glass |
| 2024 | ICLR | Listen, Think, and Understand. | Yuan Gong, Hongyin Luo, Alexander H. Liu, Leonid Karlinsky, James R. Glass |
| 2024 | ICLR | DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. | Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, Pengcheng He |
| 2024 | ICLR | Curiosity-driven Red-teaming for Large Language Models. | Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James R. Glass, Akash Srivastava, Pulkit Agrawal |
| 2024 | Interspeech | Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer. | Liming Wang, Yuan Gong, Nauman Dawalatabad, Marco Vilela, Katerina Placek, Brian Tracey, Yishu Gong, Alan Premasiri, Fernando Vieira, James R. Glass |
| 2024 | NAACL | R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces. | Heng-Jui Chang, James R. Glass |
| 2023 | ACL | Expand, Rerank, and Retrieve: Query Reranking for Open-Domain Question Answering. | Yung-Sung Chuang, Wei Fang, Shang-Wen Li, Wen-tau Yih, James R. Glass |
| 2023 | ACL | Entailment as Robust Self-Learner. | Jiaxin Ge, Hongyin Luo, Yoon Kim, James R. Glass |
| 2023 | ACL | On the Blind Spots of Model-Based Evaluation Metrics for Text Generation. | Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov |
| 2023 | ASRU | Joint Audio and Speech Understanding. | Yuan Gong, Alexander H. Liu, Hongyin Luo, Leonid Karlinsky, James R. Glass |
| 2023 | ASRU | Audio-Visual Neural Syntax Acquisition. | Cheng-I Jeff Lai, Freda Shi, Puyuan Peng, Yoon Kim, Kevin Gimpel, Shiyu Chang, Yung-Sung Chuang, Saurabhchand Bhati, David D. Cox, David Harwath, Yang Zhang, Karen Livescu, James R. Glass |
| 2023 | EACL | Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence Reasoning. | Hongyin Luo, James R. Glass |
| 2023 | EMNLP | Search Augmented Instruction Learning. | Hongyin Luo, Tianhua Zhang, Yung-Sung Chuang, Yuan Gong, Yoon Kim, Xixin Wu, Helen Meng, James R. Glass |
| 2023 | ICASSP | On Unsupervised Uncertainty-Driven Speech Pseudo-Label Filtering and Model Calibration. | Nauman Dawalatabad, Sameer Khurana, Antoine Laurent, James R. Glass |
| 2023 | ICASSP | C2KD: Cross-Lingual Cross-Modal Knowledge Distillation for Multilingual Text-Video Retrieval. | Andrew Rouditchenko, Yung-Sung Chuang, Nina Shvetsova, Samuel Thomas, Rogrio Feris, Brian Kingsbury, Leonid Karlinsky, David Harwath, Hilde Kuehne, James R. Glass |
| 2023 | ICLR | Contrastive Audio-Visual Masked Autoencoder. | Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, James R. Glass |
| 2023 | Interspeech | Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers. | Yuan Gong, Sameer Khurana, Leonid Karlinsky, James R. Glass |
| 2023 | Interspeech | Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering. | Heng-Jui Chang, Alexander H. Liu, James R. Glass |
| 2023 | Interspeech | Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages. | Andrew Rouditchenko, Sameer Khurana, Samuel Thomas, Rogrio Feris, Leonid Karlinsky, Hilde Kuehne, David Harwath, Brian Kingsbury, James R. Glass |
| 2022 | AAAI | SSAST: Self-Supervised Audio Spectrogram Transformer. | Yuan Gong, Cheng-I Lai, Yu-An Chung, James R. Glass |
| 2022 | ACL | Controlling the Focus of Pretrained Language Generation Models. | Jiabao Ji, Yoon Kim, James R. Glass, Tianxing He |
| 2022 | ACL | Cross-Modal Discrete Representation Learning. | Alexander H. Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, James R. Glass |
| 2022 | CVPR | Everything at Once - Multi-modal Fusion Transformer for Video Retrieval. | Nina Shvetsova, Brian Chen, Andrew Rouditchenko, Samuel Thomas, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Hilde Kuehne |
| 2022 | EMNLP | Detecting Dementia from Long Neuropsychological Interviews. | Nauman Dawalatabad, Yuan Gong, Sameer Khurana, Rhoda Au, James R. Glass |
| 2022 | ICASSP | Transformer-Based Multi-Aspect Multi-Granularity Non-Native English Speaker Pronunciation Assessment. | Yuan Gong, Ziyi Chen, Iek-Heng Chu, Peng Chang, James R. Glass |
| 2022 | ICASSP | Vocalsound: A Dataset for Improving Human Vocal Sounds Recognition. | Yuan Gong, Jin Yu, James R. Glass |
| 2022 | ICASSP | Repetition Assessment for Speech and Language Disorders: A Study of the Logopenic Variant of Primary Progressive Aphasia. | R'mani Haulcy, Katerina Placek, Brian Tracey, Adam P. Vogel, James R. Glass |
| 2022 | ICASSP | Magic Dust for Cross-Lingual Adaptation of Monolingual Wav2vec-2.0. | Sameer Khurana, Antoine Laurent, James R. Glass |
| 2022 | ICASSP | On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis. | Cheng-I Jeff Lai, Erica Cooper, Yang Zhang, Shiyu Chang, Kaizhi Qian, Yi-Lun Liao, Yung-Sung Chuang, Alexander H. Liu, Junichi Yamagishi, David D. Cox, James R. Glass |
| 2022 | Interspeech | Simple and Effective Unsupervised Speech Synthesis. | Alexander H. Liu, Cheng-I Lai, Wei-Ning Hsu, Michael Auli, Alexei Baevski, James R. Glass |
| 2022 | LREC | Speak: A Toolkit Using Amazon Mechanical Turk to Collect and Validate Speech Audio Recordings. | Christopher Song, David Harwath, Tuka Alhanai, James R. Glass |
| 2022 | NAACL | DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings. | Yung-Sung Chuang, Rumen Dangovski, Hongyin Luo, Yang Zhang, Shiyu Chang, Marin Soljacic, Shang-Wen Li, Scott Yih, Yoon Kim, James R. Glass |
| 2022 | NAACL | Cooperative Self-training of Machine Reading Comprehension. | Hongyin Luo, Shang-Wen Li, Mingye Gao, Seunghak Yu, James R. Glass |
| 2021 | ACL | Text-Free Image-to-Speech Synthesis Using Learned Segmental Units. | Wei-Ning Hsu, David Harwath, Tyler Miller, Christopher Song, James R. Glass |
| 2021 | CVPR | Spoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions. | Mathew Monfort, SouYoung Jin, Alexander H. Liu, David Harwath, Rogrio Feris, James R. Glass, Aude Oliva |
| 2021 | EACL | Analyzing the Forgetting Problem in Pretrain-Finetuning of Open-domain Dialogue Response Models. | Tianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James R. Glass, Fuchun Peng |
| 2021 | EMNLP | Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation? | Tianxing He, Jingzhao Zhang, Zhiming Zhou, James R. Glass |
| 2021 | ICASSP | Similarity Analysis of Self-Supervised Speech Representations. | Yu-An Chung, Yonatan Belinkov, James R. Glass |
| 2021 | ICASSP | Semi-Supervised Spoken Language Understanding via Self-Supervised Speech and Language Model Pretraining. | Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, James R. Glass |
| 2021 | ICCV | Multimodal Clustering Networks for Self-supervised Learning from Unlabeled Videos. | Brian Chen, Andrew Rouditchenko, Kevin Duarte, Hilde Kuehne, Samuel Thomas, Angie W. Boggust, Rameswar Panda, Brian Kingsbury, Rogrio Feris, David Harwath, James R. Glass, Michael Picheny, Shih-Fu Chang |
| 2021 | Interspeech | AST: Audio Spectrogram Transformer. | Yuan Gong, Yu-An Chung, James R. Glass |
| 2021 | Interspeech | CLAC: A Speech Corpus of Healthy English Speakers. | R'mani Haulcy, James R. Glass |
| 2021 | Interspeech | Non-Autoregressive Predictive Coding for Learning Speech Representations from Local Dependencies. | Alexander H. Liu, Yu-An Chung, James R. Glass |
| 2021 | Interspeech | Joint Retrieval-Extraction Training for Evidence-Aware Dialog Response Selection. | Hongyin Luo, James R. Glass, Garima Lalwani, Yi Zhang, Shang-Wen Li |
| 2021 | Interspeech | Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset. | Ian Palmer, Andrew Rouditchenko, Andrei Barbu, Boris Katz, James R. Glass |
| 2021 | Interspeech | Cascaded Multilingual Audio-Visual Learning from Videos. | Andrew Rouditchenko, Angie W. Boggust, David Harwath, Samuel Thomas, Hilde Kuehne, Brian Chen, Rameswar Panda, Rogrio Feris, Brian Kingsbury, Michael Picheny, James R. Glass |
| 2021 | Interspeech | AVLnet: Learning Audio-Visual Language Representations from Instructional Videos. | Andrew Rouditchenko, Angie W. Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogrio Schmidt Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, James R. Glass |
| 2020 | ACL | What Was Written vs. Who Read It: News Media Profiling Using Text Analysis and Social Media Context. | Ramy Baly, Georgi Karadzhov, Jisun An, Haewoon Kwak, Yoan Dinkov, Ahmed Ali, James R. Glass, Preslav Nakov |
| 2020 | ACL | Improved Speech Representations with Multi-Target Autoregressive Predictive Coding. | Yu-An Chung, James R. Glass |
| 2020 | ACL | Negative Training for Neural Dialogue Response Generation. | Tianxing He, James R. Glass |
| 2020 | ACL | Similarity Analysis of Contextual Word Representation Models. | John M. Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass |
| 2020 | EMNLP | We Can Detect Your Bias: Predicting the Political Ideology of News Articles. | Ramy Baly, Giovanni Da San Martino, James R. Glass, Preslav Nakov |
| 2020 | ICASSP | Generative Pre-Training for Speech with Autoregressive Predictive Coding. | Yu-An Chung, James R. Glass |
| 2020 | ICASSP | Learning a Subword Inventory Jointly with End-to-End Automatic Speech Recognition. | Jennifer Drexler, James R. Glass |
| 2020 | ICASSP | Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT. | Franois Grondin, Hao Tang, James R. Glass |
| 2020 | ICASSP | Trilingual Semantic Embeddings of Visually Grounded Speech with Self-Attention Mechanisms. | Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David Harwath, James R. Glass |
| 2020 | ICASSP | ADI17: A Fine-Grained Arabic Dialect Identification Dataset. | Suwon Shon, Ahmed Ali, Younes Samih, Hamdy Mubarak, James R. Glass |
| 2020 | ICLR | Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech. | David Harwath, Wei-Ning Hsu, James R. Glass |
| 2020 | IJCNLP | A Systematic Characterization of Sampling Algorithms for Open-ended Language Generation. | Moin Nadeem, Tianxing He, Kyunghyun Cho, James R. Glass |
| 2020 | Interspeech | What Does an End-to-End Dialect Identification Model Learn About Non-Dialectal Information? | Shammur A. Chowdhury, Ahmed Ali, Suwon Shon, James R. Glass |
| 2020 | Interspeech | Vector-Quantized Autoregressive Predictive Coding. | Yu-An Chung, Hao Tang, James R. Glass |
| 2020 | Interspeech | Unsupervised Methods for Evaluating Speech Representations. | Michael Gump, Wei-Ning Hsu, James R. Glass |
| 2020 | Interspeech | A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning. | Sameer Khurana, Antoine Laurent, Wei-Ning Hsu, Jan Chorowski, Adrian Lancucki, Ricard Marxer, James R. Glass |
| 2020 | Interspeech | Prototypical Q Networks for Automatic Conversational Diagnosis and Few-Shot New Disease Adaption. | Hongyin Luo, Shang-Wen Li, James R. Glass |
| 2020 | Interspeech | Pair Expansion for Learning Multilingual Semantic Embeddings Using Disjoint Visually-Grounded Speech Audio Datasets. | Yasunori Ohishi, Akisato Kimura, Takahito Kawanishi, Kunio Kashino, David Harwath, James R. Glass |
| 2020 | Interspeech | Multimodal Association for Speaker Verification. | Suwon Shon, James R. Glass |
| 2019 | AAAI | What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models. | Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, James R. Glass |
| 2019 | AAAI | NeuroX: A Toolkit for Analyzing Individual Neurons in Neural Networks. | Fahim Dalvi, Avery Nortonsmith, Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, James R. Glass |
| 2019 | ACL | Improving Neural Language Models by Segmenting, Attending, and Predicting the Future. | Hongyin Luo, Lan Jiang, Yonatan Belinkov, James R. Glass |
| 2019 | ASRU | The MGB-5 Challenge: Recognition and Dialect Identification of Dialectal Arabic Speech. | Ahmed Ali, Suwon Shon, Younes Samih, Hamdy Mubarak, Ahmed Abdelali, James R. Glass, Steve Renals, Khalid Choukri |
| 2019 | ASRU | Explicit Alignment of Text and Speech Encodings for Attention-Based End-to-End Speech Recognition. | Jennifer Drexler, James R. Glass |
| 2019 | CVPR | Grounding Spoken Words in Unlabeled Video. | Angie W. Boggust, Kartik Audhkhasi, Dhiraj Joshi, David Harwath, Samuel Thomas, Rogrio Schmidt Feris, Danny Gutfreund, Yang Zhang, Antonio Torralba, Michael Picheny, James R. Glass |
| 2019 | CVPR | Learning Words by Drawing Images. | Didac Suris, Adri Recasens, David Bau, David Harwath, James R. Glass, Antonio Torralba |
| 2019 | EMNLP | Contrastive Language Adaptation for Cross-Lingual Stance Detection. | Mitra Mohtarami, James R. Glass, Preslav Nakov |
| 2019 | EMNLP | Tanbih: Get To Know What You Are Reading. | Yifan Zhang, Giovanni Da San Martino, Alberto Barrn-Cedeo, Salvatore Romeo, Jisun An, Haewoon Kwak, Todor Staykovski, Israa Jaradat, Georgi Karadzhov, Ramy Baly, Kareem Darwish, James R. Glass, Preslav Nakov |
| 2019 | ICASSP | Towards Unsupervised Speech-to-text Translation. | Yu-An Chung, Wei-Hung Weng, Schrasing Tong, James R. Glass |
| 2019 | ICASSP | Subword Regularization and Beam Search Decoding for End-to-end Automatic Speech Recognition. | Jennifer Drexler, James R. Glass |
| 2019 | ICASSP | SVD-PHAT: A Fast Sound Source Localization Method. | Franois Grondin, James R. Glass |
| 2019 | ICASSP | Towards Visually Grounded Sub-word Speech Unit Discovery. | David Harwath, James R. Glass |
| 2019 | ICASSP | Disentangling Correlated Speaker and Noise for Speech Synthesis via Data Augmentation and Adversarial Factorization. | Wei-Ning Hsu, Yu Zhang, Ron J. Weiss, Yu-An Chung, Yuxuan Wang, Yonghui Wu, James R. Glass |
| 2019 | ICASSP | A Factorial Deep Markov Model for Unsupervised Disentangled Representation Learning from Speech. | Sameer Khurana, Shafiq Rayhan Joty, Ahmed Ali, James R. Glass |
| 2019 | ICASSP | Dialogue State Tracking with Convolutional Semantic Taggers. | Mandy Korpusik, James R. Glass |
| 2019 | ICASSP | Domain Attentive Fusion for End-to-end Dialect Identification with Unknown Target Domain. | Suwon Shon, Ahmed Ali, James R. Glass |
| 2019 | ICASSP | Noise-tolerant Audio-visual Online Person Verification Using an Attention-based Neural Network Fusion. | Suwon Shon, Tae-Hyun Oh, James R. Glass |
| 2019 | ICLR | Identifying and Controlling Important Neurons in Neural Machine Translation. | Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass |
| 2019 | ICLR | Detecting Egregious Responses in Neural Sequence-to-sequence Models. | Tianxing He, James R. Glass |
| 2019 | Interspeech | Towards Bilingual Lexicon Discovery From Visually Grounded Speech Audio. | Emmanuel Azuh, David Harwath, James R. Glass |
| 2019 | Interspeech | Analyzing Phonetic and Graphemic Representations in End-to-End Automatic Speech Recognition. | Yonatan Belinkov, Ahmed Ali, James R. Glass |
| 2019 | Interspeech | An Unsupervised Autoregressive Model for Speech Representation Learning. | Yu-An Chung, Wei-Ning Hsu, Hao Tang, James R. Glass |
| 2019 | Interspeech | A Deep Residual Network for Large-Scale Acoustic Scene Analysis. | Logan Ford, Hao Tang, Franois Grondin, James R. Glass |
| 2019 | Interspeech | Multiple Sound Source Localization with SVD-PHAT. | Franois Grondin, James R. Glass |
| 2019 | Interspeech | Transfer Learning from Audio-Visual Grounding to Speech Recognition. | Wei-Ning Hsu, David Harwath, James R. Glass |
| 2019 | Interspeech | A Comparison of Deep Learning Methods for Language Understanding. | Mandy Korpusik, Zoe Liu, James R. Glass |
| 2019 | Interspeech | Integrating Video Retrieval and Moment Detection in a Unified Corpus for Video Question Answering. | Hongyin Luo, Mitra Mohtarami, James R. Glass, Karthik Krishnamurthy, Brigitte Richardson |
| 2019 | Interspeech | MCE 2018: The 1st Multi-Target Speaker Detection and Identification Challenge Evaluation. | Suwon Shon, Najim Dehak, Douglas A. Reynolds, James R. Glass |
| 2019 | Interspeech | VoiceID Loss: Speech Enhancement for Speaker Verification. | Suwon Shon, Hao Tang, James R. Glass |
| 2019 | IROS | Fast and Robust 3-D Sound Source Localization with DSVD-PHAT. | Franois Grondin, James R. Glass |
| 2019 | NAACL | Multi-Task Ordinal Regression for Jointly Predicting the Trustworthiness and the Leading Political Ideology of News Media. | Ramy Baly, Georgi Karadzhov, Abdelrhman Saleh, James R. Glass, Preslav Nakov |
| 2019 | NAACL | Analysis Methods in Neural Language Processing: A Survey. | Yonatan Belinkov, James R. Glass |
| 2019 | NAACL | FAKTA: An Automatic End-to-End Fact Checking System. | Moin Nadeem, Wei Fang, Brian Xu, Mitra Mohtarami, James R. Glass |
| 2018 | AAAI | Fact Checking in Community Forums. | Tsvetomila Mihaylova, Preslav Nakov, Llus Mrquez, Alberto Barrn-Cedeo, Mitra Mohtarami, Georgi Karadzhov, James R. Glass |
| 2018 | ECCV | Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input. | David Harwath, Adri Recasens, Ddac Surs, Galen Chuang, Antonio Torralba, James R. Glass |
| 2018 | EMNLP | Predicting Factuality of Reporting and Bias of News Media Sources. | Ramy Baly, Georgi Karadzhov, Dimitar Alexandrov, James R. Glass, Preslav Nakov |
| 2018 | EMNLP | Learning Word Representations with Cross-Sentence Dependencyfor End-to-End Co-reference Resolution. | Hongyin Luo, James R. Glass |
| 2018 | ICASSP | Vision as an Interlingua: Learning Multilingual Semantic Embeddings of Untranscribed Speech. | David Harwath, Galen Chuang, James R. Glass |
| 2018 | ICASSP | Extracting Domain Invariant Features by Unsupervised Learning for Robust Automatic Speech Recognition. | Wei-Ning Hsu, James R. Glass |
| 2018 | ICASSP | Energy-Efficient Speaker Identification with Low-Precision Networks. | Skanda Koppula, James R. Glass, Anantha P. Chandrakasan |
| 2018 | ICASSP | Convolutional Neural Networks and Multitask Strategies for Semantic Mapping of Natural Language Input to a Structured Database. | Mandy Korpusik, James R. Glass |
| 2018 | ICASSP | Exploiting Convolutional Neural Networks for Phonotactic Based Dialect Identification. | Maryam Najafian, Sameer Khurana, Suwon Shon, Ahmed Ali, James R. Glass |
| 2018 | ICPR | A Noise-Robust Self-Adaptive Multitarget Speaker Detection System. | Siqi Zheng, Jianzong Wang, Jing Xiao, Wei-Ning Hsu, James R. Glass |
| 2018 | Interspeech | Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech. | Yu-An Chung, James R. Glass |
| 2018 | Interspeech | Detecting Depression with Audio/Text Sequence Modeling of Interviews. | Tuka Al Hanai, Mohammad M. Ghassemi, James R. Glass |
| 2018 | Interspeech | Scalable Factorized Hierarchical Variational Autoencoder Training. | Wei-Ning Hsu, James R. Glass |
| 2018 | Interspeech | Unsupervised Adaptation with Interpretable Disentangled Representations for Distant Conversational Speech Recognition. | Wei-Ning Hsu, Hao Tang, James R. Glass |
| 2018 | Interspeech | A Study of Enhancement, Augmentation and Autoencoder Methods for Domain Adaptation in Distant Speech Recognition. | Hao Tang, Wei-Ning Hsu, Franois Grondin, James R. Glass |
| 2018 | NAACL | Integrating Stance Detection and Fact Checking in a Unified Corpus. | Ramy Baly, Mitra Mohtarami, James R. Glass, Llus Mrquez, Alessandro Moschitti, Preslav Nakov |
| 2018 | NAACL | Supervised and Unsupervised Transfer Learning for Question Answering. | Yu-An Chung, Hung-yi Lee, James R. Glass |
| 2018 | NAACL | Role-specific Language Models for Processing Recorded Neuropsychological Exams. | Tuka Al Hanai, Rhoda Au, James R. Glass |
| 2018 | NAACL | Automatic Stance Detection Using End-to-End Memory Networks. | Mitra Mohtarami, Ramy Baly, James R. Glass, Preslav Nakov, Llus Mrquez, Alessandro Moschitti |
| 2017 | ACL | What do Neural Machine Translation Models Learn about Morphology? | Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, James R. Glass |
| 2017 | ACL | Learning Word-Like Units from Joint Audio-Visual Analysis. | David Harwath, James R. Glass |
| 2017 | ASRU | Spoken language biomarkers for detecting cognitive impairment. | Tuka Alhanai, Rhoda Au, James R. Glass |
| 2017 | ASRU | Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentation. | Wei-Ning Hsu, Yu Zhang, James R. Glass |
| 2017 | ASRU | Learning modality-invariant representations for speech and images. | Kenneth Leidal, David Harwath, James R. Glass |
| 2017 | ASRU | Automatic speech recognition of Arabic multi-genre broadcast media. | Maryam Najafian, Wei-Ning Hsu, Ahmed Ali, James R. Glass |
| 2017 | ASRU | MIT-QCRI Arabic dialect identification system for the 2017 multi-genre broadcast challenge. | Suwon Shon, Ahmed Ali, James R. Glass |
| 2017 | ICASSP | Semantic mapping of natural language input to database entries via convolutional neural networks. | Mandy Korpusik, Zachary Collins, James R. Glass |
| 2017 | IJCNLP | Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks. | Yonatan Belinkov, Llus Mrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, James R. Glass |
| 2017 | Interspeech | An Environmental Feature Representation for Robust Speech Recognition and for Environment Identification. | Xue Feng, Brigitte Richardson, Scott Amman, James R. Glass |
| 2017 | Interspeech | Learning Latent Representations for Speech Generation and Transformation. | Wei-Ning Hsu, Yu Zhang, James R. Glass |
| 2017 | Interspeech | QMDIS: QCRI-MIT Advanced Dialect Identification System. | Sameer Khurana, Maryam Najafian, Ahmed Ali, Tuka Al Hanai, Yonatan Belinkov, James R. Glass |
| 2017 | Interspeech | Character-Based Embedding Models and Reranking Strategies for Understanding Natural Language Meal Descriptions. | Mandy Korpusik, Zachary Collins, James R. Glass |
| 2016 | COLING | Neural Attention for Learning to Rank Questions in Community Question Answering. | Salvatore Romeo, Giovanni Da San Martino, Alberto Barrn-Cedeo, Alessandro Moschitti, Yonatan Belinkov, Wei-Ning Hsu, Yu Zhang, Mitra Mohtarami, James R. Glass |
| 2016 | ICASSP | Multilingual data selection for training stacked bottleneck features. | Ekapol Chuangsuwanich, Yu Zhang, James R. Glass |
| 2016 | ICASSP | Distributional semantics for understanding spoken meal descriptions. | Mandy Korpusik, Calvin Huang, Michael Price, James R. Glass |
| 2016 | ICASSP | Personalized mispronunciation detection and diagnosis based on unsupervised error pattern discovery. | Ann Lee, Nancy F. Chen, James R. Glass |
| 2016 | ICASSP | Prediction-adaptation-correction recurrent neural networks for low-resource language speech recognition. | Yu Zhang, Ekapol Chuangsuwanich, James R. Glass, Dong Yu |
| 2016 | ICASSP | Highway long short-term memory RNNS for distant speech recognition. | Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yao, Sanjeev Khudanpur, James R. Glass |
| 2016 | Interspeech | Automatic Dialect Detection in Arabic Broadcast Speech. | Ahmed Ali, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James R. Glass, Peter Bell, Steve Renals |
| 2016 | Interspeech | Exploiting Depth and Highway Connections in Convolutional Recurrent Deep Neural Networks for Speech Recognition. | Wei-Ning Hsu, Yu Zhang, Ann Lee, James R. Glass |
| 2016 | Interspeech | Memory-Efficient Modeling and Search Techniques for Hardware ASR Decoders. | Michael Price, Anantha P. Chandrakasan, James R. Glass |
| 2015 | ASRU | Deep multimodal semantic embeddings for speech and images. | David F. Harwath, James R. Glass |
| 2015 | CHI | Wait-Learning: Leveraging Wait Time for Second Language Education. | Carrie J. Cai, Philip J. Guo, James R. Glass, Robert C. Miller |
| 2015 | EMNLP | Arabic Diacritization with Recurrent Neural Networks. | Yonatan Belinkov, James R. Glass |
| 2015 | ICASSP | On using heterogeneous data for vehicle-based speech recognition: A DNN-based approach. | Xue Feng, Brigitte Richardson, Scott Amman, James R. Glass |
| 2015 | Interspeech | Speaker adaptation using the i-vector technique for bottleneck features. | Patrick Cardinal, Najim Dehak, Yu Zhang, James R. Glass |
| 2015 | Interspeech | Mispronunciation detection without nonnative training data. | Ann Lee, James R. Glass |
| 2015 | NAACL | A Vector Space Approach for Aspect Based Sentiment Analysis. | Abdulaziz Alghunaim, Mitra Mohtarami, Scott Cyphers, James R. Glass |
| 2014 | CHI | Wait-learning: leveraging conversational dead time for second language education. | Carrie J. Cai, Philip J. Guo, James R. Glass, Robert C. Miller |
| 2014 | CogSci | One-shot learning of generative speech concepts. | Brenden M. Lake, Chia-ying Lee, James R. Glass, Joshua B. Tenenbaum |
| 2014 | COLING | A Study of using Syntactic and Semantic Structures for Concept Segmentation and Labeling. | Iman Saleh, Scott Cyphers, James R. Glass, Shafiq R. Joty, Llus Mrquez i Villodre, Alessandro Moschitti, Preslav Nakov |
| 2014 | ICASSP | Speech feature denoising and dereverberation via deep autoencoders for noisy reverberant speech recognition. | Xue Feng, Yaodong Zhang, James R. Glass |
| 2014 | ICASSP | Extracting deep neural network bottleneck features using low-rank matrix factorization. | Yu Zhang, Ekapol Chuangsuwanich, James R. Glass |
| 2014 | Interspeech | Recent advances in ASR applied to an Arabic transcription system for Al-Jazeera. | Patrick Cardinal, Ahmed Ali, Najim Dehak, Yu Zhang, Tuka Al Hanai, Yifan Zhang, James R. Glass, Stephan Vogel |
| 2014 | Interspeech | Language ID-based training of multilingual stacked bottleneck features. | Anne Cutler, Yu Zhang, Ekapol Chuangsuwanich, James R. Glass |
| 2014 | Interspeech | Lexical modeling for Arabic ASR: a systematic approach. | Tuka Al Hanai, James R. Glass |
| 2014 | Interspeech | Speech recognition without a lexicon - bridging the gap between graphemic and phonetic systems. | David F. Harwath, James R. Glass |
| 2014 | Interspeech | Context-dependent pronunciation error pattern discovery with limited annotations. | Ann Lee, James R. Glass |
| 2014 | Interspeech | Graph-based re-ranking using acoustic feature similarity between search results for spoken term detection on low-resource languages. | Hung-yi Lee, Yu Zhang, Ekapol Chuangsuwanich, James R. Glass |
| 2014 | Interspeech | Limited labels for unlimited data: active learning for speaker recognition. | Stephen H. Shum, Najim Dehak, James R. Glass |
| 2013 | ASRU | Query understanding enhanced by hierarchical parsing structures. | Jingjing Liu, Panupong Pasupat, Yining Wang, Scott Cyphers, James R. Glass |
| 2013 | EMNLP | Joint Learning of Phonetic Units and Word Pronunciations for ASR. | Chia-ying Lee, Yu Zhang, James R. Glass |
| 2013 | ICASSP | Zero resource spoken audio corpus analysis. | David F. Harwath, Timothy J. Hazen, James R. Glass |
| 2013 | ICASSP | Mispronunciation detection via dynamic time warping on deep belief network-based posteriorgrams. | Ann Lee, Yaodong Zhang, James R. Glass |
| 2013 | ICASSP | Asgard: A portable architecture for multilingual dialogue systems. | Jingjing Liu, Panupong Pasupat, Scott Cyphers, James R. Glass |
| 2013 | Interspeech | Bayesian distance metric learning on i-vector for speaker verification. | Xiao Fang, Najim Dehak, James R. Glass |
| 2012 | ACL | A Nonparametric Bayesian Approach to Acoustic Model Discovery. | Chia-ying Lee, James R. Glass |
| 2012 | ICASSP | Evaluation of multi-level context-dependent acoustic model for large vocabulary speaker adaptation tasks. | Hung-An Chang, James R. Glass |
| 2012 | ICASSP | Handling uncertain observations in unsupervised topic-mixture language model adaptation. | Ekapol Chuangsuwanich, Shinji Watanabe, Takaaki Hori, Tomoharu Iwata, James R. Glass |
| 2012 | ICASSP | Fast spoken query detection using lower-bound Dynamic Time Warping on Graphical Processing Units. | Yaodong Zhang, Kiarash Adl, James R. Glass |
| 2012 | ICASSP | Resource configurable spoken query detection using Deep Boltzmann Machines. | Yaodong Zhang, Ruslan Salakhutdinov, Hung-An Chang, James R. Glass |
| 2012 | Interspeech | Sentence Detection Using Multiple Annotations. | Ann Lee, James R. Glass |
| 2012 | Interspeech | A Conversational Movie Search System Based on Conditional Random Fields. | Jingjing Liu, Scott Cyphers, Panupong Pasupat, Ian McGraw, James R. Glass |
| 2012 | Interspeech | Automating Crowd-supervised Learning for Spoken Language Systems. | Ian McGraw, Scott Cyphers, Panupong Pasupat, Jingjing Liu, James R. Glass |
| 2011 | ASRU | Multi-level context-dependent acoustic modeling for automatic speech recognition. | Hung-An Chang, James R. Glass |
| 2011 | ICASSP | A channel-blind system for speaker verification. | Najim Dehak, Zahi N. Karam, Douglas A. Reynolds, Rda Dehak, William M. Campbell, James R. Glass |
| 2011 | ICASSP | An inner-product lower-bound estimate for dynamic time warping. | Yaodong Zhang, James R. Glass |
| 2011 | Interspeech | Pronunciation Learning from Continuous Speech. | Ibrahim Badr, Ian McGraw, James R. Glass |
| 2011 | Interspeech | Robust Voice Activity Detector for Real World Applications Using Harmonicity and Modulation Frequency. | Ekapol Chuangsuwanich, James R. Glass |
| 2011 | Interspeech | A Transcription Task for Crowdsourcing with Automatic Quality Control. | Chia-ying Lee, James R. Glass |
| 2011 | Interspeech | An Efferent-Inspired Auditory Model Front-End for Speech Recognition. | Chia-ying Lee, James R. Glass, Oded Ghitza |
| 2011 | Interspeech | Growing a Spoken Language Interface on Amazon Mechanical Turk. | Ian McGraw, James R. Glass, Stephanie Seneff |
| 2011 | Interspeech | Exploiting Intra-Conversation Variability for Speaker Diarization. | Stephen Shum, Najim Dehak, Ekapol Chuangsuwanich, Douglas A. Reynolds, James R. Glass |
| 2011 | Interspeech | A Piecewise Aggregate Approximation Lower-Bound Estimate for Posteriorgram-Based Dynamic Time Warping. | Yaodong Zhang, James R. Glass |
| 2010 | ICASSP | Towards multi-speaker unsupervised speech pattern discovery. | Yaodong Zhang, James R. Glass |
| 2010 | Interspeech | Learning new word pronunciations from spoken examples. | Ibrahim Badr, Ian McGraw, James R. Glass |
| 2010 | LREC | Collecting Voices from the Cloud. | Ian McGraw, Chia-ying Lee, I. Lee Hetherington, Stephanie Seneff, James R. Glass |
| 2009 | ASRU | Unsupervised spoken keyword spotting via segmental DTW on Gaussian posteriorgrams. | Yaodong Zhang, James R. Glass |
| 2009 | CHI | City browser: developing a conversational automotive HMI. | Alexander Gruenstein, Jarrod Orszulak, Sean Liu, Shannon C. Roberts, Jeff Zabel, Bryan Reimer, Bruce Mehler, Stephanie Seneff, James R. Glass, Joseph F. Coughlin |
| 2009 | EACL | Syntactic Phrase Reordering for English-to-Arabic Statistical Machine Translation. | Ibrahim Badr, Rabih Zbib, James R. Glass |
| 2009 | ICASSP | Discriminative training of hierarchical acoustic models for large vocabulary continuous speech recognition. | Hung-An Chang, James R. Glass |
| 2009 | ICASSP | Language model parameter estimation using user transcriptions. | Bo-June Paul Hsu, James R. Glass |
| 2009 | ICASSP | On the phonetic information in ultrasonic microphone signals. | Karen Livescu, Bo Zhu, James R. Glass |
| 2009 | ICASSP | Speech rhythm guided syllable nuclei detection. | Yaodong Zhang, James R. Glass |
| 2009 | Interspeech | A back-off discriminative acoustic model for automatic speech recognition. | Hung-An Chang, James R. Glass |
| 2008 | ACL | Segmentation for English-to-Arabic Statistical Machine Translation. | Ibrahim Badr, Rabih Zbib, James R. Glass |
| 2008 | EMNLP | N-gram Weighting: Reducing Training Data Mismatch in Cross-Domain Language Model Estimation. | Bo-June Paul Hsu, James R. Glass |
| 2008 | ICASSP | A turbo-style algorithm for lexical baseforms estimation. | Ghinwa F. Choueiter, Mesrob I. Ohannessian, Stephanie Seneff, James R. Glass |
| 2008 | Interspeech | Iterative language model estimation: efficient data structure & algorithms. | Bo-June Paul Hsu, James R. Glass |
| 2007 | ACL | Making Sense of Sound: Unsupervised Topic Segmentation over Acoustic Input. | Igor Malioutov, Alex Park, Regina Barzilay, James R. Glass |
| 2007 | ASRU | Hierarchical large-margin Gaussian mixture models for phonetic classification. | Hung-An Chang, James R. Glass |
| 2007 | ASRU | Automatic lexical pronunciations generation and update. | Ghinwa F. Choueiter, Stephanie Seneff, James R. Glass |
| 2007 | ASRU | Speech recognition with localized time-frequency pattern detectors. | Ken Schutte, James R. Glass |
| 2007 | ICASSP | Open-Vocabulary Spoken Utterance Retrieval using Confusion Networks. | Takaaki Hori, I. Lee Hetherington, Timothy J. Hazen, James R. Glass |
| 2007 | ICASSP | Noise Robust Phonetic Classificationwith Linear Regularized Least Squares and Second-Order Features. | Ryan Rifkin, Ken Schutte, Michelle Saad, Jake V. Bouvrie, James R. Glass |
| 2007 | Interspeech | New word acquisition using subword modeling. | Ghinwa F. Choueiter, Stephanie Seneff, James R. Glass |
| 2007 | Interspeech | Recent progress in the MIT spoken lecture processing project. | James R. Glass, Timothy J. Hazen, D. Scott Cyphers, Igor Malioutov, David Huynh, Regina Barzilay |
| 2007 | Interspeech | Multimodal speech recognition with ultrasonic sensors. | Bo Zhu, Timothy J. Hazen, James R. Glass |
| 2006 | EMNLP | Style & Topic Language Model Adaptation Using HMM-LDA. | Bo-June Paul Hsu, James R. Glass |
| 2006 | ICASSP | Flexible Multi-Stream Framework for Speech Recognition using Multi-Tape Finite-State Transducers. | I. Lee Hetherington, Han Shu, James R. Glass |
| 2006 | ICASSP | Speaker Verification Over Handheld Devices with Realistic Noisy Speech Data. | Ji Ming, Timothy J. Hazen, James R. Glass |
| 2006 | ICASSP | Unsupervised Word Acquisition from Speech using Pattern Discovery. | Alex Park, James R. Glass |
| 2006 | Interspeech | Combining missing-feature theory, speech enhancement and speaker-dependent/-independent modeling for speech separation. | Ji Ming, Timothy J. Hazen, James R. Glass |
| 2005 | ICASSP | A Wavelet and Filter Bank Framework For Phonetic Classification. | Ghinwa F. Choueiter, James R. Glass |
| 2005 | ICASSP | Automatic Processing of Audio Lectures for Information Retrieval: Vocabulary Selection and Language Modeling. | Alex Park, Timothy J. Hazen, James R. Glass |
| 2005 | ICASSP | Production domain modeling of pronunciation for visual speech recognition. | Kate Saenko, Karen Livescu, James R. Glass, Trevor Darrell |
| 2005 | ICCV | Visual Speech Recognition with Loosely Synchronized Feature Streams. | Kate Saenko, Karen Livescu, Michael Siracusa, Kevin W. Wilson, James R. Glass, Trevor Darrell |
| 2005 | Interspeech | Morphing spectral envelopes using audio flow. | Tony Ezzat, Ethan Meyers, James R. Glass, Tomaso A. Poggio |
| 2005 | Interspeech | Robust detection of sonorant landmarks. | Ken Schutte, James R. Glass |
| 2005 | NAACL | The MIT Spoken Lecture Processing Project. | James R. Glass, Timothy J. Hazen, D. Scott Cyphers, Ken Schutte, Alex Park |
| 2004 | ICMI | A segment-based audio-visual speech recognizer: data collection, development, and initial experiments. | Timothy J. Hazen, Kate Saenko, Chia-Hao La, James R. Glass |
| 2004 | ICMI | Articulatory features for robust visual speech recognition. | Kate Saenko, Trevor Darrell, James R. Glass |
| 2004 | Interspeech | Feature-based pronunciation modeling with trainable asynchrony probabilities. | Karen Livescu, James R. Glass |
| 2004 | NAACL | Feature-based Pronunciation Modeling for Speech Recognition. | Karen Livescu, James R. Glass |
| 2003 | Interspeech | Hidden feature models for speech recognition using dynamic Bayesian networks. | Karen Livescu, James R. Glass, Jeff A. Bilmes |
| 2002 | Interspeech | A multi-class approach for modelling out-of-vocabulary words. | Issam Bazzi, James R. Glass |
| 2002 | Interspeech | Information-theoretic criteria for unit selection synthesis. | Jon R. W. Yi, James R. Glass |
| 2001 | Interspeech | Learning units for domain-independent out-of- vocabulary word modelling. | Issam Bazzi, James R. Glass |
| 2001 | Interspeech | Speechbuilder: facilitating spoken dialogue system development. | James R. Glass, Eugene Weinstein |
| 2001 | Interspeech | Segment-based recognition on the phonebook task: initial results and observations on duration modeling. | Karen Livescu, James R. Glass |
| 2001 | Interspeech | Mokusei: a telephone-based Japanese conversational system in the weather domain. | Mikio Nakano, Yasuhiro Minami, Stephanie Seneff, Timothy J. Hazen, D. Scott Cyphers, James R. Glass, Joseph Polifroni, Victor Zue |
| 2000 | ICASSP | Heterogeneous lexical units for automatic speech recognition: preliminary investigations. | Issam Bazzi, James R. Glass |
| 2000 | ICASSP | Lexical modeling of non-native speech for automatic speech recognition. | Karen Livescu, James R. Glass |
| 2000 | Interspeech | Modeling out-of-vocabulary words for robust speech recognition. | Issam Bazzi, James R. Glass |
| 2000 | Interspeech | Data collection and performance evaluation of spoken dialogue systems: the MIT experience. | James R. Glass, Joseph Polifroni, Stephanie Seneff, Victor Zue |
| 2000 | Interspeech | A flexible, scalable finite-state transducer architecture for corpus-based concatenative speech synthesis. | Jon R. W. Yi, James R. Glass, I. Lee Hetherington |
| 1999 | ICASSP | Real-time telephone-based speech recognition in the Jupiter domain. | James R. Glass, Timothy J. Hazen, I. Lee Hetherington |
| 1998 | Interspeech | Telephone-based conversational speech recognition in the JUPITER domain. | James R. Glass, Timothy J. Hazen |
| 1998 | Interspeech | Heterogeneous measurements and multiple classifiers for speech recognition. | Andrew K. Halberstadt, James R. Glass |
| 1998 | Interspeech | Real-time probabilistic segmentation for segment-based speech recognition. | Steven C. Lee, James R. Glass |
| 1998 | Interspeech | Confidence scoring for speech understanding systems. | Christine Pao, Philipp Schmid, James R. Glass |
| 1998 | Interspeech | Natural-sounding speech synthesis using variable-length units. | Jon R. W. Yi, James R. Glass |
| 1998 | LREC | Evaluation methodology for a telephone-based conversational system. | Joseph Polifroni, Stephanie Seneff, James R. Glass, Timothy J. Hazen |
| 1997 | Interspeech | Segmentation and modeling in segment-based recognition. | Jane W. Chang, James R. Glass |
| 1997 | Interspeech | Heterogeneous acoustic measurements for phonetic classification 1. | Andrew K. Halberstadt, James R. Glass |
| 1997 | Interspeech | A comparison of novel techniques for instantaneous speaker adaptation. | Timothy J. Hazen, James R. Glass |
| 1997 | Interspeech | MUSE: a scripting language for the development of interactive speech analysis and recognition tools. | Michael K. McCandless, James R. Glass |
| 1997 | Interspeech | YINHE: a Mandarin Chinese version of the GALAXY system. | Chao Wang, James R. Glass, Helen M. Meng, Joseph Polifroni, Stephanie Seneff, Victor W. Zue |
| 1997 | Interspeech | From interface to content: translingual access and delivery of on-line information. | Victor W. Zue, Stephanie Seneff, James R. Glass, I. Lee Hetherington, Edward Hurley, Helen M. Meng, Christine Pao, Joseph Polifroni, Rafael Schloming, Philipp Schmid |
| 1996 | Interspeech | A probabilistic framework for feature-based speech recognition. | James R. Glass, Jane W. Chang, Michael K. McCandless |
| 1996 | Interspeech | Telephone data collection using the world wide web. | Edward Hurley, Joseph Polifroni, James R. Glass |
| 1996 | Interspeech | WHEELS: a conversational system in the automobile classifieds domain. | Helen M. Meng, Senis Busayapongchai, James R. Glass, David Goddeau, I. Lee Hetherington, Edward Hurley, Christine Pao, Joseph Polifroni, Stephanie Seneff, Victor Zue |
| 1996 | Interspeech | Multilingual human-computer interactions: from information access to language learning. | Victor Zue, Stephanie Seneff, Joseph Polifroni, Helen M. Meng, James R. Glass |
| 1994 | Interspeech | Porting the bilingual voyager system to Italian. | Giovanni Flammia, James R. Glass, Michael S. Phillips, Joseph Polifroni, Stephanie Seneff, Victor W. Zue |
| 1994 | Interspeech | Multilingual language generation across multiple domains. | James R. Glass, Joseph Polifroni, Stephanie Seneff |
| 1994 | Interspeech | GALAXY: a human-language interface to on-line travel information. | David Goddeau, Eric Brill, James R. Glass, Christine Pao, Michael S. Phillips, Joseph Polifroni, Stephanie Seneff, Victor W. Zue |
| 1994 | Interspeech | Statistical trajectory models for phonetic recognition. | William Goldenthal, James R. Glass |
| 1994 | Interspeech | Empirical acquisition of language models for speech recognition. | Michael K. McCandless, James R. Glass |
| 1993 | ICASSP | A comparative study of signal representations and classification techniques for speech recognition. | Hong C. Leung, Benjamin Chigier, James R. Glass |
| 1993 | Interspeech | A bilingual Voyager system. | James R. Glass, David Goodine, Michael S. Phillips, Shinsuke Sakai, Stephanie Seneff, Victor W. Zue |
| 1993 | Interspeech | Modelling spectral dynamics for vowel classification. | William Goldenthal, James R. Glass |
| 1993 | Interspeech | A* word network search for continuous speech recognition. | I. Lee Hetherington, Michael S. Phillips, James R. Glass, Victor W. Zue |
| 1993 | Interspeech | Empirical acquisition of word and phrase classes in the atis domain. | Michael K. McCandless, James R. Glass |
| 1992 | Interspeech | Vowel classification based on analysis-by-synthesis. | Rolf Carlson, James R. Glass |
| 1992 | Interspeech | Collection and analyses of WSJ-CSR corpus at MIT. | Michael S. Phillips, James R. Glass, Joseph Polifroni, Victor Zue |
| 1992 | NAACL | Collection and Analyses of WSJ-CSR Data at MIT. | Michael S. Phillips, James R. Glass, Joseph Polifroni, Victor Zue |
| 1991 | ICASSP | Integration of speech recognition and natural language processing in the MIT VOYAGER system. | Victor Zue, James R. Glass, David Goodine, Hong C. Leung, Michael S. Phillips, Joseph Polifroni, Stephanie Seneff |
| 1991 | Interspeech | Automatic learning of lexical representations for sub-word unit based speech recognition systems. | Michael S. Phillips, James R. Glass, Victor W. Zue |
| 1991 | Interspeech | The MIT ATIS system; preliminary development, spontaneous speech data collection, and performance evaluation. | Victor W. Zue, James R. Glass, David Goodine, Lynette Hirschman, Hong C. Leung, Michael S. Phillips, Joseph Polifroni, Stephanie Seneff |
| 1991 | NAACL | Modelling Context Dependency in Acoustic-Phonetic and Lexical Representations. | Michael S. Phillips, James R. Glass, Victor Zue |
| 1990 | ICASSP | The VOYAGER speech understanding system: preliminary development and evaluation. | Victor Zue, James R. Glass, David Goodine, Hong C. Leung, Michael S. Phillips, Joseph Polifroni, Stephanie Seneff |
| 1990 | ICASSP | The SUMMIT speech recognition system: phonological modelling and lexical access. | Victor Zue, James R. Glass, David Goodine, Michael Philips, Stephanie Seneff |
| 1990 | Interspeech | Detection and classification of phonemes using context-independent error back-propagation. | Hong C. Leung, James R. Glass, Michael S. Phillips, Victor W. Zue |
| 1990 | Interspeech | Recent progress on the MIT VOYAGER spoken language system. | Victor W. Zue, James R. Glass, Dave Goddeau, David Goodine, Hong C. Leung, Michael K. McCandless, Michael S. Phillips, Joseph Polifroni, Stephanie Seneff, Dave Whitney |
| 1989 | ICASSP | Acoustic segmentation and phonetic classification in the SUMMIT system. | Victor Zue, James R. Glass, Michael Philips, Stephanie Seneff |
| 1988 | ICASSP | Multi-level acoustic segmentation of continuous speech. | James R. Glass, Victor W. Zue |
| 1986 | ICASSP | Detection and recognition of nasal consonants in American English. | James R. Glass, Victor W. Zue |
| 1985 | ICASSP | Detection of nasalized vowels in American English. | James R. Glass, Victor W. Zue |