| 2026 | EACL | Do GUI Grounders Truly Understand UI Elements? | Surgan Jandial, Yinheng Li, Justin Wagle, Kazuhito Koishida |
| 2025 | ACL | WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models. | Zheng Hui, Yinheng Li, Dan Zhao, Colby R. Banbury, Tianyi Chen, Kazuhito Koishida |
| 2025 | CVPR | Automatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression. | Xiaoyi Qu, David Aponte, Colby R. Banbury, Daniel P. Robinson, Tianyu Ding, Kazuhito Koishida, Ilya Zharkov, Tianyi Chen |
| 2025 | ICASSP | CorrGAN: Simultaneous Learning of Speech Enhancement and Perceptual Quality Loss Functions. | Vasily Zadorozhnyy, Saeed Amizadeh, Qiang Ye, Kazuhito Koishida |
| 2025 | ICLR | VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks. | Lawrence Keunho Jang, Yinheng Li, Dan Zhao, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, Kazuhito Koishida |
| 2025 | ICML | Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale. | Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Keunho Jang, Zheng Hui |
| 2024 | ICASSP | uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures. | Afrina Tabassum, Dung N. Tran, Trung Dang, Ismini Lourentzou, Kazuhito Koishida |
| 2024 | ICIP | Learned Image Compression With Text Quality Enhancement. | Chih-Yu Lai, Dung N. Tran, Kazuhito Koishida |
| 2024 | ICLR | Weakly-supervised Audio Separation via Bi-modal Semantic Similarity. | Tanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana Marculescu |
| 2024 | Interspeech | LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes. | Trung Dang, David Aponte, Dung N. Tran, Kazuhito Koishida |
| 2024 | Interspeech | ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation. | Yatong Bai, Trung Dang, Dung N. Tran, Kazuhito Koishida, Somayeh Sojoudi |
| 2023 | ICASSP | Toward A Multimodal Approach for Disfluency Detection and Categorization. | Amrit Romana, Kazuhito Koishida |
| 2023 | Interspeech | SCP-GAN: Self-Correcting Discriminator Optimization for Training Consistency Preserving Metric GAN on Speech Enhancement Tasks. | Vasily Zadorozhnyy, Qiang Ye, Kazuhito Koishida |
| 2022 | ICASSP | Training Robust Zero-Shot Voice Conversion Models with Self-Supervised Features. | Trung Dang, Dung N. Tran, Peter Chin, Kazuhito Koishida |
| 2022 | ICASSP | A Training Framework for Stereo-Aware Speech Enhancement Using Deep Neural Networks. | Bahareh Tolooshams, Kazuhito Koishida |
| 2021 | ICASSP | Cascaded Time + Time-Frequency Unet For Speech Enhancement: Jointly Addressing Clipping, Codec Distortions, And Gaps. | Arun Asokan Nair, Kazuhito Koishida |
| 2021 | Interspeech | Single-Channel Speech Enhancement Using Learnable Loss Mixup. | Oscar Chang, Dung N. Tran, Kazuhito Koishida |
| 2021 | Interspeech | INTERSPEECH 2021 Deep Noise Suppression Challenge. | Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan |
| 2020 | CVPR | Improved Active Speaker Detection based on Optical Flow. | Chong Huang, Kazuhito Koishida |
| 2020 | CVPR | MMTM: Multimodal Transfer Module for CNN Fusion. | Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, Kazuhito Koishida |
| 2020 | ICASSP | Low-Latency Single Channel Speech Enhancement Using U-Net Convolutional Neural Networks. | Ahmet Emin Bulut, Kazuhito Koishida |
| 2020 | ICASSP | AV(SE) | Michael L. Iuzzolino, Kazuhito Koishida |
| 2020 | ICASSP | Geometrically Constrained Independent Vector Analysis for Directional Speech Enhancement. | Li Li, Kazuhito Koishida |
| 2020 | ICML | Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning". | Saeed Amizadeh, Hamid Palangi, Alex Polozov, Yichen Huang, Kazuhito Koishida |
| 2020 | Interspeech | Online Directional Speech Enhancement Using Geometrically Constrained Independent Vector Analysis. | Li Li, Kazuhito Koishida, Shoji Makino |
| 2020 | Interspeech | Low-Latency Single Channel Speech Dereverberation Using U-Net Convolutional Neural Networks. | Ahmet Emin Bulut, Kazuhito Koishida |
| 2020 | Interspeech | Robust Pitch Regression with Voiced/Unvoiced Classification in Nonstationary Noise Environments. | Dung N. Tran, Uros Batricevic, Kazuhito Koishida |
| 2020 | Interspeech | Single-Channel Speech Enhancement by Subspace Affinity Minimization. | Dung N. Tran, Kazuhito Koishida |
| 2019 | ICASSP | Speech Super Resolution Generative Adversarial Network. | Sefik Emre Eskimez, Kazuhito Koishida |
| 2019 | Interspeech | Sound Event Detection in Multichannel Audio Using Convolutional Time-Frequency-Channel Squeeze and Excitation. | Wei Xia, Kazuhito Koishida |
| 2017 | ASRU | End-to-end text-independent speaker verification with flexibility in utterance duration. | Chunlei Zhang, Kazuhito Koishida |
| 2017 | Interspeech | End-to-End Text-Independent Speaker Verification with Triplet Loss on Short Utterances. | Chunlei Zhang, Kazuhito Koishida |
| 2008 | MMSP | Hybrid low bitrate audio coding using adaptive gain shape vector quantization. | Sanjeev Mehrotra, Wei-Ge Chen, Kazuhito Koishida, Naveen Thumpudi |
| 2000 | ICASSP | A 16-kbit/s bandwidth scalable audio coder based on the G.729 standard. | Kazuhito Koishida, Vladimir Cuperman, Allen Gersho |
| 2000 | ICASSP | A 1200 bps speech coder based on MELP. | Tian Wang, Kazuhito Koishida, Vladimir Cuperman, Allen Gersho, John S. Collura |
| 1998 | ICASSP | A wideband CELP speech coder at 16 kbit/s based on mel-generalized cepstral analysis. | Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi |
| 1998 | Interspeech | A 16 kbit/s wideband CELP coder using MEL-generalized cepstral analysis and its subjective evaluation. | Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi |
| 1997 | ICASSP | Efficient encoding of mel-generalized cepstrum for CELP coders. | Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
| 1996 | Interspeech | CELP coding system based on mel-generalized cepstral analysis. | Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
| 1995 | ICASSP | CELP coding based on mel-cepstral analysis. | Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |
| 1994 | Interspeech | Speech coding based on adaptive MEL-cepstral analysis for noisy channels. | Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai |