Skip to content

Kazuhito Koishida

Publication record assembled from the DBLP archive of ranked conferences.

Papers indexed

41

Venues

10

Active years

1994–2026

Best venue rank

A*

Where they publish

Papers

41 indexed papers, newest first.

YearVenueTitleAuthors
2026EACLDo GUI Grounders Truly Understand UI Elements?Surgan Jandial, Yinheng Li, Justin Wagle, Kazuhito Koishida
2025ACLWinSpot: GUI Grounding Benchmark with Multimodal Large Language Models.Zheng Hui, Yinheng Li, Dan Zhao, Colby R. Banbury, Tianyi Chen, Kazuhito Koishida
2025CVPRAutomatic Joint Structured Pruning and Quantization for Efficient Neural Network Training and Compression.Xiaoyi Qu, David Aponte, Colby R. Banbury, Daniel P. Robinson, Tianyu Ding, Kazuhito Koishida, Ilya Zharkov, Tianyi Chen
2025ICASSPCorrGAN: Simultaneous Learning of Speech Enhancement and Perceptual Quality Loss Functions.Vasily Zadorozhnyy, Saeed Amizadeh, Qiang Ye, Kazuhito Koishida
2025ICLRVideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks.Lawrence Keunho Jang, Yinheng Li, Dan Zhao, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, Kazuhito Koishida
2025ICMLWindows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Keunho Jang, Zheng Hui
2024ICASSPuaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures.Afrina Tabassum, Dung N. Tran, Trung Dang, Ismini Lourentzou, Kazuhito Koishida
2024ICIPLearned Image Compression With Text Quality Enhancement.Chih-Yu Lai, Dung N. Tran, Kazuhito Koishida
2024ICLRWeakly-supervised Audio Separation via Bi-modal Semantic Similarity.Tanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana Marculescu
2024InterspeechLiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.Trung Dang, David Aponte, Dung N. Tran, Kazuhito Koishida
2024InterspeechConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation.Yatong Bai, Trung Dang, Dung N. Tran, Kazuhito Koishida, Somayeh Sojoudi
2023ICASSPToward A Multimodal Approach for Disfluency Detection and Categorization.Amrit Romana, Kazuhito Koishida
2023InterspeechSCP-GAN: Self-Correcting Discriminator Optimization for Training Consistency Preserving Metric GAN on Speech Enhancement Tasks.Vasily Zadorozhnyy, Qiang Ye, Kazuhito Koishida
2022ICASSPTraining Robust Zero-Shot Voice Conversion Models with Self-Supervised Features.Trung Dang, Dung N. Tran, Peter Chin, Kazuhito Koishida
2022ICASSPA Training Framework for Stereo-Aware Speech Enhancement Using Deep Neural Networks.Bahareh Tolooshams, Kazuhito Koishida
2021ICASSPCascaded Time + Time-Frequency Unet For Speech Enhancement: Jointly Addressing Clipping, Codec Distortions, And Gaps.Arun Asokan Nair, Kazuhito Koishida
2021InterspeechSingle-Channel Speech Enhancement Using Learnable Loss Mixup.Oscar Chang, Dung N. Tran, Kazuhito Koishida
2021InterspeechINTERSPEECH 2021 Deep Noise Suppression Challenge.Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan
2020CVPRImproved Active Speaker Detection based on Optical Flow.Chong Huang, Kazuhito Koishida
2020CVPRMMTM: Multimodal Transfer Module for CNN Fusion.Hamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, Kazuhito Koishida
2020ICASSPLow-Latency Single Channel Speech Enhancement Using U-Net Convolutional Neural Networks.Ahmet Emin Bulut, Kazuhito Koishida
2020ICASSPAV(SE)Michael L. Iuzzolino, Kazuhito Koishida
2020ICASSPGeometrically Constrained Independent Vector Analysis for Directional Speech Enhancement.Li Li, Kazuhito Koishida
2020ICMLNeuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning".Saeed Amizadeh, Hamid Palangi, Alex Polozov, Yichen Huang, Kazuhito Koishida
2020InterspeechOnline Directional Speech Enhancement Using Geometrically Constrained Independent Vector Analysis.Li Li, Kazuhito Koishida, Shoji Makino
2020InterspeechLow-Latency Single Channel Speech Dereverberation Using U-Net Convolutional Neural Networks.Ahmet Emin Bulut, Kazuhito Koishida
2020InterspeechRobust Pitch Regression with Voiced/Unvoiced Classification in Nonstationary Noise Environments.Dung N. Tran, Uros Batricevic, Kazuhito Koishida
2020InterspeechSingle-Channel Speech Enhancement by Subspace Affinity Minimization.Dung N. Tran, Kazuhito Koishida
2019ICASSPSpeech Super Resolution Generative Adversarial Network.Sefik Emre Eskimez, Kazuhito Koishida
2019InterspeechSound Event Detection in Multichannel Audio Using Convolutional Time-Frequency-Channel Squeeze and Excitation.Wei Xia, Kazuhito Koishida
2017ASRUEnd-to-end text-independent speaker verification with flexibility in utterance duration.Chunlei Zhang, Kazuhito Koishida
2017InterspeechEnd-to-End Text-Independent Speaker Verification with Triplet Loss on Short Utterances.Chunlei Zhang, Kazuhito Koishida
2008MMSPHybrid low bitrate audio coding using adaptive gain shape vector quantization.Sanjeev Mehrotra, Wei-Ge Chen, Kazuhito Koishida, Naveen Thumpudi
2000ICASSPA 16-kbit/s bandwidth scalable audio coder based on the G.729 standard.Kazuhito Koishida, Vladimir Cuperman, Allen Gersho
2000ICASSPA 1200 bps speech coder based on MELP.Tian Wang, Kazuhito Koishida, Vladimir Cuperman, Allen Gersho, John S. Collura
1998ICASSPA wideband CELP speech coder at 16 kbit/s based on mel-generalized cepstral analysis.Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi
1998InterspeechA 16 kbit/s wideband CELP coder using MEL-generalized cepstral analysis and its subjective evaluation.Kazuhito Koishida, Gou Hirabayashi, Keiichi Tokuda, Takao Kobayashi
1997ICASSPEfficient encoding of mel-generalized cepstrum for CELP coders.Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai
1996InterspeechCELP coding system based on mel-generalized cepstral analysis.Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai
1995ICASSPCELP coding based on mel-cepstral analysis.Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai
1994InterspeechSpeech coding based on adaptive MEL-cepstral analysis for noisy channels.Kazuhito Koishida, Keiichi Tokuda, Takao Kobayashi, Satoshi Imai