| 2026 | AAAI | Multi-Metric Preference Alignment for Generative Speech Restoration. | Junan Zhang, Xueyao Zhang, Jing Yang, Yuancheng Wang, Fan Fan, Zhizheng Wu |
| 2026 | ACL | MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora. | Tao Feng, Yuxiang Wang, Yuancheng Wang, Xueyao Zhang, Dekun Chen, Chaoren Wang, Xun Guan, Zhizheng Wu |
| 2026 | ACL | Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration. | Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank |
| 2026 | ACL | Closing the Modality Reasoning Gap for Speech Large Language Models. | Chaoren Wang, Heng Lu, Xueyao Zhang, Shujie Liu, Yan Lu, Jinyu Li, Zhizheng Wu |
| 2025 | ACL | Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment. | Xueyao Zhang, Yuancheng Wang, Chaoren Wang, Ziniu Li, Zhuo Chen, Zhizheng Wu |
| 2025 | ICASSP | Data-Driven White Noise Gain Constrained Robust Superdirective Beamformer for Speech Enhancement. | Hanchen Pei, Gongping Huang, Jilu Jin, Jianbo Ma, Zhizheng Wu, Jingdong Chen, Jacob Benesty |
| 2025 | ICASSP | PicoAudio: Enabling Precise Temporal Controllability in Text-to-Audio Generation. | Zeyu Xie, Xuenan Xu, Zhizheng Wu, Mengyue Wu |
| 2025 | ICASSP | AudioTime: A Temporally-aligned Audio-text Benchmark Dataset. | Zeyu Xie, Xuenan Xu, Zhizheng Wu, Mengyue Wu |
| 2025 | ICLR | MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer. | Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, Zhizheng Wu |
| 2025 | ICLR | LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models. | Junyan Ye, Baichuan Zhou, Zilong Huang, Junan Zhang, Tianyi Bai, Hengrui Kang, Jun He, Honglin Lin, Zihao Wang, Tong Wu, Zhizheng Wu, Yiping Chen, Dahua Lin, Conghui He, Weijia Li |
| 2025 | ICLR | Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement. | Xueyao Zhang, Xiaohui Zhang, Kainan Peng, Zhenyu Tang, Vimal Manohar, Yingru Liu, Jeff Hwang, Dangna Li, Yuhao Wang, Julian Chan, Yuan Huang, Zhizheng Wu, Mingbo Ma |
| 2025 | IJCNLP | Swallowing the Poison Pills: Insights from Vulnerability Disparity Among LLMs. | Yifeng Peng, Zhizheng Wu, Chen Chen |
| 2025 | Interspeech | Neurodyne: Neural Pitch Manipulation with Representation Learning and Cycle-Consistency GAN. | Yicheng Gu, Chaoren Wang, Zhizheng Wu, Lauri Juvela |
| 2025 | Interspeech | DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation. | Jiaqi Li, Xiaolong Lin, Zhekai Li, Shixi Huang, Yuancheng Wang, Chaoren Wang, Zhenpeng Zhan, Zhizheng Wu |
| 2024 | ICASSP | Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity Vocoder. | Yicheng Gu, Xueyao Zhang, Liumeng Xue, Zhizheng Wu |
| 2024 | ICASSP | An Initial Investigation of Neural Replay Simulator for Over-The-Air Adversarial Perturbations to Automatic Speaker Verification. | Jiaqi Li, Li Wang, Liumeng Xue, Lei Wang, Zhizheng Wu |
| 2024 | ICASSP | ADVSV: An Over-the-Air Adversarial Attack Dataset for Speaker Verification. | Li Wang, Jiaqi Li, Yuhao Luo, Jiahao Zheng, Lei Wang, Hao Li, Ke Xu, Chengfang Fang, Jie Shi, Zhizheng Wu |
| 2024 | ICML | NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models. | Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Eric Liu, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu, Tao Qin, Xiangyang Li, Wei Ye, Shikun Zhang, Jiang Bian, Lei He, Jinyu Li, Sheng Zhao |
| 2023 | Interspeech | PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction Network. | Qinghua Liu, Meng Ge, Zhizheng Wu, Haizhou Li |
| 2021 | Interspeech | Cross-Lingual Voice Conversion with a Cycle Consistency Loss on Linguistic Representation. | Yi Zhou, Xiaohai Tian, Zhizheng Wu, Haizhou Li |
| 2019 | Interspeech | Building a Mixed-Lingual Neural TTS System with Only Monolingual Data. | Liumeng Xue, Wei Song, Guanghui Xu, Lei Xie, Zhizheng Wu |
| 2017 | Interspeech | Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System. | Tim Capes, Paul Coles, Alistair Conkie, Ladan Golipour, Abie Hadjitarkhani, Qiong Hu, Nancy Huddleston, Melvyn Hunt, Jiangchuan Li, Matthias Neeracher, Kishore Prahallad, Tuomo Raitio, Ramya Rasipuram, Greg Townsend, Becci Williamson, David Winarsky, Zhizheng Wu, Hepeng Zhang |
| 2016 | ICASSP | Robust TTS duration modelling using DNNS. | Gustav Eje Henter, Srikanth Ronanki, Oliver Watts, Mirjam Wester, Zhizheng Wu, Simon King |
| 2016 | ICASSP | Deep neural network-guided unit selection synthesis. | Thomas Merritt, Robert A. J. Clark, Zhizheng Wu, Junichi Yamagishi, Simon King |
| 2016 | ICASSP | Spoofing detection from a feature representation perspective. | Xiaohai Tian, Zhizheng Wu, Xiong Xiao, Eng Siong Chng, Haizhou Li |
| 2016 | ICASSP | From HMMS to DNNS: Where do the improvements come from? | Oliver Watts, Gustav Eje Henter, Thomas Merritt, Zhizheng Wu, Simon King |
| 2016 | ICASSP | Investigating gated recurrent networks for speech synthesis. | Zhizheng Wu, Simon King |
| 2016 | Interspeech | GlottDNN - A Full-Band Glottal Vocoder for Statistical Parametric Speech Synthesis. | Manu Airaksinen, Bajibabu Bollepalli, Lauri Juvela, Zhizheng Wu, Simon King, Paavo Alku |
| 2016 | Interspeech | Waveform Generation Based on Signal Reshaping for Statistical Parametric Speech Synthesis. | Felipe Espic, Cassia Valentini-Botinhao, Zhizheng Wu, Simon King |
| 2016 | Interspeech | A Template-Based Approach for Speech Synthesis Intonation Generation Using LSTMs. | Srikanth Ronanki, Gustav Eje Henter, Zhizheng Wu, Simon King |
| 2016 | Interspeech | An Investigation of Spoofing Speech Detection Under Additive Noise and Reverberant Conditions. | Xiaohai Tian, Zhizheng Wu, Xiong Xiao, Eng Siong Chng, Haizhou Li |
| 2016 | Interspeech | The Voice Conversion Challenge 2016. | Tomoki Toda, Ling-Hui Chen, Daisuke Saito, Fernando Villavicencio, Mirjam Wester, Zhizheng Wu, Junichi Yamagishi |
| 2016 | Interspeech | Analysis of the Voice Conversion Challenge 2016 Evaluation Results. | Mirjam Wester, Zhizheng Wu, Junichi Yamagishi |
| 2016 | ISCAS | Efficient architecture for soft-output massive MIMO detection with Gauss-Seidel method. | Zhizheng Wu, Chuan Zhang, Ye Xue, Shugong Xu, Xiaohu You |
| 2015 | ICASSP | Sparse representation for frequency warping based voice conversion. | Xiaohai Tian, Zhizheng Wu, Siu Wa Lee, Nguyen Quy Hy, Engsiong Chng, Minghui Dong |
| 2015 | ICASSP | SAS: A speaker verification spoofing database containing diverse attacks. | Zhizheng Wu, Ali Khodabakhsh, Cenk Demiroglu, Junichi Yamagishi, Daisuke Saito, Tomoki Toda, Simon King |
| 2015 | ICASSP | Deep neural networks employing Multi-Task Learning and stacked bottleneck features for speech synthesis. | Zhizheng Wu, Cassia Valentini-Botinhao, Oliver Watts, Simon King |
| 2015 | Interspeech | Fusion of multiple parameterisations for DNN-based sinusoidal speech synthesis with multi-task learning. | Qiong Hu, Zhizheng Wu, Korin Richmond, Junichi Yamagishi, Yannis Stylianou, Ranniery Maia |
| 2015 | Interspeech | Deep neural network context embeddings for model selection in rich-context HMM synthesis. | Thomas Merritt, Junichi Yamagishi, Zhizheng Wu, Oliver Watts, Simon King |
| 2015 | Interspeech | System fusion for high-performance voice conversion. | Xiaohai Tian, Zhizheng Wu, Siu Wa Lee, Nguyen Quy Hy, Minghui Dong, Engsiong Chng |
| 2015 | Interspeech | Towards minimum perceptual error training for DNN-based speech synthesis. | Cassia Valentini-Botinhao, Zhizheng Wu, Simon King |
| 2015 | Interspeech | Sentence-level control vectors for deep neural network speech synthesis. | Oliver Watts, Zhizheng Wu, Simon King |
| 2015 | Interspeech | Human vs machine spoofing detection on wideband and narrowband data. | Mirjam Wester, Zhizheng Wu, Junichi Yamagishi |
| 2015 | Interspeech | Minimum trajectory error training for deep neural networks, combined with stacked bottleneck features. | Zhizheng Wu, Simon King |
| 2015 | Interspeech | Automatic speaker verification spoofing and countermeasures (ASVspoof 2015): introductory talk by the organizers. | Zhizheng Wu, Tomi Kinnunen |
| 2015 | Interspeech | ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge. | Zhizheng Wu, Tomi Kinnunen, Nicholas W. D. Evans, Junichi Yamagishi, Cemal Hanili, Md. Sahidullah, Aleksandr Sizov |
| 2015 | Interspeech | A study of speaker adaptation for DNN-based speech synthesis. | Zhizheng Wu, Pawel Swietojanski, Christophe Veaux, Steve Renals, Simon King |
| 2014 | Interspeech | Introducing i-vectors for joint anti-spoofing and speaker verification. | Elie Khoury, Tomi Kinnunen, Aleksandr Sizov, Zhizheng Wu, Sbastien Marcel |
| 2014 | Interspeech | A comparative study of spectral transformation techniques for singing voice synthesis. | Siu Wa Lee, Zhizheng Wu, Minghui Dong, Xiaohai Tian, Haizhou Li |
| 2014 | Interspeech | Joint nonnegative matrix factorization for exemplar-based voice conversion. | Zhizheng Wu, Chng Eng Siong, Haizhou Li |
| 2013 | ICASSP | Synthetic speech detection using temporal modulation feature. | Zhizheng Wu, Xiong Xiao, Engsiong Chng, Haizhou Li |
| 2013 | Interspeech | Vulnerability evaluation of speaker verification under voice conversion spoofing: the effect of text constraints. | Zhizheng Wu, Anthony Larcher, Kong-Aik Lee, Engsiong Chng, Tomi Kinnunen, Haizhou Li |
| 2013 | Interspeech | Exemplar-based unit selection for voice conversion utilizing temporal information. | Zhizheng Wu, Tuomas Virtanen, Tomi Kinnunen, Engsiong Chng, Haizhou Li |
| 2012 | ICASSP | Vulnerability of speaker verification systems against voice conversion spoofing attacks: The case of telephone speech. | Tomi Kinnunen, Zhizheng Wu, Kong-Aik Lee, Filip Sedlak, Engsiong Chng, Haizhou Li |
| 2012 | Interspeech | Detecting Converted Speech and Natural Speech for anti-Spoofing Attack in Speaker Recognition. | Zhizheng Wu, Chng Eng Siong, Haizhou Li |
| 2010 | Interspeech | Text-independent F0 transformation with non-parallel data for voice conversion. | Zhizheng Wu, Tomi Kinnunen, Engsiong Chng, Haizhou Li |
| 2009 | ICASSP | Improved prosody generation by maximizing joint likelihood of state and longer units. | Yao Qian, Zhizheng Wu, Frank K. Soong |
| 2009 | Interspeech | A minimum v/u error approach to F0 generation in HMM-based TTS. | Yao Qian, Frank K. Soong, Miaomiao Wang, Zhizheng Wu |
| 2008 | Interspeech | Duration refinement by jointly optimizing state and longer unit likelihood. | Boyang Gao, Yao Qian, Zhizheng Wu, Frank K. Soong |