| 2026 | AAAI | Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding. | Rui-Chen Zheng, Wenrui Liu, Hui-Peng Du, Qinglin Zhang, Chong Deng, Qian Chen, Wen Wang, Yang Ai, Zhen-Hua Ling |
| 2026 | ACL | UniVocal: Unified Speech-Singing Code-Switching Synthesis. | Yufei Shi, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling, Yang Ai |
| 2025 | ICASSP | Recursive Feature Learning from Pre-Trained Models for Spoofing Speech Detection. | Yu Guan, Yang Ai, Zuoliang Li, Shengyu Peng, Wu Guo |
| 2025 | ICASSP | CASC-XVC: Zero-Shot Cross-Lingual Voice Conversion with Content Accordant and Speaker Contrastive Losses. | Han-Jie Guo, Hui-Peng Du, Zheng-Yan Sheng, Li-Ping Chen, Yang Ai, Zhen-Hua Ling |
| 2025 | ICASSP | Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations. | Zuoliang Li, Yang Ai, Jie Zhang, Shengyu Peng, Yu Guan, Bin Gu, Wu Guo |
| 2025 | ICASSP | Can Automated Speech Recognition Errors Provide Valuable Clues for Alzheimer's Disease Detection? | Yin-Long Liu, Rui Feng, Ye-Xin Lu, Jia-Xin Chen, Yang Ai, Jia-Hong Yuan, Zhen-Hua Ling |
| 2025 | ICASSP | Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis. | Ye-Xin Lu, Hui-Peng Du, Zheng-Yan Sheng, Yang Ai, Zhen-Hua Ling |
| 2025 | ICASSP | A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification. | Shengyu Peng, Wu Guo, Jie Zhang, Zuoliang Li, Yu Guan, Bin Gu, Yang Ai |
| 2025 | Interspeech | Vision-Integrated High-Quality Neural Speech Coding. | Yao Guo, Yang Ai, Rui-Chen Zheng, Hui-Peng Du, Xiao-Hang Jiang, Zhen-Hua Ling |
| 2025 | Interspeech | Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising. | Ye-Xin Lu, Hui-Peng Du, Fei Liu, Yang Ai, Zhen-Hua Ling |
| 2025 | Interspeech | Universal Preference-Score-based Pairwise Speech Quality Assessment. | Yufei Shi, Yang Ai, Zhen-Hua Ling |
| 2024 | ICASSP | Considering Temporal Connection between Turns for Conversational Speech Synthesis. | Kangdi Mei, Zhaoci Liu, Hui-Peng Du, Hengyu Li, Yang Ai, Liping Chen, Zhenhua Ling |
| 2024 | Interspeech | A Low-Bitrate Neural Audio Codec Framework with Bandwidth Reduction and Recovery for High-Sampling-Rate Waveforms. | Yang Ai, Ye-Xin Lu, Xiao-Hang Jiang, Zheng-Yan Sheng, Rui-Chen Zheng, Zhen-Hua Ling |
| 2024 | Interspeech | BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation. | Hui-Peng Du, Ye-Xin Lu, Yang Ai, Zhen-Hua Ling |
| 2024 | Interspeech | Refining Self-supervised Learnt Speech Representation using Brain Activations. | Hengyu Li, Kangdi Mei, Zhaoci Liu, Yang Ai, Liping Chen, Jie Zhang, Zhenhua Ling |
| 2024 | Interspeech | MultiStage Speech Bandwidth Extension with Flexible Sampling Rate Control. | Ye-Xin Lu, Yang Ai, Zheng-Yan Sheng, Zhen-Hua Ling |
| 2023 | ICASSP | Neural Speech Phase Prediction Based on Parallel Estimation Architecture and Anti-Wrapping Losses. | Yang Ai, Zhen-Hua Ling |
| 2023 | ICASSP | Zero-Shot Personalized Lip-To-Speech Synthesis with Face Image Based Voice Control. | Zhengyan Sheng, Yang Ai, Zhen-Hua Ling |
| 2023 | ICASSP | Speech Reconstruction from Silent Tongue and Lip Articulation by Pseudo Target Generation and Domain Adversarial Training. | Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling |
| 2023 | Interspeech | MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra. | Ye-Xin Lu, Yang Ai, Zhen-Hua Ling |
| 2023 | Interspeech | Incorporating Ultrasound Tongue Images for Audio-Visual Speech Enhancement through Knowledge Distillation. | Rui-Chen Zheng, Yang Ai, Zhen-Hua Ling |
| 2020 | Interspeech | Knowledge-and-Data-Driven Amplitude Spectrum Prediction for Hierarchical Neural Vocoders. | Yang Ai, Zhen-Hua Ling |
| 2020 | Interspeech | Reverberation Modeling for Source-Filter-Based Neural Vocoder. | Yang Ai, Xin Wang, Junichi Yamagishi, Zhen-Hua Ling |
| 2019 | ICASSP | Dnn-based Spectral Enhancement for Neural Waveform Generators with Low-bit Quantization. | Yang Ai, Jing-Xuan Zhang, Liang Chen, Zhen-Hua Ling |
| 2019 | Interspeech | Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling. | Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, Li-Rong Dai |
| 2018 | ICASSP | Samplernn-Based Neural Vocoder for Statistical Parametric Speech Synthesis. | Yang Ai, Hong-Chuan Wu, Zhen-Hua Ling |