| 2024 | ICASSP | Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis. | Xueyuan Chen, Xi Wang, Shaofei Zhang, Lei He, Zhiyong Wu, Xixin Wu, Helen Meng |
| 2023 | Interspeech | Large-Scale Automatic Audiobook Creation. | Brendan Walsh, Mark Hamilton, Greg Newby, Xi Wang, Serena Ruan, Sheng Zhao, Lei He, Shaofei Zhang, Eric Dettinger, William T. Freeman, Markus Weimer |
| 2023 | Interspeech | ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading. | Yujia Xiao, Shaofei Zhang, Xi Wang, Xu Tan, Lei He, Sheng Zhao, Frank K. Soong, Tan Lee |
| 2022 | Interspeech | Self-supervised Context-aware Style Representation for Expressive Speech Synthesis. | Yihan Wu, Xi Wang, Shaofei Zhang, Lei He, Ruihua Song, Jian-Yun Nie |
| 2016 | ICASSP | Exemplar-based sparse representation of timbre and prosody for voice conversion. | Huaiping Ming, Dong-Yan Huang, Lei Xie, Shaofei Zhang, Minghui Dong, Haizhou Li |
| 2015 | ACII | Fundamental frequency modeling using wavelets for emotional voice conversion. | Huaiping Ming, Dong-Yan Huang, Minghui Dong, Haizhou Li, Lei Xie, Shaofei Zhang |
| 2015 | Interspeech | Regularized non-negative matrix factorization using alternating direction method of multipliers and its application to source separation. | Shaofei Zhang, Dong-Yan Huang, Lei Xie, Engsiong Chng, Haizhou Li, Minghui Dong |