| 2025 | ASRU | Sinba: Singing-To-Accompaniment Generation With Pitch Guidance Via Mamba-Based Language Model. | Jianwei Cui, Shihao Chen, Yu Gu, Jie Zhang, Liping Chen, Na Li, Chengxing Li, Shan Yang, Li-Rong Dai |
| 2025 | ICASSP | Trusted Mamba Contrastive Network for Multi-View Clustering. | Jian Zhu, Xin Zou, Lei Liu, Zhangmin Huang, Ying Zhang, Chang Tang, Li-Rong Dai |
| 2025 | ICASSP | Dynamic SRM Curriculum for Trustworthy Multi-modal Classification. | Jian Zhu, Cui Yu, Xin Zou, Zhangmin Huang, Chenshu Hu, Jun Sun, Bo Lyu, Lei Liu, Chang Tang, Li-Rong Dai |
| 2024 | ICASSP | Adaptive Confidence Multi-View Hashing for Multimedia Retrieval. | Jian Zhu, Yu Cui, Zhangmin Huang, Xingyu Li, Lei Liu, Lingfang Zeng, Li-Rong Dai |
| 2024 | ICDM | Adaptive Loss-aware Modulation for Multimedia Retrieval. | Jian Zhu, Yu Cui, Lei Liu, Zeyi Sun, Yuyang Dai, Xi Wang, Cheng Luo, Li-Rong Dai |
| 2023 | ICASSP | Stargan-vc Based Cross-Domain Data Augmentation for Speaker Verification. | Hang-Rui Hu, Yan Song, Jian-Tao Zhang, Li-Rong Dai, Ian McLoughlin, Zhu Zhuo, Yu Zhou, Yu-Hong Li, Hui Xue |
| 2023 | ICASSP | AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer. | Kang Li, Yan Song, Li-Rong Dai, Ian McLoughlin, Xin Fang, Lin Liu |
| 2023 | ICASSP | A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement. | Haitao Xu, Liangfa Wei, Jie Zhang, Jianming Yang, Yannan Wang, Tian Gao, Xin Fang, Li-Rong Dai |
| 2023 | ICASSP | Joint Generative-Contrastive Representation Learning for Anomalous Sound Detection. | Xiao-Min Zeng, Yan Song, Zhu Zhuo, Yu Zhou, Yu-Hong Li, Hui Xue, Li-Rong Dai, Ian McLoughlin |
| 2023 | ICASSP | Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning. | Qiu-Shi Zhu, Long Zhou, Jie Zhang, Shujie Liu, Yu-Chen Hu, Li-Rong Dai |
| 2023 | ICIG | Vision-Language Adaptive Mutual Decoder for OOV-STR. | Jinshui Hu, Chenyu Liu, Qiandong Yan, Xuyang Zhu, Jiajia Wu, Jun Du, Li-Rong Dai |
| 2023 | ICIG | A Multimodal Text Block Segmentation Framework for Photo Translation. | Jiajia Wu, Anni Li, Kun Zhao, Zhengyan Yang, Bing Yin, Cong Liu, Li-Rong Dai |
| 2023 | ICIG | End-to-End Multilingual Text Recognition Based on Byte Modeling. | Jiajia Wu, Kun Zhao, Zhengyan Yang, Bing Yin, Cong Liu, Li-Rong Dai |
| 2023 | Interspeech | Fine-tuning Audio Spectrogram Transformer with Task-aware Adapters for Sound Event Detection. | Kang Li, Yan Song, Ian McLoughlin, Lin Liu, Jin Li, Li-Rong Dai |
| 2023 | Interspeech | CASA-ASR: Context-Aware Speaker-Attributed ASR. | Mohan Shi, Zhihao Du, Qian Chen, Fan Yu, Yangze Li, Shiliang Zhang, Jie Zhang, Li-Rong Dai |
| 2023 | Interspeech | Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction. | Mohan Shi, Yuchun Shu, Lingyun Zuo, Qian Chen, Shiliang Zhang, Jie Zhang, Li-Rong Dai |
| 2023 | Interspeech | Real-Time Causal Spectro-Temporal Voice Activity Detection Based on Convolutional Encoding and Residual Decoding. | Jingyuan Wang, Jie Zhang, Li-Rong Dai |
| 2023 | Interspeech | Robust Prototype Learning for Anomalous Sound Detection. | Xiao-Min Zeng, Yan Song, Ian McLoughlin, Lin Liu, Li-Rong Dai |
| 2022 | ICASSP | Self-Supervised Representation Learning for Unsupervised Anomalous Sound Detection Under Domain Shift. | Han Chen, Yan Song, Li-Rong Dai, Ian McLoughlin, Lin Liu |
| 2022 | ICASSP | Reference Microphone Selection and Low-Rank Approximation Based Multichannel Wiener Filter with Application to Speech Recognition. | Xing-Yu Chen, Jie Zhang, Li-Rong Dai |
| 2022 | ICASSP | Supervised and Self-Supervised Pretraining Based Covid-19 Detection Using Acoustic Breathing/Cough/Speech Signals. | Xing-Yu Chen, Qiu-Shi Zhu, Jie Zhang, Li-Rong Dai |
| 2022 | ICASSP | Domain Robust Deep Embedding Learning for Speaker Recognition. | Hang-Rui Hu, Yan Song, Ying Liu, Li-Rong Dai, Ian McLoughlin, Lin Liu |
| 2022 | ICASSP | Frontend Attributes Disentanglement for Speech Emotion Recognition. | Yuxuan Xi, Yan Song, Li-Rong Dai, Ian McLoughlin, Lin Liu |
| 2022 | ICASSP | A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech Recognition. | Qiu-Shi Zhu, Jie Zhang, Zi-qiang Zhang, Ming-Hui Wu, Xin Fang, Li-Rong Dai |
| 2022 | Interspeech | Class-Aware Distribution Alignment based Unsupervised Domain Adaptation for Speaker Verification. | Hang-Rui Hu, Yan Song, Li-Rong Dai, Ian McLoughlin, Lin Liu |
| 2022 | Interspeech | Differential Time-frequency Log-mel Spectrogram Features for Vision Transformer Based Infant Cry Recognition. | Hai-tao Xu, Jie Zhang, Li-Rong Dai |
| 2021 | ICASSP | An Effective Deep Embedding Learning Method Based on Dense-Residual Networks for Speaker Verification. | Ying Liu, Yan Song, Ian McLoughlin, Lin Liu, Li-Rong Dai |
| 2021 | ICASSP | An Improved Mean Teacher Based Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection. | Xu Zheng, Yan Song, Ian McLoughlin, Lin Liu, Li-Rong Dai |
| 2021 | Interspeech | Automatic Lip-Reading with Hierarchical Pyramidal Convolution and Self-Attention for Image Sequences with No Word Boundaries. | Hang Chen, Jun Du, Yu Hu, Li-Rong Dai, Bao-Cai Yin, Chin-Hui Lee |
| 2021 | Interspeech | A Weight Moving Average Based Alternate Decoupled Learning Algorithm for Long-Tailed Language Identification. | Hui Wang, Lin Liu, Yan Song, Lei Fang, Ian McLoughlin, Li-Rong Dai |
| 2021 | Interspeech | An Effective Mutual Mean Teaching Based Domain Adaptation Method for Sound Event Detection. | Xu Zheng, Yan Song, Li-Rong Dai, Ian McLoughlin, Lin Liu |
| 2021 | Interspeech | UnitNet-Based Hybrid Speech Synthesis. | Xiao Zhou, Zhen-Hua Ling, Li-Rong Dai |
| 2021 | Interspeech | An Improved Wav2Vec 2.0 Pre-Training Approach Using Enhanced Local Dependency Modeling for Speech Recognition. | Qiu-Shi Zhu, Jie Zhang, Ming-Hui Wu, Xin Fang, Li-Rong Dai |
| 2020 | ICASSP | An Online Speaker-aware Speech Separation Approach Based on Time-domain Representation. | Hui Wang, Yan Song, Zengxi Li, Ian McLoughlin, Li-Rong Dai |
| 2020 | ICASSP | Task-Aware Mean Teacher Method for Large Scale Weakly Labeled Semi-Supervised Sound Event Detection. | Jie Yan, Yan Song, Li-Rong Dai, Ian McLoughlin |
| 2020 | ICASSP | Extracting Unit Embeddings Using Sequence-To-Sequence Acoustic Models for Unit Selection Speech Synthesis. | Xiao Zhou, Zhen-Hua Ling, Li-Rong Dai |
| 2020 | Interspeech | An Effective Speaker Recognition Method Based on Joint Identification and Verification Supervisions. | Ying Liu, Yan Song, Yiheng Jiang, Ian McLoughlin, Lin Liu, Li-Rong Dai |
| 2020 | Interspeech | Recognition-Synthesis Based Non-Parallel Voice Conversion with Adversarial Learning. | Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai |
| 2020 | Interspeech | Semi-Supervised End-to-End ASR via Teacher-Student Learning with Conditional Posterior Distribution. | Zi-qiang Zhang, Yan Song, Jianshu Zhang, Ian McLoughlin, Li-Rong Dai |
| 2020 | Interspeech | An Effective Perturbation Based Semi-Supervised Learning Method for Sound Event Detection. | Xu Zheng, Yan Song, Jie Yan, Li-Rong Dai, Ian McLoughlin, Lin Liu |
| 2019 | ICASSP | A Region Based Attention Method for Weakly Supervised Sound Event Detection and Classification. | Jie Yan, Yan Song, Wu Guo, Li-Rong Dai, Ian McLoughlin, Liang Chen |
| 2019 | ICASSP | Improving Sequence-to-sequence Voice Conversion by Adding Text-supervision. | Jing-Xuan Zhang, Zhen-Hua Ling, Yuan Jiang, Li-Juan Liu, Chen Liang, Li-Rong Dai |
| 2019 | Interspeech | Neural Text Clustering with Document-Level Attention Based on Dynamic Soft Labels. | Zhi Chen, Wu Guo, Li-Rong Dai, Zhen-Hua Ling, Jun Du |
| 2019 | Interspeech | A Chinese Dataset for Identifying Speakers in Novels. | Jia-Xiang Chen, Zhen-Hua Ling, Li-Rong Dai |
| 2019 | Interspeech | Improving Aggregation and Loss Function for Better Embedding Learning in End-to-End Speaker Verification System. | Zhifu Gao, Yan Song, Ian McLoughlin, Pengcheng Li, Yiheng Jiang, Li-Rong Dai |
| 2019 | Interspeech | An Effective Deep Embedding Learning Architecture for Speaker Verification. | Yiheng Jiang, Yan Song, Ian McLoughlin, Zhifu Gao, Li-Rong Dai |
| 2019 | Interspeech | Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling. | Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, Li-Rong Dai |
| 2019 | Interspeech | Multi-Task Learning with High-Order Statistics for x-Vector Based Text-Independent Speaker Verification. | Lanhua You, Wu Guo, Li-Rong Dai, Jun Du |
| 2019 | Interspeech | Deep Neural Network Embeddings with Gating Mechanisms for Text-Independent Speaker Verification. | Lanhua You, Wu Guo, Li-Rong Dai, Jun Du |
| 2018 | ICASSP | Densely Connected Progressive Learning for LSTM-Based Speech Enhancement. | Tian Gao, Jun Du, Li-Rong Dai, Chin-Hui Lee |
| 2018 | ICASSP | Source-Aware Context Network for Single-Channel Multi-Speaker Speech Separation. | Zengxi Li, Yan Song, Li-Rong Dai, Ian McLoughlin |
| 2018 | ICASSP | Forward Attention in Sequence- To-Sequence Acoustic Modeling for Speech Synthesis. | Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai |
| 2018 | Interspeech | WaveNet Vocoder with Limited Training Data for Voice Conversion. | Li-Juan Liu, Zhen-Hua Ling, Yuan Jiang, Ming Zhou, Li-Rong Dai |
| 2018 | Interspeech | Learning and Modeling Unit Embeddings for Improving HMM-based Unit Selection Speech Synthesis. | Xiao Zhou, Zhen-Hua Ling, Zhi-Ping Zhou, Li-Rong Dai |
| 2017 | ASRU | The USTC system for blizzard machine learning challenge 2017-ES2. | Ya-Jun Hu, Li-Juan Liu, Chuang Ding, Zhen-Hua Ling, Li-Rong Dai |
| 2017 | ICASSP | Adaptation of PLDA for multi-source text-independent speaker verification. | Liping Chen, Kong-Aik Lee, Bin Ma, Long Ma, Haizhou Li, Li-Rong Dai |
| 2017 | ICASSP | Extracting structural spectral features using what-where auto-encoders for statistical parametric speech synthesis. | Ya-Jun Hu, Zhen-Hua Ling, Li-Rong Dai |
| 2017 | IJCNN | An investigation of high-resolution modeling units of deep neural networks for acoustic scene classification. | Xiao Bao, Tian Gao, Jun Du, Li-Rong Dai |
| 2017 | Interspeech | Gaussian Prediction Based Attention for Online End-to-End Speech Recognition. | Junfeng Hou, Shiliang Zhang, Li-Rong Dai |
| 2017 | Interspeech | End-to-End Language Identification Using High-Order Utterance Representation with Bilinear Pooling. | Ma Jin, Yan Song, Ian Vince McLoughlin, Wu Guo, Li-Rong Dai |
| 2017 | Interspeech | A Maximum Likelihood Approach to Deep Neural Network Based Nonlinear Spectral Mapping for Single-Channel Speech Separation. | Yannan Wang, Jun Du, Li-Rong Dai, Chin-Hui Lee |
| 2016 | ICASSP | Content-aware local variability vector for speaker verification with short utterance. | Liping Chen, Kong-Aik Lee, Eng Siong Chng, Bin Ma, Haizhou Li, Li-Rong Dai |
| 2016 | ICASSP | Speaker adaptation OF RNN-BLSTM for speech recognition based on speaker code. | Zhiying Huang, Jian Tang, Shaofei Xue, Li-Rong Dai |
| 2016 | ICASSP | Deep belief network-based post-filtering for statistical parametric speech synthesis. | Ya-Jun Hu, Zhen-Hua Ling, Li-Rong Dai |
| 2016 | ICASSP | Modulation spectrum compensation for HMM-based speech synthesis using line spectral pairs. | Zhen-Hua Ling, Xiao-Hui Sun, Li-Rong Dai, Yu Hu |
| 2016 | ICASSP | Compact convolutional neural network transfer learning for small-scale image classification. | Zengxi Li, Yan Song, Ian McLoughlin, Li-Rong Dai |
| 2016 | ICASSP | Modeling spectral envelopes using deep conditional restricted Boltzmann machines for statistical parametric speech synthesis. | Xiang Yin, Zhen-Hua Ling, Ya-Jun Hu, Li-Rong Dai |
| 2016 | Interspeech | The USTC System for Voice Conversion Challenge 2016: Neural Network Based Approaches for Spectrum, Aperiodicity and F | Ling-Hui Chen, Li-Juan Liu, Zhen-Hua Ling, Yuan Jiang, Li-Rong Dai |
| 2016 | Interspeech | SNR-Based Progressive Learning of Deep Neural Network for Speech Enhancement. | Tian Gao, Jun Du, Li-Rong Dai, Chin-Hui Lee |
| 2016 | Interspeech | Speech Bandwidth Extension Using Bottleneck Features and Deep Recurrent Neural Networks. | Yu Gu, Zhen-Hua Ling, Li-Rong Dai |
| 2016 | Interspeech | Articulatory-to-Acoustic Conversion with Cascaded Prediction of Spectral and Excitation Features Using Neural Networks. | Zheng-Chen Liu, Zhen-Hua Ling, Li-Rong Dai |
| 2016 | Interspeech | Future Context Attention for Unidirectional LSTM Based Acoustic Model. | Jian Tang, Shiliang Zhang, Si Wei, Li-Rong Dai |
| 2016 | Interspeech | Compact Feedforward Sequential Memory Networks for Large Vocabulary Continuous Speech Recognition. | Shiliang Zhang, Hui Jiang, Shifu Xiong, Si Wei, Li-Rong Dai |
| 2016 | Interspeech | RNN-BLSTM Based Multi-Pitch Estimation. | Jianshu Zhang, Jian Tang, Li-Rong Dai |
| 2016 | VCIP | Image classification with CNN-based Fisher vector coding. | Yan Song, Xinhai Hong, Ian McLoughlin, Li-Rong Dai |
| 2015 | ACL | The Fixed-Size Ordinally-Forgetting Encoding Method for Neural Network Language Models. | Shiliang Zhang, Hui Jiang, Mingbin Xu, Junfeng Hou, Li-Rong Dai |
| 2015 | ASRU | An information fusion approach to recognizing microphone array speech in the CHiME-3 challenge based on a deep learning framework. | Jun Du, Qing Wang, Yanhui Tu, Xiao Bao, Li-Rong Dai, Chin-Hui Lee |
| 2015 | ICASSP | Channel adaptation of plda for text-independent speaker verification. | Liping Chen, Kong-Aik Lee, Bin Ma, Wu Guo, Haizhou Li, Li-Rong Dai |
| 2015 | ICASSP | Joint training of front-end and back-end deep neural networks for robust speech recognition. | Tian Gao, Jun Du, Li-Rong Dai, Chin-Hui Lee |
| 2015 | ICASSP | Spectral conversion using deep neural networks trained with multi-source speakers. | Li-Juan Liu, Ling-Hui Chen, Zhen-Hua Ling, Li-Rong Dai |
| 2015 | ICASSP | Improved language identification using deep bottleneck network. | Yan Song, Ruilian Cui, Xinhai Hong, Ian McLoughlin, Jiong Shi, Li-Rong Dai |
| 2015 | ICASSP | Speech Separation based on signal-noise-dependent deep neural networks for robust speech recognition. | Yanhui Tu, Jun Du, Li-Rong Dai, Chin-Hui Lee |
| 2015 | ICASSP | Unsupervised speaker adaptation of deep neural network based on the combination of speaker codes and singular value decomposition for speech recognition. | Shaofei Xue, Hui Jiang, Li-Rong Dai, Qingfeng Liu |
| 2015 | ICDAR | Writer adaptive feature extraction based on convolutional neural networks for online handwritten Chinese character recognition. | Jun Du, Jian-Fang Zhai, Jin-Shui Hu, Bo Zhu, Si Wei, Li-Rong Dai |
| 2015 | Interspeech | Phone-centric local variability vector for text-constrained speaker verification. | Liping Chen, Kong-Aik Lee, Bin Ma, Wu Guo, Haizhou Li, Li-Rong Dai |
| 2015 | Interspeech | Automatic phrase boundary labeling of speech synthesis database using context-dependent HMMs and n-gram prior distributions. | Qian Chen, Zhen-Hua Ling, Chen-Yu Yang, Li-Rong Dai |
| 2015 | Interspeech | Deep bottleneck network based i-vector representation for language identification. | Yan Song, Xinhai Hong, Bing Jiang, Ruilian Cui, Ian McLoughlin, Li-Rong Dai |
| 2015 | Interspeech | A universal VAD based on jointly trained deep neural networks. | Qing Wang, Jun Du, Xiao Bao, Zi-Rui Wang, Li-Rong Dai, Chin-Hui Lee |
| 2015 | Interspeech | High-resolution acoustic modeling and compact language modeling of language-universal speech attributes for spoken language identification. | Yannan Wang, Jun Du, Li-Rong Dai, Chin-Hui Lee |
| 2015 | Interspeech | Multi-objective learning and mask-based post-processing for deep neural network based speech enhancement. | Yong Xu, Jun Du, Zhen Huang, Li-Rong Dai, Chin-Hui Lee |
| 2015 | Interspeech | Rectified linear neural networks with tied-scalar regularization for LVCSR. | Shiliang Zhang, Hui Jiang, Si Wei, Li-Rong Dai |
| 2014 | ICASSP | Minimum divergence estimation of speaker prior in multi-session PLDA scoring. | Liping Chen, Kong-Aik Lee, Bin Ma, Wu Guo, Haizhou Li, Li-Rong Dai |
| 2014 | ICASSP | Synthesized stereo mapping via deep neural networks for noisy speech recognition. | Jun Du, Li-Rong Dai, Qiang Huo |
| 2014 | ICASSP | Using bidirectional associative memories for joint spectral envelope modeling in voice conversion. | Li-Juan Liu, Ling-Hui Chen, Zhen-Hua Ling, Li-Rong Dai |
| 2014 | ICASSP | Lattice based optimization of bottleneck feature extractor with linear transformation. | Diyuan Liu, Si Wei, Wu Guo, Yebo Bao, Shifu Xiong, Li-Rong Dai |
| 2014 | ICASSP | Direct adaptation of hybrid DNN/HMM model for fast speaker adaptation in LVCSR based on speaker code. | Shaofei Xue, Ossama Abdel-Hamid, Hui Jiang, Li-Rong Dai |
| 2014 | ICASSP | Spectral modeling using neural autoregressive distribution estimators for statistical parametric speech synthesis. | Xiang Yin, Zhen-Hua Ling, Li-Rong Dai |
| 2014 | ICASSP | Improving deep neural networks for LVCSR using dropout and shrinking structure. | Shiliang Zhang, Yebo Bao, Pan Zhou, Hui Jiang, Li-Rong Dai |
| 2014 | ICASSP | Sequence training of multiple deep neural networks for better performance and faster training speed. | Pan Zhou, Li-Rong Dai, Hui Jiang |
| 2014 | ICFHR | Writer Adaptation Using Bottleneck Features and Discriminative Linear Regression for Online Handwritten Chinese Character Recognition. | Jun Du, Jin-Shui Hu, Bo Zhu, Si Wei, Li-Rong Dai |
| 2014 | ICPR | A Study of Designing Compact Classifiers Using Deep Neural Networks for Online Handwritten Chinese Character Recognition. | Jun Du, Jin-Shui Hu, Bo Zhu, Si Wei, Li-Rong Dai |
| 2014 | Interspeech | Formant-controlled speech synthesis using hidden trajectory model. | Ming-Qi Cai, Zhen-Hua Ling, Li-Rong Dai |
| 2014 | Interspeech | Voice conversion using generative trained deep neural networks with multiple frame spectral envelopes. | Ling-Hui Chen, Zhen-Hua Ling, Li-Rong Dai |
| 2014 | Interspeech | Robust speech recognition with speech enhanced deep neural networks. | Jun Du, Qing Wang, Tian Gao, Yong Xu, Li-Rong Dai, Chin-Hui Lee |
| 2014 | Interspeech | Task-aware deep bottleneck features for spoken language identification. | Bing Jiang, Yan Song, Si Wei, Ian Vince McLoughlin, Li-Rong Dai |
| 2014 | Interspeech | Concept-to-speech generation by integrating syntagmatic features into HMM-based speech synthesis. | Xin Wang, Zhen-Hua Ling, Li-Rong Dai |
| 2014 | Interspeech | Dynamic noise aware training for speech enhancement based on deep neural networks. | Yong Xu, Jun Du, Li-Rong Dai, Chin-Hui Lee |
| 2014 | Interspeech | Modeling DCT parameterized F0 trajectory at intonation phrase level with DNN or decision tree. | Xiang Yin, Ming Lei, Yao Qian, Frank K. Soong, Lei He, Zhen-Hua Ling, Li-Rong Dai |
| 2013 | ICASSP | Incoherent training of deep neural networks to de-correlate bottleneck features for speech recognition. | Yebo Bao, Hui Jiang, Li-Rong Dai, Cong Liu |
| 2013 | ICASSP | Phoneme variation based synthesized speech discrimination for speaker verification. | LianWu Chen, Wu Guo, Yan Song, Li-Rong Dai |
| 2013 | ICASSP | Exemplar based language recognition method for short-duration speech segments. | Meng-Ge Wang, Yan Song, Bing Jiang, Li-Rong Dai, Ian McLoughlin |
| 2013 | ICASSP | Unsupervised prosodic phrase boundary labeling of Mandarin speech synthesis database using context-dependent HMM. | Chen-Yu Yang, Zhen-Hua Ling, Li-Rong Dai |
| 2013 | ICASSP | A cluster-based multiple deep neural networks method for large vocabulary continuous speech recognition. | Pan Zhou, Cong Liu, Qingfeng Liu, Li-Rong Dai, Hui Jiang |
| 2013 | Interspeech | Joint spectral distribution modeling using restricted boltzmann machines for voice conversion. | Ling-Hui Chen, Zhen-Hua Ling, Yan Song, Li-Rong Dai |
| 2012 | Interspeech | Exemplar-Based Sparse Representation for Language Recognition on I-Vectors. | Bing Jiang, Yan Song, Wu Guo, Li-Rong Dai |
| 2012 | Interspeech | Considering Global Variance of the Log Power Spectrum Derived from Mel-Cepstrum in HMM-based Parametric Speech Synthesis. | Xiang Yin, Zhen-Hua Ling, Ming Lei, Li-Rong Dai |
| 2011 | ICASSP | Non-parallel training for voice conversion based on FT-GMM. | Ling-Hui Chen, Zhen-Hua Ling, Li-Rong Dai |
| 2011 | ICASSP | Preserve ordering property of generated LSPS for minimum generation error training in HMM-based speech synthesis. | Ming Lei, Zhen-Hua Ling, Li-Rong Dai |
| 2011 | ICASSP | Speaker characterization using spectral subband energy ratio based on Harmonic plus Noise Model. | Yanhua Long, Zhi-Jie Yan, Frank K. Soong, Li-Rong Dai, Wu Guo |
| 2011 | ICASSP | Building HMM based unit-selection speech synthesis system using synthetic speech naturalness evaluation score. | Heng Lu, Zhen-Hua Ling, Li-Rong Dai, Ren-Hua Wang |
| 2011 | ICASSP | Factored covariance modeling for text-independent speaker verification. | Eryu Wang, Kong-Aik Lee, Bin Ma, Haizhou Li, Wu Guo, Li-Rong Dai |
| 2011 | Interspeech | Estimation of Window Coefficients for Dynamic Feature Extraction for HMM-Based Speech Synthesis. | Ling-Hui Chen, Yoshihiko Nankaku, Heiga Zen, Keiichi Tokuda, Zhen-Hua Ling, Li-Rong Dai |
| 2011 | Interspeech | Formant-Controlled HMM-Based Speech Synthesis. | Ming Lei, Junichi Yamagishi, Korin Richmond, Zhen-Hua Ling, Simon King, Li-Rong Dai |
| 2011 | Interspeech | Improvements in Speaker Characterization Using Spectral Subband Energy Based on Harmonic plus Noise Model. | Yanhua Long, Zhi-Jie Yan, Frank K. Soong, Li-Rong Dai, Wu Guo |
| 2010 | ICASSP | HMM-based pseudo-clean speech synthesis for splice algorithm. | Jun Du, Yu Hu, Li-Rong Dai, Ren-Hua Wang |
| 2010 | ICASSP | N-gram nearest neighbor algorithm for voice password system. | Wu Guo, Zhao Zhang, Yanhua Long, Li-Rong Dai |
| 2010 | ICASSP | Minimum generation error training with weighted Euclidean distance on LSP for HMM-based speech synthesis. | Ming Lei, Zhen-Hua Ling, Li-Rong Dai |
| 2010 | ICASSP | A bounded trust region optimization for discriminative training of HMMS in speech recognition. | Cong Liu, Yu Hu, Hui Jiang, Li-Rong Dai |
| 2010 | Interspeech | A hierarchical F0 modeling method for HMM-based speech synthesis. | Ming Lei, Yi-Jian Wu, Frank K. Soong, Zhen-Hua Ling, Li-Rong Dai |
| 2010 | Interspeech | Global variance modeling on the log power spectrum of LSPs for HMM-based speech synthesis. | Zhen-Hua Ling, Yu Hu, Li-Rong Dai |
| 2010 | Interspeech | Effects of the phonological relevance in speaker verification. | Yanhua Long, Li-Rong Dai, Bin Ma, Wu Guo |
| 2010 | Interspeech | Automatic error detection for unit selection speech synthesis using log likelihood ratio based SVM classifier. | Heng Lu, Zhen-Hua Ling, Si Wei, Li-Rong Dai, Ren-Hua Wang |
| 2010 | Interspeech | HMM based TTS for mixed language text. | Zhiwei Shuang, Shiyin Kang, Yong Qin, Li-Rong Dai, Lianhong Cai |
| 2010 | Interspeech | The estimation and kernel metric of spectral correlation for text-independent speaker verification. | Eryu Wang, Kong-Aik Lee, Bin Ma, Haizhou Li, Wu Guo, Li-Rong Dai |
| 2009 | ICASSP | iFLY system for the NIST 2008 speaker recognition evaluation. | Wu Guo, Yanhua Long, Yijie Li, Lei Pan, Eryu Wang, Li-Rong Dai |
| 2009 | ICASSP | The I4U system in NIST 2008 speaker recognition evaluation. | Haizhou Li, Bin Ma, Kong-Aik Lee, Hanwu Sun, Donglai Zhu, Khe Chai Sim, Changhuai You, Rong Tong, Ismo Krkkinen, Chien-Lin Huang, Vladimir Pervouchine, Wu Guo, Yijie Li, Li-Rong Dai, Mohaddeseh Nosratighods, Tharmarajah Thiruvaran, Julien Epps, Eliathamby Ambikairajah, Chng Eng Siong, Tanja Schultz, Qin Jin |
| 2009 | ICASSP | Exploiting prosodic information for Speaker Recognition. | Yanhua Long, Bin Ma, Haizhou Li, Wu Guo, Chng Eng Siong, Li-Rong Dai |
| 2009 | ICASSP | Full covariance state duration modeling for HMM-based speech synthesis. | Heng Lu, Yi-Jian Wu, Keiichi Tokuda, Li-Rong Dai, Ren-Hua Wang |
| 2009 | Interspeech | Asynchronous F0 and spectrum modeling for HMM-based speech synthesis. | Cheng-Cheng Wang, Zhen-Hua Ling, Li-Rong Dai |
| 2008 | ICASSP | Minumum generation error linear regression based model adaptation for HMM-based speech synthesis. | Long Qin, Yi-Jian Wu, Zhen-Hua Ling, Ren-Hua Wang, Li-Rong Dai |
| 2008 | ICASSP | Minimum generation error criterion considering global/local variance for HMM-based speech synthesis. | Long Qin, Yi-Jian Wu, Zhen-Hua Ling, Ren-Hua Wang, Li-Rong Dai |
| 2007 | ICASSP | An Interactive Video Annotation Frameowrk with Multiple Modalities. | Meng Wang, Xian-Sheng Hua, Yan Song, Li-Rong Dai, Ren-Hua Wang |
| 2007 | MMM | An Efficient Automatic Video Shot Size Annotation Scheme. | Meng Wang, Xian-Sheng Hua, Yan Song, Wei Lai, Li-Rong Dai, Ren-Hua Wang |
| 2006 | CVPR | Video Annotation by Active Learning and Cluster Tuning. | Guo-Jun Qi, Yan Song, Xian-Sheng Hua, Hong-Jiang Zhang, Li-Rong Dai |
| 2006 | ICASSP | An Automatic Video Semantic Annotation Scheme Based on Combination of Complementary Predictors. | Yan Song, Xian-Sheng Hua, Li-Rong Dai, Meng Wang, Ren-Hua Wang |
| 2006 | ICDM | Semi-Supervised Kernel Regression. | Meng Wang, Xian-Sheng Hua, Yan Song, Li-Rong Dai, HongJiang Zhang |
| 2006 | ISCAS | Automatic video annotation based on co-adaptation and label correction. | Meng Wang, Xian-Sheng Hua, Yan Song, Li-Rong Dai, Shipeng Li |
| 2005 | ICASSP | Sliding Window Smoothing For Maximum Entropy Based Intonational Phrase Prediction In Chinese. | Jianfeng Li, Guoping Hu, Ren-Hua Wang, Li-Rong Dai |
| 2005 | ICASSP | An Improved Spectral and Prosodic Transformation Method in STRAIGHT-based Voice Conversion. | Long Qin, Gao Peng Chen, Zhen-Hua Ling, Li-Rong Dai |
| 2004 | ICASSP | A complexity reduction of ETSI advanced front-end for DSR. | Jin-Yu Li, Bo Liu, Ren-Hua Wang, Li-Rong Dai |
| 2004 | ICIP | A region based multiple frame-rate tradeoff of video streaming. | Wei Lai, Xiaodong Gu, Ren-Hua Wang, Li-Rong Dai, HongJiang Zhang |