| 2026 | ACL | Discovering and Causally Validating Emotion-Sensitive Neurons in Large Audio-Language Models. | Xiutian Zhao, Bjrn W. Schuller, Berrak Sisman |
| 2025 | EMNLP | Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech Synthesis. | Zhenqi Jia, Rui Liu, Berrak Sisman, Haizhou Li |
| 2025 | Interspeech | Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset. | Rui Liu, Pu Gao, Jiatian Xi, Berrak Sisman, Carlos Busso, Haizhou Li |
| 2025 | Interspeech | EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast. | Shreeram Suresh Chandra, Lucas Goncalves, Junchen Lu, Carlos Busso, Berrak Sisman |
| 2025 | Interspeech | Can Emotion Fool Anti-spoofing? | Aurosweta Mahapatra, Ismail Rasim Ulgen, Abinay Reddy Naini, Carlos Busso, Berrak Sisman |
| 2025 | Interspeech | The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions. | Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ulgen, Thomas Thebaud, Laureano Moro-Velzquez, Leibny Paola Garca, Najim Dehak, Berrak Sisman, Carlos Busso |
| 2025 | Interspeech | Advancing Pediatric ASR: The Role of Voice Generation in Disordered Speech. | Karen Rosero, Ali N. Salman, Shreeram Suresh Chandra, Berrak Sisman, Cortney Van't Slot, Alex A. Kane, Rami R. Hallac, Carlos Busso |
| 2024 | ICASSP | Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition. | Ismail Rasim Ulgen, Zongyang Du, Carlos Busso, Berrak Sisman |
| 2024 | Interspeech | Unsupervised Domain Adaptation for Speech Emotion Recognition using K-Nearest Neighbors Voice Conversion. | Pravin Mote, Berrak Sisman, Carlos Busso |
| 2024 | Interspeech | Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline. | Ali N. Salman, Zongyang Du, Shreeram Suresh Chandra, Ismail Rasim lgen, Carlos Busso, Berrak Sisman |
| 2024 | Tencon | SNIPER Training: Single-Shot Sparse Training for Text-to-Speech. | Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman, Dorien Herremans |
| 2024 | Tencon | Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder. | Jan Melechovsk, Ambuj Mehrish, Berrak Sisman, Dorien Herremans |
| 2024 | Tencon | Accent Conversion in Text-to-Speech Using Multi-Level VAE and Adversarial Training. | Jan Melechovsk, Ambuj Mehrish, Berrak Sisman, Dorien Herremans |
| 2023 | Interspeech | SlothSpeech: Denial-of-service Attack Against Speech Recognition Models. | Mirazul Haque, Rutvij Shah, Simin Chen, Berrak Sisman, Cong Liu, Wei Yang |
| 2023 | Interspeech | High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units. | Junchen Lu, Berrak Sisman, Mingyang Zhang, Haizhou Li |
| 2022 | ICASSP | Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over. | Junchen Lu, Berrak Sisman, Rui Liu, Mingyang Zhang, Haizhou Li |
| 2022 | Interspeech | Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning. | Rui Liu, Berrak Sisman, Bjrn W. Schuller, Guanglai Gao, Haizhou Li |
| 2022 | Interspeech | Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion. | Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li |
| 2022 | Interspeech | EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models. | Perry Lam, Huayun Zhang, Nancy F. Chen, Berrak Sisman |
| 2021 | ASRU | Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style Transfer. | Zongyang Du, Berrak Sisman, Kun Zhou, Haizhou Li |
| 2021 | ASRU | DEEPA: A Deep Neural Analyzer for Speech and Singing Vocoding. | Sergey Nikonorov, Berrak Sisman, Mingyang Zhang, Haizhou Li |
| 2021 | ICASSP | Graphspeech: Syntax-Aware Graph Attention Network for Neural Speech Synthesis. | Rui Liu, Berrak Sisman, Haizhou Li |
| 2021 | ICASSP | Seen and Unseen Emotional Style Transfer for Voice Conversion with A New Emotional Speech Dataset. | Kun Zhou, Berrak Sisman, Rui Liu, Haizhou Li |
| 2021 | Interspeech | Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion Discriminability. | Rui Liu, Berrak Sisman, Haizhou Li |
| 2021 | Interspeech | Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-Stage Sequence-to-Sequence Training. | Kun Zhou, Berrak Sisman, Haizhou Li |
| 2021 | SIGdial | Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue. | Haizhou Li, Gina-Anne Levow, Zhou Yu, Chitralekha Gupta, Berrak Sisman, Siqi Cai, David Vandyke, Nina Dethlefs, Yan Wu, Junyi Jessy Li |
| 2020 | ICASSP | Teacher-Student Training For Robust Tacotron-Based TTS. | Rui Liu, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, Haizhou Li |
| 2020 | Interspeech | Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion. | Kun Zhou, Berrak Sisman, Mingyang Zhang, Haizhou Li |
| 2019 | ASRU | On the Study of Generative Adversarial Networks for Cross-Lingual Voice Conversion. | Berrak Sisman, Mingyang Zhang, Minghui Dong, Haizhou Li |
| 2019 | Interspeech | VQVAE Unsupervised Unit Discovery and Multi-Scale Code2Spec Inverter for Zerospeech Challenge 2019. | Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, Satoshi Nakamura |
| 2018 | Interspeech | Wavelet Analysis of Speaker Dependent and Independent Prosody for Voice Conversion. | Berrak Sisman, Haizhou Li |
| 2018 | Interspeech | A Voice Conversion Framework with Tandem Feature Sparse Representation and Speaker-Adapted WaveNet Vocoder. | Berrak Sisman, Mingyang Zhang, Haizhou Li |
| 2017 | ASRU | Sparse representation of phonetic features for voice conversion with and without parallel data. | Berrak Sisman, Haizhou Li, Kay Chen Tan |
| 2016 | WCNC | Energy and data cooperation in energy harvesting multiple access channel. | Berk Gurakan, Berrak Sisman, Onur Kaya, Sennur Ulukus |
| 2016 | WCNC | Energy and data cooperation in energy harvesting multiple access channel. | Berk Gurakan, Berrak Sisman, Onur Kaya, Sennur Ulukus |