Skip to content

Interspeech (combined EuroSpeech and ICSLP in 2000)

Interspeech

A

CORE rank

CORE rank (raw)

A

Acceptance rate

48.0% (2024)

Fields of research

Artificial Intelligence

Papers indexed

28,540

1987–2025

Papers per year

19871,180 peak2025

Interspeech papers

28,540 records sourced from DBLP. Search titles, filter by year, sort by recency.

YearTitleAuthors
2025Meta-Learning Approaches for Speaker-Dependent Voice Fatigue Models.Roseline Polle, Agnes Norbury, Alexandra Livia Georgescu, Nicholas Cummins, Stefano Goria
2025Constrained LDDMM for Dynamic Vocal Tract Morphing: Integrating Volumetric and Real-Time MRI.Tharinda Piyadasa, Joan Glauns, Amelia Gully, Michael Proctor, Kirrie J. Ballard, Tnde Szalay, Naeim Sanaei, Sheryl Foster, David Waddington, Craig T. Jin
2025Multimodal Assessment of Speech Impairment in Amyotrophic Lateral Sclerosis Using Audio-Visual and Machine Learning Approaches.Francesco Pierotti, Andrea Bandini
2025Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models.Tuan Dat Phuong, Long-Vu Hoang, Huy Dat Tran
2025Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches.Bornali Phukon, Xiuwen Zheng, Mark Hasegawa-Johnson
2025Towards Machine Unlearning for Paralinguistic Speech Processing.Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Shubham Singh, Swarup Ranjan Behera, Vandana Rajan, Muskaan Singh, Arun Balaji Buduru, Rajesh Sharma
2025Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction.Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma
2025HYFuse: Aligning Heterogeneous Speech Pre-Trained Representations in Hyperbolic Space for Speech Emotion Recognition.Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma
2025Towards Fusion of Neural Audio Codec-based Representations with Spectral for Heart Murmur Classification via Bandit-based Cross-Attention Mechanism.Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Priyabrata Mallick, Santanu Roy, Arun Balaji Buduru, Rajesh Sharma
2025Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models.Orchid Chetia Phukan, Girish, Mohd Mujtaba Akhtar, Swarup Ranjan Behera, Priyabrata Mallick, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma
2025SNIFR : Boosting Fine-Grained Child Harmful Content Detection Through Audio-Visual Alignment with Cascaded Cross-Transformer.Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Abu Osama Siddiqui, Sarthak Jain, Priyabrata Mallick, Jaya Sai Kiran Patibandla, Pailla Balakrishna Reddy, Arun Balaji Buduru, Rajesh Sharma
2025PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition.Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera, Jaya Sai Kiran Patibandla, Arun Balaji Buduru, Rajesh Sharma
2025Model as Loss: A Self-Consistent Training Paradigm.Saisamarth Rajesh Phaye, Milos Cernak, Andrew Harper
2025Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish.Nhan Phan, Mikko Kuronen, Maria Kautonen, Riikka Ullakonoja, Anna von Zansen, Yaroslav Getman, Ekaterina Voskoboinik, Tams Grsz, Mikko Kurimo
2025RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling.Long-Khanh Pham, Thanh V. T. Tran, Minh-Tan Pham, Van Nguyen
2025Fifteen Years of Child-Centered Long-Form Recordings: Promises, Resources, and Remaining Challenges to Validity.Loann Peurey, Marvin Lavechin, Tarek Kunze, Manel Khentout, Lucas Gautheron, Emmanuel Dupoux, Alejandrina Cristi
2025Automatic Detection and Sub-typing of Primary Progressive Aphasia from Speech: Integrating Task-Specific Features and Spatio-Semantic Graphs.Fritz Peters, W. Richard Bevan-Jones, Grace Threlfall, Jenny M. Harris, Julie S. Snowden, Matthew Jones, Jennifer C. Thompson, Daniel J. Blackburn, Heidi Christensen
2025EnCodecMAE: leveraging neural codecs for universal audio representation learning.Leonardo Pepino, Pablo Riera, Luciana Ferrer
2025Parameter-Efficient Fine-tuning with Instance-Aware Prompt and Parallel Adapters for Speaker Verification.Shengyu Peng, Wu Guo, Jie Zhang, Yu Guan, Lipeng Dai, Zuoliang Li
2025FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems.Yizhou Peng, Yi-Wen Chao, Dianwen Ng, Yukun Ma, Chongjia Ni, Bin Ma, Eng Siong Chng
2025Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation.Varsha Pendyala, Pedro Morgado, William A. Sethares
2025Optimized Real-time Speech Enhancement with Deep SSMs on Raw Audio.Yan Ru Pei, Ritik Shrivastava, Sidharth
2025TinyClick: Single-Turn Agent for Empowering GUI Automation.Pawel Pawlowski, Krystian Zawistowski, Wojciech Lapacz, Adam Wiacek, Marcin Skorupa, Sebastien Postansque, Jakub Hoscilowicz
2025Evaluating the suitability of acoustic parameters for capturing breathy voice in non-pathological female speakers.Chloe Patman, Paul Foulkes, Kirsty McDougall
2025Kinship in Speech: Leveraging Linguistic Relatedness for Zero-Shot TTS in Indian Languages.Utkarsh Pathak, Chandra Sai Krishna Gunda, Anusha Prakash, Keshav Agarwal, Hema A. Murthy
401425 of 28,540← PreviousNext →

Comparable venues

Other A*/A conferences filed under the same field of research.