IEEE Automatic Speech Recognition and Understanding Workshop
ASRU
C
CORE rank
CORE rank (raw)
C
Fields of research
Artificial Intelligence
Papers indexed
1,323
2007–2025
Papers per year
2007209 peak2025
Most published authors
ASRU papers
1,323 records sourced from DBLP. Search titles, filter by year, sort by recency.
| Year | Title | Authors |
|---|---|---|
| 2025 | Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM. | Chiori Hori, Yoshiki Masuyama, Siddarth Jain, Radu Corcodel, Devesh K. Jha, Diego Romeres, Jonathan Le Roux |
| 2025 | Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization? | Shota Horiguchi, Naohiro Tawara, Takanori Ashihara, Atsushi Ando, Marc Delcroix |
| 2025 | EmoTale: An Enacted Speech-emotion Dataset in Danish. | Maja J. Hjuler, Harald V. Skat-Rrdam, Line H. Clemmensen, Sneha Das |
| 2025 | PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation. | Jiajun He, Naoki Sawada, Koichi Miyazaki, Tomoki Toda |
| 2025 | SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compresison in Speaker Verification. | Jungwoo Heo, Hyun-seo Shin, Chan-yeong Lim, Kyo-Won Koo, Seung-bin Kim, Jisoo Son, Ha-Jin Yu |
| 2025 | ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching. | Zhu Han, Wei Kang, Zengwei Yao, Liyong Guo, Fangjun Kuang, Zhaoqing Li, Weiji Zhuang, Long Lin, Daniel Povey |
| 2025 | EmoBiMamba-TTS: Bidirectional State Space Model for Emotion-Intensity Controllable Text-to-Speech. | Insung Ham, Bonwha Ku, Hanseok Ko |
| 2025 | ASTAR-NTU solution to AudioMOS Challenge 2025 Track1. | Fabian Ritter Gutierrez, Yi-Cheng Lin, Jui-Chiang Wei, Jeremy H. M. Wong, Nancy F. Chen, Hung-Yi Lee |
| 2025 | A correlation-permutation approach for speech-music encoders model merging. | Fabian Ritter Gutierrez, Yi-Cheng Lin, Jeremy H. M. Wong, Hung-Yi Lee, Eng Siong Chng, Nancy F. Chen |
| 2025 | Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation. | Hongming Guo, Ruibo Fu, Yizhong Geng, Shuchen Shi, Tao Wang, Chunyu Qiang, Ya Li, Zhengqi Wen, Yukun Liu, Xuefei Liu, Chenxing Li |
| 2025 | Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting. | Ramesh Gundluru, Shubham Gupta, K. Sri Rama Murty |
| 2025 | On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts. | Kashaf Gulzar, Dominik Wagner, Sebastian P. Bayerl, Florian Hnig, Tobias Bocklet, Korbinian Riedhammer |
| 2025 | Omni-Router: Sharing Routing Decisions in Sparse Mixture-of-Experts for Speech Recognition. | Zijin Gu, Tatiana Likhomanenko, Navdeep Jaitly |
| 2025 | Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement. | Heitor R. Guimares, Ke Tan, Juan Azcarreta, Jesus Alvarez, Prabhav Agrawal, Ashutosh Pandey, Buye Xu |
| 2025 | FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities. | Lilit Grigoryan, Vladimir Bataev, Nikolay Karpov, Andrei Andrusenko, Vitaly Lavrukhin, Boris Ginsburg |
| 2025 | Graph Connectionist Temporal Classification for Phoneme Recognition. | Henry Graf, Hugo Van hamme |
| 2025 | Post-training for Deepfake Speech Detection. | Wanying Ge, Xin Wang, Xuechen Liu, Junichi Yamagishi |
| 2025 | Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts. | Ashi Garg, Zexin Cai, Henry Li Xinyuan, Leibny Paola Garca-Perera, Sanjeev Khudanpur, Matthew Wiesner, Nicholas Andrews |
| 2025 | WST: Weakly Supervised Transducer for Automatic Speech Recognition. | Dongji Gao, Chenda Liao, Changliang Liu, Matthew Wiesner, Leibny Paola Garca-Perera, Daniel Povey, Sanjeev Khudanpur, Jian Wu |
| 2025 | Predictive ASR and Turn-taking Prediction at Once: Towards More Responsive Spoken Dialog System. | Ryo Fukuda, Takatomo Kano, Naohiro Tawara, Marc Delcroix, Atsunori Ogawa, Yuya Chiba, Atsushi Ando |
| 2025 | Multi-Target Backdoor Attacks Against Speaker Recognition. | Alexandrine Fortier, Sonal Joshi, Thomas Thebaud, Jess Antonio Villalba Lpez, Najim Dehak, Patrick Cardinal |
| 2025 | Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning. | Yangui Fang, Jing Peng, Xu Li, Yu Xi, Chengwei Zhang, Guohui Zhong, Kai Yu |
| 2025 | Beyond Modality Limitations: A Unified MLLM Approach to Automated Speaking Assessment with Effective Curriculum Learning. | Yu-Hsuan Fang, Tien-Hong Lo, Yao-Ting Sung, Berlin Chen |
| 2025 | Fewer Hallucinations, More Verification: A Three-Stage LLM-Based Framework for ASR Error Correction. | Yangui Fang, Baixu Cheng, Jing Peng, Xu Li, Yu Xi, Chengwei Zhang, Guohui Zhong |
| 2025 | WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction. | Jakaria Islam Emon, Md Abu Salek, Kazi Tamanna Alam |
151–175 of 1,323← PreviousNext →
Comparable venues
Other A*/A conferences filed under the same field of research.
- A*AAAINational Conference of the American Association for Artificial Intelligence
- A*ICRAIEEE International Conference on Robotics and Automation
- AInterspeechInterspeech (combined EuroSpeech and ICSLP in 2000)
- AIROSIEEE/RSJ International Conference on Intelligent Robots and Systems
- A*ACLAssociation for Computational Linguistics
- A*IJCAIInternational Joint Conference on Artificial Intelligence
- A*EMNLPEmpirical Methods in Natural Language Processing
- AGECCOGenetic and Evolutionary Computations